Audio generation method, device, computer equipment, storage medium and product

By obtaining chord arrays and using neural network models to automatically generate pitch arrays and chord arrays, the existing LoFi music production methods are solved, and the automation and efficiency of audio production is achieved.

CN116168668BActive Publication Date: 2025-05-09TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310150317.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-09
Publication Date
2025-05-09
Estimated Expiration
2043-02-09

AI Technical Summary

Technical Problem

The existing LoFi music production methods rely on manual synthesis, which are cumbersome and have low efficiency, making it difficult to meet users' needs for convenient and efficient audio production.

Method used

By obtaining the first chord array and generating multiple pitch arrays and second chord arrays associated with it, the target audio is generated. This method uses a neural network model to automatically generate pitch arrays and chord arrays, and arranges and processes them in combination with users' personalized needs to realize automatic audio generation.

Benefits of technology

It realizes the automation and intelligence of audio production, improves the efficiency and flexibility of audio production, and meets users' needs for convenient and efficient audio production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116168668B_ABST
    Figure CN116168668B_ABST
Patent Text Reader

Abstract

The present application proposes an audio generation method, device, computer equipment, storage medium and product. The method comprises: obtaining a first chord array, and generating k pitch arrays and k second chord arrays associated with the first chord array, where k is a positive integer; arranging and processing the k pitch arrays to generate first track data; generating second track data based on the k second chord arrays and the first track data; generating target audio based on the first track data and the second track data. The present application can conveniently perform audio production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an audio generation method, an audio generation device, a computer equipment, a computer-readable storage medium, and a computer program product. Background Art

[0002] As the music market flourishes, slow-paced pure music is gradually returning to the mainstream music market, and users are increasingly demanding LoFi (Low Fidelity) music, which is widely used in focused scenarios such as study and work. LoFi music can not only help users block out external noise, but its regular rhythm can also easily put the brain into a cycle, reducing the possibility of distraction. Secondly, slow music and repetitive and predictable rhythms can relax people, while frequent drum beats can stimulate hearing, so it is loved by the general public.

[0003] At present, the production of LoFi music mainly relies on manual synthesis, which is cumbersome and inefficient. Summary of the invention

[0004] The embodiments of the present application propose an audio generation method, device, equipment, medium, and product, which can conveniently perform audio production.

[0005] On the one hand, an embodiment of the present application provides an audio generation method, the method comprising:

[0006] Get a first chord array, and generate k pitch arrays and k second chord arrays associated with the first chord array, where k is a positive integer;

[0007] Arrange and process the k pitch arrays to generate first audio track data;

[0008] According to the audio duration of the first audio track data, generating second audio track data matching the audio duration of the first audio track data by the k second chord arrays;

[0009] Generate target audio according to the first audio track data and the second audio track data.

[0010] In one aspect, an embodiment of the present application provides an audio generation device, the device comprising:

[0011] An acquisition unit, used for acquiring a first chord array, and generating k pitch arrays and k second chord arrays associated with the first chord array, where k is a positive integer;

[0012] A processing unit, used for arranging and processing the k pitch arrays to generate first audio track data;

[0013] The processing unit is further used to generate second track data according to the k second chord arrays and the first track data;

[0014] The processing unit is further used to generate target audio according to the first audio track data and the second audio track data.

[0015] In a possible implementation, the processing unit generates k pitch arrays and k second chord arrays associated with the first chord array, for performing the following operations:

[0016] Inputting the first chord array into a neural network model, and running the neural network model k times to obtain k pitch arrays and k second chord arrays associated with the first chord array;

[0017] Wherein, each time the neural network model is run, a pitch array and a second chord array are generated.

[0018] In a possible implementation, the processing unit arranges the k pitch arrays to generate first track data for performing the following operations:

[0019] Arrange the k pitch arrays according to M musical form structures to obtain main melody audio data, where M is a positive integer;

[0020] Audio attribute information is set for the main melody audio data, and the main melody audio data with the audio attribute information is determined as the first track data; wherein the audio attribute information includes: any one or more of instrument information, number of beats, and rhythm texture.

[0021] In a possible implementation, the processing unit generates second track data according to the k second chord arrays and the first track data, and is used to perform the following operations:

[0022] Obtaining the audio duration of the first audio track data;

[0023] According to the audio duration of the first audio track data, the k second chord arrays are cyclically copied to obtain second audio track data; wherein the audio duration of the second audio track data matches the audio duration of the first audio track data.

[0024] In a possible implementation, the processing unit is further configured to perform the following operations:

[0025] According to the number of beats corresponding to each chord in the second chord array, adapting the duration of each note included in each chord in the second audio track data;

[0026] Performing rhythmic arrangement processing on each note after the duration adaptation processing according to the note pitch information of each chord in the second chord array to obtain third track data;

[0027] The processing unit generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0028] Generate target audio according to the first audio track data, the second audio track data and the third audio track data.

[0029] In a possible implementation, the processing unit is further configured to perform the following operations:

[0030] Acquire target drum beat data selected from a drum beat database, and combine and process the target drum beat data according to preset rhythm pattern data to obtain fourth track data;

[0031] The processing unit generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0032] Generate target audio according to the first audio track data, the second audio track data and the fourth audio track data.

[0033] In a possible implementation, the processing unit is further configured to perform the following operations:

[0034] Acquire target background audio selected from a background audio database, and obtain fifth audio track data according to the target background audio;

[0035] The processing unit generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0036] Generate target audio according to the first audio track data, the second audio track data and the fifth audio track data.

[0037] In a possible implementation, the processing unit is further configured to perform the following operations:

[0038] Acquire a target note of the lowest pitch information from each chord included in the second audio track data, and obtain sixth audio track data according to the acquired multiple target notes;

[0039] The processing unit generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0040] Generate target audio according to the first audio track data, the second audio track data and the sixth audio track data.

[0041] In a possible implementation, the processing unit generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0042] Performing audio distortion processing according to the first audio track data and the second audio track data to obtain audio track post-processing data;

[0043] Performing reverberation rendering processing on the audio track post-processing data to obtain audio track reference data;

[0044] The reference data of each audio track are merged to generate target audio.

[0045] In a possible implementation, the processing unit generates target audio according to the plurality of audio track reference data, and is used to perform the following operations:

[0046] Merging the multiple audio track reference data to obtain total audio data to be processed;

[0047] The total audio data to be processed is subjected to low-fidelity simulation processing to obtain target audio.

[0048] On the one hand, an embodiment of the present application provides a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the above-mentioned audio generation method.

[0049] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is read and executed by a processor of a computer device, the computer device executes the above-mentioned audio generation method.

[0050] On the one hand, the embodiments of the present application provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned audio generation method.

[0051] In an embodiment of the present application, first, a first chord array can be obtained, and k pitch arrays and k second chord arrays associated with the first chord array can be generated, where k is a positive integer; then, the k pitch arrays can be arranged and processed to generate first track data; and, based on the audio duration of the first track data, second track data that matches the audio duration of the first track data can be generated by the k second chord arrays; finally, target audio can be generated based on the first track data and the second track data. It can be seen that in an embodiment of the present application, multiple pitch arrays associated with the first chord information can be generated, and information arrangement processing can be performed on the pitch array and the second chord array, thereby meeting more flexible audio production requirements, and the entire audio generation process is automated, meeting the automation and intelligent requirements of audio production, and improving the efficiency of the audio production process. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0053] Figure 1 It is a schematic diagram of the architecture of an audio generation system provided in an embodiment of the present application;

[0054] Figure 2 It is a flowchart of an audio generation method provided in an embodiment of the present application;

[0055] Figure 3 It is a structural schematic diagram of a chord array provided in an embodiment of the present application;

[0056] Figure 4 It is a schematic diagram of a process of calling a neural network model provided in an embodiment of the present application;

[0057] Figure 5 It is a schematic diagram of a process of arranging a pitch array provided in an embodiment of the present application;

[0058] Figure 6 is a flowchart of another audio generation method provided in an embodiment of the present application;

[0059] Figure 7 is a schematic diagram of rhythmic data provided by an embodiment of the present application;

[0060] Figure 8 is a schematic diagram of generating bass track data provided by an embodiment of the present application;

[0061] Fig. 9It is a schematic diagram of a process of generating target audio provided in an embodiment of the present application;

[0062] Fig.10 is a structural schematic diagram of an audio generating device provided in an embodiment of the present application;

[0063] Fig.11 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0064] Here, exemplary embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application.

[0065] The embodiment of the present application proposes an audio generation scheme that supports custom selection of audio information such as chord arrays, rhythmic textures, number of beats (Beats Per Minute, BPM), and instrument information. The audio production process is more flexible, convenient and automated, which can improve the efficiency of audio production. Among them, the principle of the audio generation scheme mainly includes: first, a first chord array can be obtained, and then a neural network can be called to generate pitch array information associated with the first chord array, and a second chord array that has been rearranged and combined with the first chord array. The pitch array information and the second chord array generated by the neural network can be used as audio materials such as the main melody (pitch array information) and accompaniment (second chord array) for generating audio. Among them, the pitch array information generated by the neural network can be stored based on the music format of a MIDI (Music Instrument Digital Interface) file. Next, the pitch array information and the second chord array can be customized according to the personalized needs of the user, so as to automatically adapt the MIDI file. The adaptation process is used to indicate that the audio signals of multiple audio tracks are to be arranged and processed, so as to obtain the audio signals of multiple audio tracks (hereinafter also referred to as track data). Finally, the target audio is generated according to the track data of multiple audio tracks.

[0066] It can be seen that in this application, audio materials such as the main melody and accompaniment can be automatically generated based on the neural network; then, in the process of arranging the pitch array information and the second chord array, the audio attributes (such as BPM, chord array, rhythm texture, music information, etc.) can be adapted accordingly according to the personalized needs of the user, so as to meet more flexible audio production needs. In addition, the entire audio generation process is automated, thereby improving the efficiency of the audio production process.

[0067] Next, the relevant technical terms and main application scenarios involved are introduced in detail in combination with the relevant principles of the audio generation solution provided by this application:

[0068] 1. Artificial Intelligence

[0069] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence.

[0070] In one possible implementation, the present application can be combined with machine learning technology in the field of artificial intelligence. Specifically, machine learning technology can be used to train a neural network model, thereby calling the neural network model to generate k pitch arrays and k second chord arrays associated with the first chord array. Among them, the so-called machine learning (ML) is a multi-field cross-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and formula-based learning.

[0071] 2. Cloud Technology:

[0072] Cloud computing refers to the delivery and use model of IT infrastructure, which means obtaining required resources through the network in an on-demand and easily scalable manner; in a broad sense, cloud computing refers to the delivery and use model of services, which means obtaining required services through the network in an on-demand and easily scalable manner. This service can be related to IT and software, the Internet, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing.

[0073] In the present application, "k pitch arrays are arranged and processed to generate first audio track data; and, based on the audio duration of the first audio track data, second audio track data that matches the audio duration of the first audio track data is generated by k second chord arrays." The above process involves large-scale calculations and requires large computing power and storage space. Therefore, in one possible implementation of the present application, a computer device can obtain sufficient computing power and storage space through cloud computing technology to execute the specific process of generating the target audio involved in the present application.

[0074] It is particularly important to note that in the subsequent specific implementation methods of this application, related data such as object information is involved. When the above embodiments of this application are applied to specific products or technologies, it is necessary to obtain the object's permission or consent, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0075] Next, the architecture diagram of the audio generation system involved in this application is described accordingly. Figure 1 , Figure 1 Schematic diagram of the architecture of an audio generation system provided in an embodiment of the present application. Figure 1 As shown, the system architecture diagram may include at least: a server 104 and a terminal device cluster, wherein the terminal device cluster may include at least: a terminal device 101, a terminal device 102, a terminal device 103, etc. Any terminal device in the terminal device cluster may be directly or indirectly connected to the server 104 via wired or wireless communication, and this application does not limit this.

[0076] in, Figure 1The server 104 shown can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.

[0077] Figure 1 Any of the terminal devices shown can be a mobile phone, a tablet computer, a laptop computer, a PDA, a mobile Internet device (MID), a vehicle, a vehicle-mounted device, a roadside device, an aircraft, a wearable device such as a smart watch, a smart bracelet, a pedometer, etc., and other smart devices with audio generation capabilities.

[0078] In a possible implementation, taking the terminal device 101 as an example, the audio generation scheme provided in the embodiment of the present application is further elaborated. Specifically, when it is necessary to make audio, the terminal device 101 can obtain a first chord array. Then the terminal device 101 can send the first chord array to the server 104, and the server 104 can generate k pitch arrays and k second chord arrays associated with the first chord array, where k is a positive integer. Then, the server 104 can arrange and process the k pitch arrays to generate the first track data. And, the server 104 generates the second track data matching the audio duration of the first track data by the k second chord arrays according to the audio duration of the first track data. Finally, the server 104 can generate the target audio according to the first track data and the second track data. Subsequently, the server 104 can send the target audio to the terminal device 101 so that the terminal device 101 plays the target audio in the audio client.

[0079] It should be understood that the above is only an exemplary description of the specific operations performed by the terminal device 101 and the server 104. In another possible implementation, the above-mentioned audio generation scheme can also be executed separately by the server in the audio generation system or any terminal device in the terminal device cluster, and the embodiment of the present application does not specifically limit this.

[0080] In a possible implementation, the audio generation system provided in the embodiment of the present application can be deployed on a blockchain. For example, the terminal device 101, the terminal device 102, and the server 103 can all be regarded as node devices of the blockchain to jointly constitute a blockchain network. Therefore, the audio generation process in the present application can be executed on the blockchain, which can not only ensure the fairness and justice of the audio generation process, but also make the audio generation process traceable, thereby improving the security of the audio generation process.

[0081] It can be understood that the system architecture diagram described in the embodiment of the present application is for more clearly illustrating the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided in the embodiment of the present application. A person of ordinary skill in the art can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.

[0082] Based on the above description of the audio generation scheme and the audio generation system, the present application embodiment proposes an audio generation method. Figure 2 As shown, Figure 2 is a flow chart of an audio generation method provided in an embodiment of the present application. The audio generation method can be performed by the above Figure 1 The terminal device or server in the audio generation system mentioned above is executed. For the sake of convenience, the present application embodiment takes the execution of a computer device as an example. The audio generation method may include the following steps S201 to S204:

[0083] S201. Obtain a first chord array, and generate k pitch arrays and k second chord arrays associated with the first chord array, where k is a positive integer.

[0084] In the embodiment of the present application, the so-called chord refers to a group of pitches with a certain interval relationship, that is, three or more notes are combined vertically according to the overlapping relationship of thirds or non-thirds, which is called a chord. Figure 3 , Figure 3 is a schematic diagram of the structure of a chord array provided in an embodiment of the present application. Figure 3 As shown, the chord array can be composed of n chords, namely, chord 1, chord 2...chord n, wherein each chord is composed of 3 or more notes with an interval relationship, for example, chord 1 is a group of pitches composed of 3 notes, chord 2 is a group of pitches composed of 4 notes, and chord 3 is a group of pitches composed of 5 notes.

[0085] In a possible implementation, the computer device may call a neural network model to identify a first chord array and generate a pitch array and a second chord array associated with the first chord array. Figure 4 , Figure 4 Schematic diagram of a process of calling a neural network model provided by an embodiment of the present application. Figure 4 As shown, the first chord array is input into the neural network model, and after recognition processing by the neural network model, the pitch array and the second chord array associated with the first chord array are output.

[0086] It should be understood that each time the neural network model runs, a pitch array and a second chord array will be generated. In the present application, the value k can be preset, that is, the value k depends on the number of times the neural network model runs. Specifically, the number of times the neural network model runs is equal to the number k of the generated pitch array and the second chord array. In addition, the second chord array refers to: a chord array with the same number of chords as the first chord array, but a chord array with a different arrangement order from the first chord array. Specifically, the second chord array refers to a chord array obtained by rearranging the chord order based on the neural network model. For example, the first chord array includes: chord 1, chord 2, chord 3... chord n; the second chord array includes: chord n, chord 2... chord 3, chord 1. It can be understood that the second chord array stores the chord information after the first chord array is rearranged. Subsequently, generating the target audio based on the rearranged second chord array can enrich the audio structure and enhance the diversity of audio production.

[0087] Specifically, music theory knowledge shows that there is a certain harmonious or conflicting mapping relationship between the chord array and the pitch information. For example, there is a relatively conflicting mapping relationship between the main chord (1 (do) -3 (mi) -5 (so)) and 4 (Fa), 7 (Si) in C major, but the mapping relationship between the main chord (1 (do) -3 (mi) -5 (so)) and 2 (Re), 6 (La) in C major is relatively harmonious. Therefore, in combination with the common chord array and pitch information in songs (audio), the neural network model is used in the embodiment of the present application to find the mapping relationship between chords and pitches. Among them, the neural network model mentioned above can be a deep neural network (Deep Neural Networks, DNN) model, and the deep neural network model can include but is not limited to: LSTM (Long Short Term Memory Network) model, ResNet (Residual Network) model, CNN (Convolutional Neural Networks) model, RNN (Recurrent Neural Networks), etc. This application does not specifically limit the model structure of the deep neural network model.

[0088] For example, in the embodiment of the present application, the LSTM model of the RNN mechanism can be selected, which has the advantage that the past music information can be used to influence the current output music information, thereby forming the pitch information related to the past and the present. Specifically, the embodiment of the present application can be explained by taking the improved RNN model as an example. Figure 3 As shown, the first chord array mentioned in the embodiment of the present application can be an array composed of n chords (chord 1, chord 2...chord n). Then, by calling the improve RNN model to identify the first chord array, two arrays, array 1 and array 2, can be output. Among them, array 1 can be a pitch array, array 2 can be a second chord array, and both arrays contain n bytes of chords. In addition, array 1 stores the pitch information generated by the improve RNN model mapping, and array 2 stores the rearranged chord information.

[0089] It should be noted that in the process of inputting the first chord array into the improve RNN model, some music information description will be involved. For example, usually a chord based on a specific pitch is represented by a capital letter. The major chord based on a certain tone is directly represented by the letter. For example, C is used to represent the major chord (1-3-5), and the minor chord based on a certain tone needs to add a lowercase m next to the letter to distinguish it from the major chord. For example, Am is used to represent a sixth-degree minor triad (6-1-3). It can be understood that all information standards related to pitch in the embodiment of the present application are consistent with the general MIDI standard, and the details of the MIDI standard can be found in Table 1 below:

[0090] Table 1. MIDI standard music format

[0091] C C# D D# E F F# G G# A A# B -1 0 1 2 3 4 5 6 7 8 9 10 11 0 12 13 14 15 16 17 18 19 20 21 22 23 1 24 25 26 27 28 29 30 31 32 33 34 35 2 36 37 38 39 40 41 42 43 44 45 46 47 3 48 49 50 51 52 53 54 55 56 57 58 59 4 60 61 62 63 64 65 66 67 68 69 70 71 5 72 73 74 75 76 77 78 79 80 81 82 83 6 84 85 86 87 88 89 90 91 92 93 94 95 7 96 97 98 99 100 101 102 103 104 105 106 107 8 108 109 110 111 112 113 114 115 116 117 118 119 9 120 121 123 124 125 126 127

[0092] In a possible implementation, the computer device may input the first chord array into the neural network model, and run the neural network model k times to obtain k pitch arrays and k second chord arrays associated with the first chord array. Wherein, each time the neural network model is run once, a pitch array and a second chord array associated with the first chord array are generated. Moreover, the first chord array and any pitch array and any second chord array may have the same number of bytes, that is, assuming that the first chord array includes n chords (i.e., the number of bytes is n), then the embodiment of the present application may run the improve RNN model k times to obtain the pitch information of n*k bars (i.e., k pitch arrays, each pitch array is composed of the pitch of n bars), and at the same time, n*k bars of the second chord array may be obtained (i.e., k second chord arrays, each second chord array is also composed of the chord of n bars). Based on this approach, running the neural network model multiple times to obtain multiple pitch arrays and multiple second chord arrays can increase the musical structure information of the note string, thereby improving the accuracy of the audio production process.

[0093] S202: Arrange and process k pitch arrays to generate first audio track data.

[0094] In a specific implementation, arranging the k pitch arrays specifically refers to: arranging the k pitch arrays according to a preset musical form, wherein the so-called structural arrangement process may include: arranging the arrangement order of the k pitch arrays, copying some of the k pitch arrays, and the like. Then, the first audio track data is generated according to the multiple (greater than or equal to k) pitch arrays after arrangement. The so-called first audio track data refers to the track data stored in the first audio track.

[0095] In a possible implementation, the computer device arranges k pitch arrays to generate first track data, which may include: first, arranging the k pitch arrays according to M-segment musical form structures to obtain main melody audio data, where M is a positive integer; then, setting audio attribute information for the main melody audio data, and determining the main melody audio data with the audio attribute information as the first track data. The audio attribute information includes any one or more of instrument information, number of beats, and rhythm texture.

[0096] Specifically, taking k=3, M=3 as an example, the improve RNN model generates three pitch arrays of length n, which are respectively recorded as array 1 (array1), array 2 (array2), and array 3 (array3). Figure 5 , Figure 5Schematic diagram of a process of arranging a pitch array provided by an embodiment of the present application. Figure 5 As shown, according to the rules of composition paragraph structure, array1, array2, array3 can be arranged according to the rule of M-paragraph form structure: array1-array2-array1-array3-array1-array2-array1. Among them, the so-called three-paragraph form can also be called ABA form, which is the most common form of music in musical works. It consists of two equally important structural music to form a body with three "paragraphs", in which the first and third paragraphs can be exactly the same or almost all the same content. Among them, in Figure 5 In the arrangement shown, array1 is defined as the theme segment, which is repeated three times, and array1, array2 and array3 form two three-segment musical forms, namely array1-array2-array1 and array1-array3-array1. In this way, the result obtained after the arrangement processing can be used as the main melody of the audio (main melody audio data).

[0097] Furthermore, the instrument information of the main melody audio data can be set according to the user's preferences and personalized needs, such as piano, violin, guitar and other instruments, the playback speed (BPM, number of beats) of the main melody audio data can be set, and the rhythm texture and other information can be set to obtain the first audio track data of the first audio track. In this way, after the main melody audio data is arranged and processed, the corresponding audio attribute information can be adaptively added according to the user's personalized needs, so that the audio creation process can be made more flexible, rich and diverse, and the interest of the creation is enhanced.

[0098] S203. Generate second audio track data matching the audio duration of the first audio track data using k second chord arrays according to the audio duration of the first audio track data.

[0099] Specifically, the second audio track data can be generated by k second chord arrays, so that the audio duration of the second audio track data matches the audio duration of the first audio track data. More specifically, the audio duration of the second audio track data matches the audio duration of the first audio track data, which can mean that the audio duration of the second audio track data is the same as the audio duration of the first audio track data. The so-called second audio track data refers to the track data stored in the second audio track.

[0100] In a possible implementation, a computer device generates second track data according to k second chord arrays and first track data, including: first, obtaining the audio duration of the first track data; then, according to the audio duration of the first track data, performing a cyclic copy process on the k second chord arrays to obtain the second track data. The audio duration of the second track data matches the audio duration of the first track data. The so-called matching may mean that the audio duration of the second track data is the same as the audio duration of the first track data.

[0101] Specifically, the computer device can perform a cyclic copy process on the k second chord arrays generated by the aforementioned call to the neural network model to obtain second audio track data, and the audio duration of the second audio track data is consistent with the audio duration of the first audio track data. In other words, the purpose of the cyclic copy is to keep the audio duration of the respective track data in the first audio track and the second audio track consistent. Specifically, the cyclic copy process of the k second chord arrays includes: copying part or designated chords included in the k second chord arrays, for example, the k second chord arrays include: chords 1, 2, 3, then after cyclic copying, the following can be obtained: chords 1, 1, 2, 2, 3, 3, or chords 1, 2, 3, 1, 2, 3.

[0102] S204: Generate target audio according to the first audio track data and the second audio track data.

[0103] In a possible implementation, the computing device generates the target audio according to the first audio track data and the second audio track data, which may include: ① directly merging the first audio track data and the second audio track data to obtain the target audio. ② firstly performing corresponding processing on the first audio track data and the second audio track data; then merging the processed first audio track data and the second audio track data to obtain the target audio. ③ firstly merging the first audio track data and the second audio track data to obtain the total audio data; then performing corresponding processing on the total audio data to obtain the target audio.

[0104] In a possible implementation, a computing device generates a target audio based on the first audio track data and the second audio track data, which may include: first, performing audio distortion processing on the first audio track data and the second audio track data respectively to obtain the first audio track post-processing data and the second audio track post-processing data; then, performing reverberation rendering processing on the first audio track data and the second audio track data to obtain the first audio track reference data and the second audio track reference data; merging the first audio track reference data and the second audio track reference data to obtain the target audio. Among them, the audio distortion processing is to enhance the distortion of the audio; the reverberation rendering processing is to better meet the reverberation effect of the audio. In this way, the audio distortion processing, reverberation rendering processing and other operations are performed on the track data of each audio track, which can increase the distortion of the audio, thereby producing a target audio with a higher standard.

[0105] In a possible implementation, the computer device combines the first audio track data and the second audio track data to obtain total audio data, and performs clipping processing on the total audio data to obtain target audio. In this way, clipping processing on the total data can simulate the sound quality effect of low-fidelity audio distortion, further increase the distortion of the audio, and improve the accuracy of the low-fidelity audio.

[0106] It should be understood that in the embodiment of the present application, custom addition of audio tracks is supported, for example, a third audio track, a fourth audio track, a fifth audio track, and a sixth audio track, etc. can be added, and the present application does not specifically limit the number of audio track channels. Then, the track data of the corresponding audio track can be used as the third audio track data, the fourth audio track data, the fifth audio track data, and the sixth audio track data, respectively. Subsequently, the target audio is generated according to the generated audio track data.

[0107] In an embodiment of the present application, first, a first chord array can be obtained, and k pitch arrays and k second chord arrays associated with the first chord array can be generated, where k is a positive integer; then, the k pitch arrays can be arranged and processed to generate first track data; and, based on the audio duration of the first track data, second track data that matches the audio duration of the first track data can be generated by the k second chord arrays; finally, target audio can be generated based on the first track data and the second track data. It can be seen that in an embodiment of the present application, multiple pitch arrays associated with the first chord information can be generated, and in the process of information arrangement and processing of the pitch array and the second chord array, customized arrangement and processing can be performed in combination with the personalized needs of the user, thereby meeting more flexible audio production needs. In addition, the entire audio generation process is automated, thereby improving the efficiency of the audio production process.

[0108] See also Figure 6 , Figure 6is a flow chart of another audio generation method provided by the embodiment of the application. The audio generation method can be Figure 1 The terminal device or server in the audio generation system mentioned above is executed. For the sake of convenience, the present application embodiment takes the execution of a computer device as an example. Among them, the audio generation method may include the following steps S601-S608:

[0109] S601. Obtain a first chord array, and generate k pitch arrays and k second chord arrays associated with the first chord array, where k is a positive integer.

[0110] S602: Arrange and process k pitch arrays to generate first audio track data.

[0111] S603: Generate second audio track data according to the k second chord arrays and the first audio track data.

[0112] It should be noted that the specific steps performed by the computer device in steps S601-S603 in the embodiment of the present application can be referred to above. Figure 2 The relevant processes corresponding to steps S201-S203 in the embodiment are not further described in detail in the embodiment of the present application.

[0113] S604: Generate third audio track data.

[0114] In a possible implementation, the audio track data used to generate the target audio may also include third audio track data; each chord in any second chord array includes pitch information of multiple notes. The steps for the computer device to generate the third audio track data are as follows: first, the number of beats corresponding to each chord is obtained; then, according to the number of beats corresponding to each chord, the duration of each note included in each chord in the second audio track data is adapted; finally, the notes after the duration adaptation are rhythmically arranged according to the pitch information to obtain the third audio track data.

[0115] Specifically, if the number of beats included in each chord is m, where m is a positive integer, then the duration of the notes in each chord can be adapted to 1 / m of the chord duration. For example, chord 1 includes 4 pitch information, and the number of beats corresponding to chord 1 is 4 (i.e., one measure includes 4 beats). Then the duration of each pitch information can be changed to 1 / 4 of the duration of each measure, and the four notes can be arranged from low to high according to the pitch information, thereby forming a new rhythmic third track data. It should be noted that in the process of adapting the duration of the notes, the number of beats of the chord and the number of notes included in the chord may be equal or unequal.

[0116] S605: Generate fourth audio track data.

[0117] In a possible implementation, the audio track data used to generate the target audio may further include fourth audio track data. The computer device generates the fourth audio track data in the following steps: obtaining target drum beat data selected from the drum beat database, and combining the target drum beat data according to preset rhythm type data to obtain fourth audio track data. Specifically, the user can select soft, standard, intense and other rhythm type data from the drum beat database, and select different drum set timbres and freely combine them. Figure 7 , Figure 7 It is a schematic diagram of a rhythm type data provided in an embodiment of the present application, taking three preset rhythm type data as an example, wherein soft rhythm type data, standard rhythm type data, and intense rhythm type data correspond to quarter notes, eighth notes, and sixteenth notes, respectively.

[0118] S606: Generate fifth audio track data.

[0119] In a possible implementation, the audio track data used to generate the target audio may further include fifth audio track data. The step of the computer device generating the fifth audio track data is as follows: obtaining the target background audio selected from the background audio database, and using the target background audio as the fifth audio track data. The background audio may include, but is not limited to, the sound of rain, the sound of tides, and the sound of burning flames. In this way, the user can set a specific background audio to create a unique listening atmosphere.

[0120] S607: Generate sixth audio track data.

[0121] In a possible implementation, the audio track data for generating the target audio may further include sixth audio track data. The computer device generates the sixth audio track data in the following steps: obtaining a target note with the lowest pitch information from each chord included in the second audio track data, and using the obtained multiple target notes as the fourth audio track data.

[0122] In specific implementation, users can choose whether to add bass track data according to their own preferences, and use the added bass track data as the sixth track data. For details, please refer to Figure 8 , Figure 8 Schematic diagram of generating bass track data provided by an embodiment of the present application. Figure 8 As shown, assuming that the second chord array includes n chords, namely Chord 1, Chord 2... Chord n, then the target note with the lowest pitch information of each measure (a chord) can be extracted to form the bass track data, and the extracted n target notes are formed into the bass track data, wherein the duration of each target note in the bass track data remains unchanged.

[0123] It should be noted that the generation order of the first track data, the second track data, the third track data, the fourth track data, the fifth track data and the sixth track data is not limited to Figure 6 As shown, it can also be any other one-to-one serial execution order, or the six audio track data can be generated in parallel, or any one or more audio track data can be generated in parallel and generated serially with the remaining other audio track data, etc. The embodiments of the present application do not limit the generation order.

[0124] In summary, see Fig. 9 , Fig. 9 FIG. 1 is a flow chart of generating target audio provided in an embodiment of the present application. Fig. 9 As shown, the first audio track data refers to the track data of the first audio track (track 1), the second audio track data refers to the track data of the second audio track (track 2), the third audio track data refers to the track data of the third audio track (track 3), the fourth audio track data refers to the track data of the fourth audio track (track 4), the fifth audio track data refers to the track data of the fifth audio track (track 5), and the sixth audio track data refers to the track data of the sixth audio track (track 6).

[0125] S608. Generate target audio according to the first audio track data, the second audio track data, the third audio track data, the fourth audio track data, the fifth audio track data, and the sixth audio track data.

[0126] In one possible implementation, the computer device performs audio distortion processing on each track data to obtain multiple track post-processing data; performs reverberation rendering processing on each track post-processing data to obtain multiple track reference data; and generates target audio based on the multiple track reference data. Among them, effects such as reverberator, clipping, and bandpass filter can be added to simulate LoFi music. For example, reverberation effect rendering can be performed on each track data to enhance the integration of various instruments.

[0127] In a specific implementation, the computer device generates the target audio according to multiple audio track reference data, which may include: merging the multiple audio track reference data to obtain the total audio data to be processed; performing low-fidelity simulation processing on the total audio data to be processed to obtain the target audio. More specifically, performing low-fidelity simulation processing on the total audio data to be processed to obtain the target audio may include: performing clipping processing on the total audio data to be processed to obtain the target audio; or, using the total audio data to be processed as bypass data, and performing bandpass filtering on the bypass data to obtain the target audio. That is, the track data (first audio track data, second audio track data, third audio track data, fourth audio track data, fifth audio track data, and sixth audio track data) of each audio track may be merged to obtain the total audio data, wherein the so-called merging processing may specifically include: superimposing the data of each audio track according to the audio duration to obtain the total audio data. Then, clipping processing may be performed on the total audio data to simulate the sound quality effect of low-fidelity music distortion. In addition, the total audio data can be copied as bypass data, and delayed and bandpass (500-5kHz) filtered to simulate the low-fidelity effect of an old tape machine. In this way, low-fidelity music can be simulated to a certain extent and the distortion of the target audio can be enhanced.

[0128] Furthermore, the track data of each audio track mentioned above are stored in the format of a midi file. In the embodiment of the present application, a synthesizer FluidSynth can be used to realize the conversion between midi files and audio files. Fig. 9 The instrument setting of the main melody track (track 1) shown above selects the corresponding timbre configuration file, the file format of which can be sf2 file or sf3 file. The conversion process between midi file and audio file is as follows:

[0129] default_sound_font=' / user / … / Yamaha C5 Grand-v2.4.sf2'

[0130] fs = FluidSynth()

[0131] fs.midi_to_audio(midi_path,saved_audio_path)

[0132] In the embodiment of the present application, the main melody data is generated through the neural network model, and the track data of each audio track, such as drum data, bass data, background sound, etc., is supported to be customized and added, so that the entire audio production process is automated, saving labor costs, and can efficiently produce ultra-long and uninterrupted Lo-Fi music content. In addition, it can also support users to customize the corresponding musical instruments, rhythm types, BPM and other audio attribute information to meet personalized creation needs, greatly stimulating the user's desire to create.

[0133] See also Fig.10 , Fig.10 Schematic diagram of the structure of an audio generating device provided in an embodiment of the present application. Fig.10 As shown, the audio generating device 1000 can be applied to the computer device (terminal device or server) mentioned in the above embodiments. Specifically, the audio generating device 1000 can be a computer program (including program code) running in a computer device, for example, the audio generating device 1000 is an application software; the audio generating device 1000 can be used to execute the corresponding steps in the video matching processing method provided in the embodiment of the present application. The audio generating device 1000 includes:

[0134] An acquisition unit 1001 is used to acquire a first chord array and generate k pitch arrays and k second chord arrays associated with the first chord array, where k is a positive integer;

[0135] The processing unit 1002 is used to arrange the k pitch arrays to generate first audio track data;

[0136] The processing unit 1002 is further configured to generate second audio track data matching the audio duration of the first audio track data from the k second chord arrays according to the audio duration of the first audio track data;

[0137] The processing unit 1002 is further configured to generate target audio according to the first audio track data and the second audio track data.

[0138] In a possible implementation, the processing unit 1002 generates k pitch arrays and k second chord arrays associated with the first chord array, for performing the following operations:

[0139] Inputting the first chord array into a neural network model, and running the neural network model k times to obtain k pitch arrays and k second chord arrays associated with the first chord array;

[0140] Wherein, each time the neural network model is run, a pitch array and a second chord array are generated.

[0141] In a possible implementation, the processing unit 1002 arranges the k pitch arrays to generate first track data for performing the following operations:

[0142] Arrange the k pitch arrays according to M musical form structures to obtain main melody audio data, where M is a positive integer;

[0143] Audio attribute information is set for the main melody audio data, and the main melody audio data with the audio attribute information is determined as the first track data; wherein the audio attribute information includes: any one or more of instrument information, number of beats, and rhythm texture.

[0144] In a possible implementation, the processing unit 1002 is further configured to perform the following operations:

[0145] According to the number of beats corresponding to each chord in the second chord array, adapting the duration of each note included in each chord in the second audio track data;

[0146] Performing rhythmic arrangement processing on each note after the duration adaptation processing according to the note pitch information of each chord in the second chord array to obtain third track data;

[0147] The processing unit 1002 generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0148] Generate target audio according to the first audio track data, the second audio track data and the third audio track data.

[0149] In a possible implementation, the processing unit 1002 is further configured to perform the following operations:

[0150] Acquire target drum beat data selected from a drum beat database, and combine and process the target drum beat data according to preset rhythm pattern data to obtain fourth track data;

[0151] The processing unit 1002 generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0152] Generate target audio according to the first audio track data, the second audio track data and the fourth audio track data.

[0153] In a possible implementation, the processing unit 1002 is further configured to perform the following operations:

[0154] Acquire target background audio selected from a background audio database, and obtain fifth audio track data according to the target background audio;

[0155] The processing unit 1002 generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0156] Generate target audio according to the first audio track data, the second audio track data and the fifth audio track data.

[0157] In a possible implementation, the processing unit 1002 is further configured to perform the following operations:

[0158] Acquire a target note of the lowest pitch information from each chord included in the second audio track data, and obtain sixth audio track data according to the acquired multiple target notes;

[0159] The processing unit generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0160] Generate target audio according to the first audio track data, the second audio track data and the sixth audio track data.

[0161] In a possible implementation, the processing unit 1002 generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0162] Performing audio distortion processing according to the first audio track data and the second audio track data to obtain audio track post-processing data;

[0163] Performing reverberation rendering processing on the audio track post-processing data to obtain multi-audio track reference data;

[0164] The reference data of each audio track are merged to generate target audio.

[0165] In a possible implementation, the processing unit 1002 generates target audio according to the plurality of audio track reference data, and is used to perform the following operations:

[0166] Merging the multiple audio track reference data to obtain total audio data to be processed;

[0167] The total audio data to be processed is subjected to low-fidelity simulation processing to obtain target audio.

[0168] In an embodiment of the present application, first, a first chord array can be obtained, and k pitch arrays and k second chord arrays associated with the first chord array can be generated, where k is a positive integer; then, the k pitch arrays can be arranged and processed to generate first track data; and, based on the audio duration of the first track data, second track data that matches the audio duration of the first track data can be generated by the k second chord arrays; finally, target audio can be generated based on the first track data and the second track data. It can be seen that in an embodiment of the present application, multiple pitch arrays associated with the first chord information can be generated, and in the process of information arrangement and processing of the pitch array and the second chord array, customized arrangement and processing can be performed in combination with the personalized needs of the user, thereby meeting more flexible audio production needs. In addition, the entire audio generation process is automated, thereby improving the efficiency of the audio production process.

[0169] See also Fig.11 , Fig.11 11 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device 1100 is used to execute the steps performed by the computer device in the aforementioned method embodiment, and the computer device 1100 includes: one or more processors 1110; one or more input devices 1120, one or more output devices 1130 and a memory 1140. The above-mentioned processors 1110, input devices 1120, output devices 1130 and memory 1140 are connected via a bus 1150. Specifically, the memory 1140 is used to store a computer program, and the computer program includes program instructions. The processor 1110 is used to call the program instructions stored in the memory 1140 to perform the following operations:

[0170] Get a first chord array, and generate k pitch arrays and k second chord arrays associated with the first chord array, where k is a positive integer;

[0171] Arrange and process the k pitch arrays to generate first audio track data;

[0172] According to the audio duration of the first audio track data, generating second audio track data matching the audio duration of the first audio track data by the k second chord arrays;

[0173] Generate target audio according to the first audio track data and the second audio track data.

[0174] In a possible implementation, the processor 1110 generates k pitch arrays and k second chord arrays associated with the first chord array, for performing the following operations:

[0175] Inputting the first chord array into a neural network model, and running the neural network model k times to obtain k pitch arrays and k second chord arrays associated with the first chord array;

[0176] Wherein, each time the neural network model is run, a pitch array and a second chord array are generated.

[0177] In a possible implementation, the processor 1110 arranges the k pitch arrays to generate first track data for performing the following operations:

[0178] Arrange the k pitch arrays according to M musical form structures to obtain main melody audio data, where M is a positive integer;

[0179] Audio attribute information is set for the main melody audio data, and the main melody audio data with the audio attribute information is determined as the first track data; wherein the audio attribute information includes: any one or more of instrument information, number of beats, and rhythm texture.

[0180] In a possible implementation, the processor 1110 generates second track data according to the k second chord arrays and the first track data, and is used to perform the following operations:

[0181] Obtaining the audio duration of the first audio track data;

[0182] According to the audio duration of the first audio track data, the k second chord arrays are cyclically copied to obtain second audio track data; wherein the audio duration of the second audio track data matches the audio duration of the first audio track data.

[0183] In a possible implementation, the processor 1110 is further configured to perform the following operations:

[0184] According to the number of beats corresponding to each chord in the second chord array, adapting the duration of each note included in each chord in the second audio track data;

[0185] Performing rhythmic arrangement processing on each note after the duration adaptation processing according to the note pitch information of each chord in the second chord array to obtain third track data;

[0186] The processor 1110 generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0187] Generate target audio according to the first audio track data, the second audio track data and the third audio track data.

[0188] In a possible implementation, the processor 1110 is further configured to perform the following operations:

[0189] Acquire target drum beat data selected from a drum beat database, and combine and process the target drum beat data according to preset rhythm pattern data to obtain fourth track data;

[0190] The processor 1110 generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0191] Generate target audio according to the first audio track data, the second audio track data and the fourth audio track data.

[0192] In a possible implementation, the processor 1110 is further configured to perform the following operations:

[0193] Acquire target background audio selected from a background audio database, and obtain fifth audio track data according to the target background audio;

[0194] The processor 1110 generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0195] Generate target audio according to the first audio track data, the second audio track data and the fifth audio track data.

[0196] In a possible implementation, the processor 1110 is further configured to perform the following operations:

[0197] Acquire a target note of the lowest pitch information from each chord included in the second audio track data, and obtain sixth audio track data according to the acquired multiple target notes;

[0198] The processing unit generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0199] Generate target audio according to the first audio track data, the second audio track data and the sixth audio track data.

[0200] In a possible implementation, the processor 1110 generates target audio according to the first audio track data and the second audio track data, and is used to perform the following operations:

[0201] Performing audio distortion processing according to the first audio track data and the second audio track data to obtain audio track post-processing data;

[0202] Performing reverberation rendering processing on the audio track post-processing data to obtain audio track reference data;

[0203] The reference data of each audio track are merged to generate target audio.

[0204] In a possible implementation, the processor 1110 generates target audio according to the plurality of audio track reference data, and is used to perform the following operations:

[0205] Merging the multiple audio track reference data to obtain total audio data to be processed;

[0206] The total audio data to be processed is subjected to low-fidelity simulation processing to obtain target audio.

[0207] In an embodiment of the present application, first, a first chord array can be obtained, and k pitch arrays and k second chord arrays associated with the first chord array can be generated, where k is a positive integer; then, the k pitch arrays can be arranged and processed to generate first track data; and, based on the audio duration of the first track data, second track data that matches the audio duration of the first track data can be generated by the k second chord arrays; finally, target audio can be generated based on the first track data and the second track data. It can be seen that in an embodiment of the present application, multiple pitch arrays associated with the first chord information can be generated, and in the process of information arrangement and processing of the pitch array and the second chord array, customized arrangement and processing can be performed in combination with the personalized needs of the user, thereby meeting more flexible audio production needs. In addition, the entire audio generation process is automated, thereby improving the efficiency of the audio production process.

[0208] In addition, it should be pointed out here that: the embodiment of the present application also provides a computer storage medium, and a computer program is stored in the computer storage medium, and the computer program includes program instructions. When the processor executes the above program instructions, the method in the corresponding embodiment of the foregoing text can be executed, so it will not be repeated here. For technical details not disclosed in the computer storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application. As an example, the program instructions can be deployed on a computer device, or executed on multiple computer devices located in one place, or, executed on multiple computer devices distributed in multiple locations and interconnected by a communication network.

[0209] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes a computer instruction, and the computer instruction is stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device can execute the method in the above corresponding embodiment, so it will not be repeated here.

[0210] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the above-mentioned program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The above-mentioned storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0211] The above disclosure is only the preferred embodiment of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. An audio generation method, characterized in that: include: Obtain a first chord array, and generate k pitch arrays and k second chord arrays associated with the first chord array, where k is a positive integer; wherein the k pitch arrays and k second chord arrays associated with the first chord array are obtained by inputting the first chord array into a neural network model and running the neural network model k times, and each time the neural network model is run once, one pitch array and one second chord array are generated; Arranging the k pitch arrays to generate first track data; arranging the k pitch arrays includes: arranging the k pitch arrays according to M musical forms, where M is a positive integer; According to the audio duration of the first audio track data, generating second audio track data matching the audio duration of the first audio track data by the k second chord arrays; Generate target audio according to the first audio track data and the second audio track data.

2. The method according to claim 1, characterized in that Arranging the k pitch arrays to generate first track data includes: Arrange the k pitch arrays according to M musical form structures to obtain main melody audio data, where M is a positive integer; Audio attribute information is set for the main melody audio data, and the main melody audio data with the audio attribute information is determined as the first track data; wherein the audio attribute information includes: any one or more of instrumentation information, number of beats, and rhythm texture.

3. The method according to claim 1, characterized in that Also includes: According to the number of beats corresponding to each chord in the second chord array, adapting the duration of each note included in each chord in the second audio track data; Performing rhythmic arrangement processing on each note after the duration adaptation processing according to the note pitch information of each chord in the second chord array to obtain third track data; Generating target audio according to the first audio track data and the second audio track data includes: Generate target audio according to the first audio track data, the second audio track data and the third audio track data.

4. The method according to claim 1, characterized in that Also includes: Acquire target drum beat data selected from a drum beat database, and combine and process the target drum beat data according to preset rhythm pattern data to obtain fourth track data; Generating target audio according to the first audio track data and the second audio track data includes: Generate target audio according to the first audio track data, the second audio track data and the fourth audio track data.

5. The method according to claim 1, characterized in that Also includes: Acquire target background audio selected from a background audio database, and obtain fifth audio track data according to the target background audio; Generating target audio according to the first audio track data and the second audio track data includes: Generate target audio according to the first audio track data, the second audio track data and the fifth audio track data.

6. The method according to claim 1, characterized in that Also includes: Acquire a target note of the lowest pitch information from each chord included in the second audio track data, and obtain sixth audio track data according to the acquired multiple target notes; Generating target audio according to the first audio track data and the second audio track data includes: Generate target audio according to the first audio track data, the second audio track data and the sixth audio track data.

7. The method according to claim 1, characterized in that Generating target audio according to the first audio track data and the second audio track data, including: Performing audio distortion processing according to the first audio track data and the second audio track data to obtain audio track post-processing data; Performing reverberation rendering processing on the audio track post-processing data to obtain audio track reference data; The reference data of each audio track are merged to generate target audio.

8. A computer device, characterized in that: include: storage devices and processors; a memory storing one or more computer programs; A processor, used to load the one or more computer programs to implement the audio generation method as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the audio generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • LSTM multi-track music generation method based on alignment harmony relationship

    CN112017621A