Audio Data Processing Method, Apparatus, Device, Storage Medium and Product

By acquiring and analyzing the characteristic information of multi-track audio data, and adjusting the target audio data to be generated using the audio generation model, the problem of lack of automation and intelligence in the music creation process is solved, and efficient music generation is achieved.

CN115331648BActive Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210935243.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-03
Publication Date
2025-07-08
Estimated Expiration
2042-08-03

AI Technical Summary

Technical Problem

In the prior art, the music creation and recording process lacks automation and intelligence, resulting in inefficient music generation.

Method used

By obtaining sample multi-track audio data and its annotated audio feature information, using the initial audio generation model to predict and adjust the audio feature information, the target multi-track audio data is generated, and the automatic and intelligent generation of audio data is achieved.

Benefits of technology

The automation and intelligence of the music generation process are realized, and the efficiency and accuracy of music creation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115331648B_ABST
    Figure CN115331648B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an audio data processing method, apparatus, device, storage medium and product, including: obtaining sample multi-track audio data and annotation audio feature information respectively corresponding to N audio segments; determining predicted audio feature information of audio segment N1 according to the annotation audio feature information of audio segment N1; using an initial audio generation model to predict the predicted audio feature information of audio segment N according to the annotation audio feature information of the audio segments in the audio segment set. i If the predicted audio feature information respectively corresponding to the N audio segments is obtained, the initial audio generation model is adjusted according to the annotation audio feature information respectively corresponding to the N audio segments and the predicted audio feature information respectively corresponding to the N audio segments, and the adjusted initial audio generation model is used to generate target multi-track audio data, so as to realize the automatic and intelligent generation of multi-track audio data based on artificial intelligence technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and particularly relates to an audio data processing method, apparatus, device, storage medium and product. Background Art

[0002] Audio, such as music, is used for people's daily leisure and entertainment. For example, for music, the musical scores are manually created by composers themselves. Some singers sing based on the musical scores and lyrics, and record during the singing process to generate the song. However, this method is not automated and intelligent enough. Summary of the Invention

[0003] Embodiments of the present application provide an audio data processing method, apparatus, device and storage medium, which can realize the automated and intelligent generation of audio data.

[0004] In a first aspect, embodiments of the present application provide an audio data processing method, including:

[0005] Obtain sample multi-track audio data and annotation audio feature information corresponding to N audio segments respectively; the sample multi-track audio data includes the N audio segments generated by at least two playing musical instruments; N is an integer greater than or equal to 1;

[0006] Determine the predicted audio feature information of the audio segment N1 according to the annotation audio feature information of the audio segment N1; the audio segment N1 is the audio segment with the earliest playing time among the N audio segments;

[0007] Use an initial audio generation model to predict the predicted audio feature information of the audio segment N i according to the annotation audio feature information of the audio segments in the audio segment set; the audio segment N i belongs to the audio segments other than the audio segment N1 among the N audio segments, and i is a positive integer greater than 1 and less than or equal to N; the audio segment set includes all the audio segments among the N audio segments whose playing time is before the audio segment N i ;

[0008] If the predicted audio feature information corresponding to the N audio segments is obtained, then adjust the initial audio generation model according to the annotation audio feature information corresponding to the N audio segments and the predicted audio feature information corresponding to the N audio segments, and determine the adjusted initial audio generation model as the target audio generation model for generating target multi-track audio data.

[0009] In a second aspect, embodiments of the present application provide an audio data processing apparatus, including:

[0010] An acquisition module, configured to acquire sample multi-track audio data and labeled audio feature information corresponding to N audio segments respectively; the sample multi-track audio data includes the N audio segments generated by at least two performing musical instruments; N is an integer greater than or equal to 1;

[0011] A determination module, configured to determine predicted audio feature information of the audio segment N1 according to the labeled audio feature information of the audio segment N1; the audio segment N1 is the audio segment with the earliest playback time among the N audio segments;

[0012] A prediction module, configured to use an initial audio generation model to predict the predicted audio feature information of the audio segment N i according to the labeled audio feature information of the audio segments in the audio segment set; the audio segment N i belongs to the audio segments other than the audio segment N1 among the N audio segments, and i is a positive integer greater than 1 and less than or equal to N; the audio segment set includes all the audio segments among the N audio segments whose playback time is before the audio segment N i ;

[0013] An adjustment module, configured to, if the predicted audio feature information corresponding to the N audio segments is obtained, adjust the initial audio generation model according to the labeled audio feature information corresponding to the N audio segments and the predicted audio feature information corresponding to the N audio segments, and determine the adjusted initial audio generation model as the target audio generation model for generating target multi-track audio data.

[0014] In a third aspect, an embodiment of the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in the first aspect are implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and characterized in that when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0017] In summary, the computer device can obtain the sample multi-track audio data and the labeled audio feature information corresponding to each of the N audio segments; determine the predicted audio feature information of the audio segment N1 according to the labeled audio feature information of the audio segment N1; use the initial audio generation model to predict the audio segment N i according to the labeled audio feature information of the audio segments in the audio segment set. If the predicted audio feature information corresponding to each of the N audio segments is obtained, the initial audio generation model is adjusted according to the labeled audio feature information corresponding to each of the N audio segments and the predicted audio feature information corresponding to each of the N audio segments, and the adjusted initial audio generation model is determined as the target audio generation model for generating the target multi-track audio data, thereby realizing the automated and intelligent generation of audio data. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a schematic structural diagram of a multimedia data processing system provided by an embodiment of the present application;

[0020] Figure 2 is a schematic flowchart of an audio processing method provided by an embodiment of the present application;

[0021] Figure 3 is an example of sample audio feature information provided by an embodiment of the present application;

[0022] Figure 4 is a schematic diagram of an audio processing process provided by an embodiment of the present application;

[0023] Figure 5 is a schematic structural diagram of an audio processing device provided by an embodiment of the present application;

[0024] Figure 6 is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The following will describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application.

[0026] Music is divided into single-track music and multi-track music.

[0027] Multi-track music refers to music played by multiple instruments. For example, Music 1 is played by a violin and a piano together, so Music 1 is multi-track music. Music 2 is played by a guitar, a bass, and a drum together, so Music 2 is also multi-track music.

[0028] The audio files of multi-track music are divided into sound files and Musical Instrument Digital Interface (MIDI) files. The sound files and MIDI files of multi-track music can be converted into each other. That is, the sound file of multi-track music can be transcribed into the MIDI file of multi-track music, and the MIDI file of multi-track music can also be reverse-transcribed into the sound file of multi-track music. Among them, the formats of sound files include but are not limited to Wave, AIF, Audio, MPEG, etc. The format of MIDI files is MIDI.

[0029] A sound file is an audio file recorded by a recording device, which records the binary sampling data of multi-track music. The binary sampling data is obtained by converting the sound of the recorded multi-track music from an analog signal to a digital signal.

[0030] A MIDI file is a music file synthesized by a computer, which records information such as the digital control signals of each note of each instrument in multi-track music. The MIDI file describes music in a language that a computer can understand. The MIDI file describes music in the form of bytes.

[0031] The MIDI file records the music data of each measure of music. The music data includes information such as the instrument information of at least one instrument participating in the performance in the measure, and the note information of each note played by each instrument participating in the performance in the measure. The instrument information is used to identify the instrument. The note information includes note type, pronunciation duration, and pronunciation intensity.

[0032] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0033] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0034] The key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS), and voiceprint recognition technology. Enabling the computer to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.

[0035] Natural language processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph, and other technologies.

[0036] The solution provided in the embodiments of this application involves technologies such as the speech technology and natural speech processing technology of artificial intelligence.

[0037] To facilitate a clearer understanding of this application, first, a multimedia data processing system for implementing the multimedia data processing method of this application will be introduced, as Figure 1 shown, as Figure 1 shown, the multimedia data processing system includes a server 10 and a terminal cluster. The terminal cluster can include one or more terminals, and the number of terminals will not be limited here. As Figure 1 shown, the terminal cluster can specifically include terminal 1, terminal 2,..., terminal n; it can be understood that terminal 1, terminal 2, terminal 3,..., terminal n can all be network-connected to the server 10 so that each terminal can perform data interaction with the server 10 through the network connection.

[0038] The terminal is installed with a multi-audio production platform for users. The audio production platform can refer to web pages, applets, audio production programs, etc. The terminal can automatically generate audio data through a target audio generation model in the audio production platform. The audio data can include single-track audio data or multi-track audio data. The single-track audio data can be generated by one playing instrument, and the multi-track audio data can be generated by at least two playing instruments.

[0039] It can be understood that the server 10 can refer to a device for providing backend services for the audio production platform. For example, the server 10 can train an initial audio generation model to obtain a target audio generation model for generating audio data, and send the target audio generation model to the terminal to provide services for the terminal to generate audio data.

[0040] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of at least two physical servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can specifically refer to in-vehicle terminals, smart phones, tablet computers, laptop computers, desktop computers, smart speakers, screen speakers, smart watches, etc., but is not limited thereto. Each terminal and the server can be directly or indirectly connected through wired or wireless communication methods. At the same time, the number of terminals and servers can be one or at least two, and this application does not make any restrictions here.

[0041] Please refer to Figure 2 , which is a schematic flowchart of a method for processing audio data provided by an embodiment of this application. As Figure 1 shown, this method can be executed by the terminal in Figure 1 , or can be executed by the server in Figure 1 , or can be jointly executed by the terminal and the server in Figure 1 . The devices used to execute this method in this application can be collectively referred to as computer devices. Specifically, this method includes the following steps:

[0042] S201. Obtain the sample multi-track audio data and the labeled audio feature information corresponding to N audio segments respectively.

[0043] Among them, the sample multi-track audio data can be one or more. The sample multi-track audio data can be the audio data of the sample multi-track audio. The sample multi-track audio can be sample multi-track music. For example, the sample multi-track music is multi-track music 1, and the sample multi-track audio data is the audio data of multi-track music 1, or the sample multi-track music is multi-track music 2, and the sample multi-track audio data can be the audio data of multi-track music 2 as the sample multi-track music data.

[0044] Among them, the sample multi-track music is jointly played by at least two playing instruments. At least two means two or more. For example, the sample multi-track music is multi-track music 1, and multi-track music 1 is jointly played by a violin and a piano, and multi-track music 2 is jointly played by a guitar, a drum set, and a bass.

[0045] The sample multi-track audio data can include N audio segments. N is an integer greater than or equal to 1. The labeled audio feature information of the audio segment is the true audio feature information of the music segment. The labeled audio feature information can include information that can reflect the audio features of the audio segment.

[0046] In one embodiment, the way for the computer device to obtain the labeled audio feature information corresponding to each of the N audio segments can be: the computer device can perform measure recognition on the sample multi-track audio data to obtain M audio measures of the sample multi-track audio data, and perform data analysis on the M audio measures to obtain the audio segments corresponding to the M audio measures respectively, and obtain the labeled audio feature information of the audio segments corresponding to the M audio measures respectively. Among them, the N audio segments can include the audio segments corresponding to the M audio measures respectively. In one embodiment, the N audio segments can also only include the audio segments of audio measure 1 - audio measure p respectively. Where p is greater than or equal to 1, and less than M, and p is a positive integer. That is to say, the N audio segments can only include some of the audio segments corresponding to the M audio measures respectively. Or, the N audio segments can also include the audio segments of audio measure p - M respectively.

[0047] In one embodiment, the computer device can identify musical measures for the sample multi-track audio data to obtain M audio measures of the sample multi-track audio data in the following way: the computer device performs beat detection on the sample multi-track audio data to obtain a beat detection result, and determines M audio measures of the sample multi-track audio data based on the beat detection result. The beat detection result may include all the musical beats of the sample multi-track audio data and the occurrence times of all the musical beats respectively, where a musical beat refers to a unit beat. Since the division of audio measures is related to the occurrence positions of strong beats, M audio measures of the sample multi-track audio data can be determined based on the beat detection result. For example, if all the audio beats of the sample multi-track audio data are strong, weak, strong, weak, strong, weak in sequence, then based on all these audio beats, 3 measures of the sample audio data can be determined.

[0048] Next, taking the audio measure M j as an example, the method for performing data analysis on M audio measures to obtain the audio segments corresponding to the M audio measures respectively, and the labeled audio feature information of the audio segments corresponding to the M audio measures respectively will be introduced. Among them, the audio measure M j can be any one of the M audio measures. j is a positive integer less than or equal to M.

[0049] Specifically, the computer device can perform note recognition on the audio measure M j to obtain the audio segment corresponding to the audio measure M j and the basic audio attributes of the audio segment corresponding to the audio measure M j ; the computer device determines the labeled audio feature information of the audio segment corresponding to the audio measure M j based on the basic audio attributes of the audio segment corresponding to the audio measure M j . Among them, the audio measure M j may include one or more audio segments. One note in the audio measure M j corresponds to one audio segment.

[0050] Among them, the labeled audio feature information may include audio beat category, pronunciation speed, note type, chord feature, playing instrument category, and note information. To facilitate distinguishing the audio bars where each audio segment is located, the labeled audio feature information may include bar identifier, audio beat category, pronunciation speed, chord feature, playing instrument category, and note information. Among them, the note information may include note type, pronunciation duration, and pronunciation intensity. The bar identifier is used to identify the bar where the audio segment is located. The bar identifier includes but is not limited to bar name or bar code. The audio beat category represents the beat type used in the bar, which is represented by a time signature. The audio beat category includes but is not limited to 1 / 16 beat, 3 / 16 beat, 5 / 16 beat. Referring to Table 1, the audio beat category may be 1 / 16 note position*(1 - 16). Among them, 1–16 represents 1 to 16. 1 / 16 note position represents a 16th note. The pronunciation speed is used to measure the speed of the rhythm, which is the number of beats per unit time, such as the number of beats per minute. Referring to Table 1, the pronunciation speed may be one of 180 bpm, specifically one of 30–209 bpm. The chord feature refers to the chord type adopted in the bar. The chord type here is composed of the root note and the quality. Referring to Table 1, the chord feature may be one of 132 chords. The combination of 12 root notes and 11 qualities can form 132 chords. The playing instrument category refers to the instrument category to which the instrument playing the audio segment belongs. Referring to Table 1, the playing instrument category may be one of 10 instrument categories. The 10 instrument categories include drums, pianos, percussion instruments, organs, guitars, basses, strings, brass, woodwinds, and flutes. Among them, the pronunciation duration is the note duration, and the note duration refers to the note duration of the note in the audio segment, which is one of 64 note durations. The expression of the note duration is 1 / 32 note duration*(1 - 64). 1 / 32 note duration represents a 32nd note. (1 - 64) is 1 to 64. The pronunciation intensity is the note strength. The note strength of the note in the audio segment may be one of 32 note strengths.

[0051] Table 1

[0052]

[0053] In one embodiment, the basic audio attributes of the target audio segment of the audio bar M j include the note type, pronunciation intensity, pronunciation duration, timbre, and audio beat of the target audio segment. The target audio segment is any audio segment in the audio segment corresponding to the audio bar M j Below, based on the basic audio attributes of the audio segment corresponding to the audio bar M j determine the audio bar M jThe method of marking audio feature information of the corresponding audio clip is introduced.

[0054] In one embodiment, in order to reduce the dimension of the data processed by the model, the initial pronunciation intensity of the target audio segment can be classified. The initial pronunciation intensity is the pronunciation intensity obtained after note recognition. Specifically, the computer device can determine the pronunciation intensity range in which the initial pronunciation intensity is located according to the initial pronunciation intensity of the target audio segment, and determine the target pronunciation intensity corresponding to the pronunciation intensity range as the pronunciation intensity corresponding to the target audio segment. Since there are 32 types of intensity, the intensity range 0-127 is evenly mapped to 1-32, for example, 0-3 is mapped to 1. For example, the initial pronunciation intensity of the target audio segment is 1, and it can be determined that the initial pronunciation intensity 1 belongs to 0-3, and the target pronunciation intensity corresponding to 0-3 is 4, and the pronunciation intensity of the target audio segment can be determined to be 4. For another example, the initial pronunciation intensity of the target audio segment is 5, and it can be determined that the initial pronunciation intensity 5 belongs to 5-7, and the target pronunciation intensity corresponding to 5-7 is 8, and the pronunciation intensity of the target audio segment can be determined to be 8.

[0055] In one embodiment, when the audio feature information includes the music beat category, the computer device generates a signal according to the audio bar M. j The basic audio properties of the corresponding audio clip determine the audio section M j The method of marking the audio feature information of the corresponding audio segment may be: the computer device performs the audio segment M j The pronunciation intensity of all corresponding audio clips is detected and the audio section M is determined. j The corresponding pronunciation intensity distribution characteristics are based on the audio section M j The corresponding pronunciation intensity distribution characteristics are used to determine the audio beat category of the target audio segment. j The corresponding pronunciation intensity distribution feature, the method of determining the audio beat category of the target audio segment can be based on the audio section M j The corresponding pronunciation intensity distribution characteristics determine the M of the audio section j The corresponding music beat is based on the M of the audio measure. j The corresponding music beats determine the M of the audio measure. j In one embodiment, in addition to determining the audio beat category of the target audio segment in the above manner, the computer device may also determine the audio beat category of the target audio segment according to the target measure M. j The beat detection result determines the audio beats corresponding to the target measure, and determines the audio beat category of the target measure according to the audio beats corresponding to the target measure, and determines the audio beat category of the target measure as the audio beat category of the target audio segment.

[0056] In one embodiment, when the labeled audio feature information includes chord features, the computer device determines the labeled audio feature information of the audio segment corresponding to the audio measure M based on the basic audio attributes of the audio segment corresponding to the audio measure M. j The way to determine the labeled audio feature information of the audio segment corresponding to the audio measure M can be as follows: The computer device determines the chord features of the target audio segment based on the note types of all the audio segments corresponding to the audio measure M; the chord features between different audio segments corresponding to the audio measure M are the same. That is to say, the above process can determine the chord features corresponding to the audio measure M and use the chord features corresponding to the audio measure M as the chord features of the target audio segment. j In one embodiment, when the labeled audio feature information includes the playing instrument category, the computer device determines the labeled audio feature information of the audio segment corresponding to the audio measure M based on the basic audio attributes of the audio segment corresponding to the audio measure M. j The way to determine the labeled audio feature information of the audio segment corresponding to the audio measure M can be as follows: Based on the timbre of the target audio segment, determine the playing instrument category corresponding to the target audio segment. Among them, the timbre can be used to distinguish one instrument from other instruments. j In one embodiment, in order to reduce the dimension of the data processed by the model, the playing instruments determined according to the target audio segment can be classified. Specifically, the way for the computer device to determine the playing instrument category corresponding to the target audio segment based on the timbre of the target audio segment can be to determine the initial playing instrument corresponding to the target audio segment based on the timbre of the target audio segment; query the playing instrument category to which the initial playing instrument belongs from the instrument category mapping table, and determine the playing instrument category to which the initial playing instrument belongs as the playing instrument category corresponding to the target audio segment. For example, when the initial playing instrument corresponding to the target audio segment is determined to be a bass drum based on the timbre recognized from the target audio segment, if there is a mapping relationship between the bass drum and the percussion instrument in the instrument category table, then the percussion instrument is determined as the playing instrument category corresponding to the target audio segment. j In one embodiment, when the labeled audio feature information includes the pronunciation speed, the computer device determines the labeled audio feature information of the audio segment corresponding to the audio measure M based on the basic audio attributes of the audio segment corresponding to the audio measure M. j The way to determine the labeled audio feature information of the audio segment corresponding to the audio measure M can be as follows: Based on the pronunciation duration of all the audio segments corresponding to the audio measure M and the...

[0057] In one embodiment, when the labeled audio feature information includes the playing instrument category, the computer device determines the labeled audio feature information of the audio segment corresponding to the audio measure M based on the basic audio attributes of the audio segment corresponding to the audio measure M. j The way to determine the labeled audio feature information of the audio segment corresponding to the audio measure M can be as follows: Based on the timbre of the target audio segment, determine the playing instrument category corresponding to the target audio segment. Among them, the timbre can be used to distinguish one instrument from other instruments. j In one embodiment, in order to reduce the dimension of the data processed by the model, the playing instruments determined according to the target audio segment can be classified. Specifically, the way for the computer device to determine the playing instrument category corresponding to the target audio segment based on the timbre of the target audio segment can be to determine the initial playing instrument corresponding to the target audio segment based on the timbre of the target audio segment; query the playing instrument category to which the initial playing instrument belongs from the instrument category mapping table, and determine the playing instrument category to which the initial playing instrument belongs as the playing instrument category corresponding to the target audio segment. For example, when the initial playing instrument corresponding to the target audio segment is determined to be a bass drum based on the timbre recognized from the target audio segment, if there is a mapping relationship between the bass drum and the percussion instrument in the instrument category table, then the percussion instrument is determined as the playing instrument category corresponding to the target audio segment.

[0058] In one embodiment, when the labeled audio feature information includes the pronunciation speed, the computer device determines the labeled audio feature information of the audio segment corresponding to the audio measure M based on the basic audio attributes of the audio segment corresponding to the audio measure M.

[0059] In one embodiment, when the labeled audio feature information includes the pronunciation speed, the computer device determines the labeled audio feature information of the audio segment corresponding to the audio measure M based on the basic audio attributes of the audio segment corresponding to the audio measure M. j The way to determine the labeled audio feature information of the audio segment corresponding to the audio measure M can be as follows: Based on the pronunciation duration of all the audio segments corresponding to the audio measure M and the... j The way to determine the labeled audio feature information of the audio segment corresponding to the audio measure M can be as follows: Based on the pronunciation duration of all the audio segments corresponding to the audio measure M, and the... j Based on the pronunciation duration of all the audio segments corresponding to the audio measure M, and the pronunciation duration of the audio measure M... jThe audio beats of all corresponding audio segments are used to determine the pronunciation speed of the target audio segment; audio measure M j The pronunciation speeds between different corresponding audio segments are the same. In one embodiment, the computer device can, according to audio measure M j The pronunciation durations of all corresponding audio segments are used to determine the pronunciation duration of audio measure M j of, and according to the pronunciation duration of audio measure M j and the audio beats of all corresponding audio segments of audio measure M j the pronunciation duration of a single audio beat is determined, so that the pronunciation speed calculated based on the pronunciation duration of a single beat is used as the pronunciation speed of the target audio segment. It should be noted that the method of determining the pronunciation speed of the target audio segment according to the pronunciation durations of all corresponding audio segments of audio measure M j and the audio beats of all corresponding audio segments of audio measure M j includes but is not limited to this method. In one embodiment, the way for the computer device to determine the pronunciation durations of all corresponding audio segments of audio measure M j can also be that in one embodiment, the computer device can also determine the time intervals between every two adjacent beats in audio measure M j according to the beat detection result, so that the pronunciation speed calculated based on the time intervals between every two adjacent beats in audio measure M j is used as the pronunciation speed of the target music segment. In one embodiment, the computer device can select any one of the time intervals between every two adjacent beats in audio measure M j for calculating the pronunciation speed, or calculate the average value between the time intervals between every two adjacent beats, and calculate the pronunciation speed according to the calculated average value. It should be noted that the method of calculating the pronunciation speed according to the time intervals between every two adjacent beats in audio measure M j includes but is not limited to the above methods.

[0060] In one embodiment, the computer device may also obtain the initial pronunciation duration of the target audio segment; adjust the initial pronunciation duration of the target audio segment according to the pronunciation speed of the target audio segment to obtain the pronunciation duration of the target audio segment. The initial pronunciation duration of the target audio segment refers to the pronunciation duration of the target audio segment without being adjusted by the pronunciation speed. Due to the error in the rhythm of human playing music, the sum of the pronunciation durations of at least one note corresponding to one audio beat may be different from the sum of the pronunciation durations of at least one note corresponding to another audio beat, or the time values between the notes within one audio beat may be inconsistent. In this solution, after determining the speed, the duration of a unit audio beat can be determined, and then the duration of the notes can be adjusted using the duration of the unit beat. The adjustment method can be to increase or decrease the duration of the notes. In one embodiment, the computer device may also generate bar attribute information for the music bar M j to indicate that the music bar M j is a bar, generate instrument attribute information for the playing instrument type (to indicate that the playing instrument type is an instrument), generate rhythm attribute information for the chord feature, audio beat type, and pronunciation speed (indicating that this set of data is related to rhythm), and generate note attribute information for the audio type, pronunciation intensity, and pronunciation speed (indicating that this set of data is related to notes); In one embodiment, the labeled audio feature information includes, in addition to the audio beat category, chord feature, playing instrument category, note type, pronunciation intensity, pronunciation duration, and pronunciation speed, the above-mentioned attribute information. If there are multiple audio segments among the N audio segments that belong to the same audio bar, the bar identifiers corresponding to the multiple audio segments are the same, the audio beat types corresponding to the multiple audio segments are the same, the chord features corresponding to the multiple audio segments are the same, and the pronunciation speeds corresponding to the multiple audio segments are the same.

[0061] S202. Determine the labeled audio feature information of the audio segment N1 as the predicted audio feature information of the audio segment N1.

[0062] S203. Use the initial audio generation model to predict the predicted audio feature information of the audio segment N i according to the labeled audio feature information of the audio segments in the audio segment set.

[0063] S204. If the predicted audio feature information corresponding to the N audio segments is obtained, then adjust the initial audio generation model according to the labeled audio feature information corresponding to the N audio segments and the predicted audio feature information corresponding to the N audio segments, and determine the adjusted initial audio generation model as the target audio generation model for generating the target multi-track audio data.

[0064] Among them, the audio segment N1 is the audio segment with the earliest playback time among the N audio segments. N i belongs to the N audio segments, and i is a positive integer greater than or equal to 1 and less than N; the audio segment set includes all the audio segments among the N audio segments whose playback time is before the audio segment N i preceding ones.

[0065] In steps S202 - S204, the computer device can input the annotation feature information of N audio segments into the initial audio generation model. The computer device can use the initial audio generation model to predict the predicted audio feature information of audio segment 2 based on the annotated audio feature information of audio segment N1; use the initial audio generation model to predict the predicted audio feature information of audio segment 3 based on the annotated audio feature information of audio segment N1 and the annotated audio feature information of audio segment N2, and so on, until the predicted audio feature information of audio segment N is obtained. The computer device can determine the audio feature prediction error of the initial audio generation model according to the annotated audio feature information corresponding to each of the N audio segments and the predicted audio feature information corresponding to each of the N audio segments. If the audio feature prediction error is not in a converged state, the initial audio generation model is adjusted according to the audio feature prediction error to obtain an adjusted initial audio generation model.

[0066] Among them, the way for the computer device to determine the audio feature prediction error of the initial audio generation model according to the annotated audio feature information corresponding to each of the N audio segments and the predicted audio feature information corresponding to each of the N audio segments can be that the computer device calculates the audio prediction error between the annotated audio feature information corresponding to each of the N audio segments and the predicted audio feature information corresponding to this audio segment respectively to obtain N audio prediction errors, and then determines the sum of the N audio prediction errors as the audio feature prediction error of the initial audio generation model. In the above process, the computer device can calculate an audio prediction error according to the predicted audio feature information corresponding to each audio segment every time a predicted audio feature information corresponding to an audio segment is predicted. Or, it can also calculate the audio prediction error between the annotated audio feature information corresponding to each audio segment and the predicted feature information corresponding to this audio segment after obtaining the predicted audio feature information corresponding to each of the N audio segments.

[0067] In one embodiment, the initial audio generation model may adopt a Transformer model, for example, it may be a linear Transformer model. The Transformer model includes an encoder, and in this application, the encoder of the Transformer model can be trained using the labeled audio feature information corresponding to N audio segments respectively, and the trained encoder structure can be used to predict the audio feature information of the audio segments. The number of encoders included in the Transformer model can be one or more. In one embodiment, the encoder includes a word embedding encoding layer, a position encoding layer, and also includes a multi-head attention layer, a first addition and normalization layer connected to the multi-head attention layer. The first addition and normalization layer is also connected to a fully connected neural network, and the fully connected neural network is also connected to a second addition and normalization layer. The last layer of at least one encoder is connected to a fully connected neural network. The fully connected neural network includes multiple prediction heads, including a prediction module for predicting the measure identifier, a prediction module for predicting the audio beat category, a prediction module for predicting the pronunciation speed, a prediction module for predicting the chord feature, a prediction module for predicting the playing instrument category, and a prediction module for predicting the note information. Among them, the prediction module for predicting the note information is respectively a prediction module for predicting the note type, a prediction module for predicting the pronunciation duration, and a prediction module for predicting the pronunciation intensity. In one embodiment, the multiple prediction heads may further include a prediction module for predicting the aforementioned attribute information on this basis.

[0068] In one embodiment, the computer device may obtain the audio feature information of the reference audio segment, and use the target audio generation model to identify the audio feature information of the reference audio segment, obtaining the audio feature information corresponding to K audio segments; fuse the audio feature information of the reference audio segment and the audio feature information corresponding to the K audio segments to obtain the fused audio feature information, and generate the target multi-track audio data according to the fused audio feature information. Among them, the reference audio segment is the audio segment to be predicted. Specifically, the computer device uses the target audio generation model to predict the audio feature information corresponding to the first audio segment according to the audio feature information corresponding to the reference audio segment, and then the target audio generation model predicts the audio feature information corresponding to the second audio segment according to the audio feature information corresponding to the reference audio segment and the audio feature information corresponding to the first audio segment, and so on until the prediction task end condition is met, obtaining the audio feature information corresponding to the Kth audio segment. After obtaining the audio feature information corresponding to the Kth audio segment, the audio feature information corresponding to the Kth audio segment may be output, and / or the audio feature information corresponding to the reference audio segment and the audio feature information corresponding to the K audio segments may be output, or the target multi-track music data may be generated according to the audio feature information corresponding to the reference audio segment and the audio feature information corresponding to the K audio segments and then output. Specifically, it may be to fuse the audio feature information of the reference audio segment and the audio feature information corresponding to the K audio segments to obtain the fused audio feature information, and generate the target multi-track audio data according to the fused audio feature information. It should be noted that according to different application scenarios, other prediction stop conditions may also be set, which will not be elaborated one by one in this application.

[0069] In one embodiment, if there are multiple audio segments among the N audio segments that belong to the same measure, the measure identifiers corresponding to the multiple audio segments are the same, the audio beat types corresponding to the multiple audio segments are the same, the chord features corresponding to the multiple audio segments are the same, and the pronunciation speeds corresponding to the multiple audio segments are the same. Therefore, in the embodiments of the present application, the labeled audio feature information corresponding to the multiple audio segments in the same measure can be fused into Figure 3 the form shown, where the labeled audio feature information corresponding to the multiple audio segments in different measures also refers to Figure 3 the form shown for fusion. In Figure 3 it, Figure 3 the first column of data shown is measure 1. Measure 1 is the measure identifier of the first measure. Figure 3 The second column of data shown, from bottom to top, is the beat (1 / 16), speed, and chord. The beat (1 / 16) is the audio beat type mentioned in the embodiments of the present application, the speed is the pronunciation speed mentioned in the embodiments of the present application, and the chord is the chord feature mentioned in the embodiments of the present application. Figure 3The data in the third column shown is the musical instrument (1), and the musical instrument (1) is the type of musical instrument mentioned in the embodiments of the present application. Figure 3 The fourth to sixth columns shown represent the note information of the notes sequentially played by the musical instrument (1) in measure 1. Figure 3 The data in the fourth column shown, from bottom to top, are pitch, duration, and intensity. Here, the pitch is the note type corresponding to the note at this position, the duration is the sounding duration of the corresponding note, and the intensity is the sounding intensity of the corresponding note. Figure 3 The data in the fifth column shown, from bottom to top, are pitch, duration, and intensity. Here, the pitch is the note type corresponding to the note at this position, the duration is the sounding duration of the corresponding note at this position, and the intensity is the sounding intensity of the corresponding note at this position. Figure 3 The data in the sixth column shown, from bottom to top, are pitch, duration, and intensity. Here, the pitch is the note type corresponding to the note at this position, the duration is the sounding duration of the corresponding note at this position, and the intensity is the sounding intensity of the corresponding note at this position. Regarding Figure 3 For the data in other columns except the first to seventh columns, reference can be made to the above description for understanding, and details will not be elaborated here. In one embodiment, each column of data shown as Figure 3 can be constructed as a compound word and input into the initial audio generation model to train the initial audio generation model. Accordingly, the target audio generation model can then generate other columns of data according to a certain column of data newly input by the user, such as generating other columns of data according to the measure identifier input by the user, that is, realizing the automatic generation of the audio feature information of the audio segment.

[0070] The computer device can obtain the sample multi-track audio data and the labeled audio feature information corresponding to N audio segments respectively; determine the predicted audio feature information of audio segment N1 according to the labeled audio feature information of audio segment N1; use the initial audio generation model to predict the predicted audio feature information of audio segment N according to the labeled audio feature information of the audio segments in the audio segment set i ; if the predicted audio feature information corresponding to N audio segments is obtained, then adjust the initial audio generation model according to the labeled audio feature information corresponding to N audio segments and the predicted audio feature information corresponding to N audio segments, and determine the adjusted initial audio generation model as the target audio generation model for generating the target multi-track audio data, thereby realizing the automatic generation of audio data.

[0071] See Figure 4 for an elaboration of a process for processing audio data provided by the embodiments of the present application.

[0072] A computer device can obtain the sound data of multi-track music (corresponding to the sample multi-track audio data) and convert the audio data of the multi-track music into MIDI data. That is to say, the computer device can obtain the sound file of the multi-track music (including the audio data) and transcribe the audio file of the multi-track music into a MIDI file (including MIDI data). Among them, the computer device can use the MIDI automated transcription technology to transcribe the sample multi-track audio data into MIDI data.

[0073] The computer device can obtain the partial audio feature information corresponding to N audio segments of the multi-track music from the MIDI file (corresponding to part of the information in the sample audio feature information). The computer device can also determine the remaining partial audio feature information corresponding to N audio segments of the multi-track music according to the sound file. Here, it is assumed that one audio segment is the audio segment where a note is located.

[0074] Since the partial audio feature information of each audio segment is recorded in the MIDI file, the partial audio feature information of each audio segment can be obtained from the MIDI file. It should be noted that the partial audio feature information read from the MIDI file can be subjected to data conversion to obtain the audio feature information finally input to the initial audio generation model. Because the MIDI file records a series of instructions, it can be converted into the information finally input to the initial audio generation model shown in Table 1 to improve the interpretability of the audio feature information prediction process. Among them, the partial audio feature information of each audio segment obtained by the computer device from the MIDI file includes: the type of playing instrument, the type of note, the pronunciation intensity, and the pronunciation duration.

[0075] Since the MIDI file does not contain the relevant identification information of the measure identifier and the indication information related to the audio beat type, this solution can determine the measure identifier and the audio beat type of the audio measure corresponding to each music segment according to the sound file. Specifically, for the method of generating the measure identifier and the audio beat type of the audio measure, reference can be made to the previous description and will not be elaborated here.

[0076] In addition, since MIDI files do not contain chord-related identification information, the present application can generate chord features based on MIDI files in the following manner. Specifically, an audio editing interface for the MIDI file is displayed; the audio editing interface includes a timeline composed of the multiple audios, and a note sequence of at least one instrument associated with the multi-track music, and the note sequence includes note images of the at least one instrument sequentially playing each note according to the timeline; a sliding window is moved in the audio editing interface, and an image within the window is obtained each time the window is moved. According to the image obtained each time within the window, the note type corresponding to the image that appears within the window is determined; according to the note type corresponding to the image that appears within the window, chord features corresponding to multiple audio bars are determined. In addition, since MIDI files do not include tempo information, the tempo information can refer to the acquisition method described above and will not be elaborated here.

[0077] Thus, the audio feature information corresponding to each music segment can be obtained and input into the transformer model (original audio generation model) for training. The training process can refer to the above and will not be elaborated here. After obtaining the trained transformer model, it can be used to generate the audio feature information of the audio segment. By reverse-transcribing the predicted audio feature information of the audio segment, a MIDI file can be obtained.

[0078] Please refer to Figure 5 , which is a schematic structural diagram of an audio data processing device provided by an embodiment of the present application. This audio processing device can be applied to the aforementioned terminal. Specifically, the device includes an acquisition module 501, a determination module 502, a prediction module 503, and an adjustment module 504. Among them:

[0079] The acquisition module is used to acquire sample multi-track audio data and the labeled audio feature information corresponding to N audio segments respectively; the sample multi-track audio data includes the N audio segments generated by at least two performing instruments; N is an integer greater than or equal to 1;

[0080] The determination module is used to determine the predicted audio feature information of the audio segment N1 according to the labeled audio feature information of the audio segment N1; the audio segment N1 is the audio segment with the earliest playing time among the N audio segments;

[0081] The prediction module is used to use the initial audio generation model to predict the predicted audio feature information of the audio segment N i according to the labeled audio feature information of the audio segments in the audio segment set; the audio segment N iAudio segments other than the audio segment N1 among the N audio segments, where i is a positive integer greater than 1 and less than or equal to N; the audio segment set includes all audio segments among the N audio segments whose playback times are before the audio segment N i All audio segments before;

[0082] An adjustment module, configured to, if the predicted audio feature information corresponding to the N audio segments is obtained, adjust the initial audio generation model according to the labeled audio feature information corresponding to the N audio segments and the predicted audio feature information corresponding to the N audio segments, and determine the adjusted initial audio generation model as the target audio generation model for generating target multi-track audio data.

[0083] In one embodiment, the obtaining module obtaining the labeled audio feature information corresponding to the N audio segments includes:

[0084] Performing beat detection on the sample multi-track audio data to obtain M audio bars of the sample multi-track audio data; M is an integer greater than or equal to 1;

[0085] Performing note recognition on the audio bar M j To obtain the audio segment corresponding to the audio bar M j And the basic audio attributes of the audio segment corresponding to the audio bar M j Where j is a positive integer less than or equal to M, and one note in the audio bar M j Corresponds to one audio segment, and the number of audio segments corresponding to the M audio bars is N;

[0086] According to the basic audio attributes of the audio segment corresponding to the audio bar M j Determine the labeled audio feature information of the audio segment corresponding to the audio bar M j j

[0087] In one embodiment, the basic audio attributes of the target audio segment corresponding to the audio bar M j Include the note type, pronunciation intensity, pronunciation duration, timbre, and audio beat of the target audio segment; the target audio segment is any audio segment in the audio segment corresponding to the audio bar M j j

[0088] In one embodiment, the determining module determines the labeled audio feature information of the audio segment corresponding to the audio bar M j According to the basic audio attributes of the audio segment corresponding to the audio bar M j Including:

[0089] Performing on the audio bar Mj Perform a distribution detection on the pronunciation intensities of all corresponding audio segments to determine the audio measure M j Corresponding pronunciation intensity distribution characteristics;

[0090] Based on the audio measure M j Corresponding pronunciation intensity distribution characteristics, determine the audio beat category of the target audio segment;

[0091] Based on the audio measure M j Corresponding note types of all audio segments, determine the chord characteristics of the target audio segment; the audio measure M j The chord characteristics between different corresponding audio segments are the same;

[0092] Based on the timbre of the target audio segment, determine the performance instrument category corresponding to the target audio segment;

[0093] Based on the audio measure M j Corresponding pronunciation durations of all audio segments, and the audio measure M j Corresponding audio beats of all audio segments, determine the pronunciation speed of the target audio segment; the audio measure M j The pronunciation speeds between different corresponding audio segments are the same;

[0094] Determine the audio beat category, chord characteristics, performance instrument category, note type, pronunciation intensity, pronunciation duration, and pronunciation speed corresponding to the target audio segment as the labeled audio feature information of the target audio segment.

[0095] In one embodiment, the determining module determines the performance instrument category corresponding to the target audio segment based on the timbre of the target audio segment, including:

[0096] Based on the timbre of the target audio segment, determine the initial performance instrument corresponding to the target audio segment;

[0097] Query the performance instrument category to which the initial performance instrument belongs from the instrument category mapping table, and determine the performance instrument category to which the initial performance instrument belongs as the performance instrument category corresponding to the target audio segment.

[0098] In one embodiment, the determining module determines the pronunciation speed of the target audio segment based on the pronunciation durations of all audio segments corresponding to the audio measure M j And the audio beats of all audio segments corresponding to the audio measure M j Including:

[0099] Based on the audio measure M jFor all the audio beats of the corresponding audio segments, count the audio beats M j in the total number of audio beats;

[0100] According to the pronunciation duration of all the audio segments corresponding to the audio measure Mj, calculate the total pronunciation duration corresponding to the audio measure Mj;

[0101] Determine the pronunciation speed of the target audio segment based on the ratio between the total pronunciation duration and the total number of audio beats.

[0102] In one embodiment, the acquisition module is further configured to acquire the initial pronunciation duration of the target audio segment.

[0103] In one embodiment, the adjustment module is further configured to adjust the initial pronunciation duration of the target audio segment according to the pronunciation speed of the target audio segment to obtain the pronunciation duration of the target audio segment.

[0104] In one embodiment, the adjustment module is further configured to adjust the initial audio generation model according to the labeled audio feature information and the predicted audio feature information respectively corresponding to the N audio segments, including:

[0105] Determine the audio feature prediction error of the initial audio generation model according to the labeled audio feature information and the predicted audio feature information respectively corresponding to the N audio segments;

[0106] If the audio feature prediction error is not in a converged state, adjust the initial audio generation model according to the audio feature prediction error to obtain the adjusted initial audio generation model.

[0107] In one embodiment, the prediction module is further configured to:

[0108] Acquire the audio feature information of the reference audio segment;

[0109] Use the target audio generation model to identify the audio feature information of the reference audio segment to obtain the audio feature information corresponding to K audio segments;

[0110] Fuse the audio feature information of the reference audio segment and the audio feature information corresponding to the K audio segments to obtain the fused audio feature information;

[0111] Generate target multi-track audio data according to the fused audio feature information.

[0112] It can be seen that the audio processing device can obtain the sample multi-track audio data and the labeled audio feature information corresponding to N audio segments respectively; determine the predicted audio feature information of the audio segment N1 according to the labeled audio feature information of the audio segment N1; use the initial audio generation model to predict the audio segment N according to the labeled audio feature information of the audio segments in the audio segment set i 's predicted audio feature information; if the predicted audio feature information corresponding to N audio segments is obtained respectively, then according to the labeled audio feature information corresponding to N audio segments respectively, and the predicted audio feature information corresponding to N audio segments respectively, adjust the initial audio generation model, and determine the adjusted initial audio generation model as the target audio generation model for generating the target multi-track audio data, so as to realize the automatic generation of audio data.

[0113] Please refer to Figure 6 , which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 6 shown, the above computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above computer device 1000 may further include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory, for example, at least one disk memory. Optionally, the memory 1005 may also be at least one storage device far from the aforementioned processor 1001. As Figure 6 shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0114] In Figure 6 the computer device 1000 shown, the network interface 1004 can provide network communication functions; while the user interface 1003 is mainly used to provide an input interface; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:

[0115] Obtain the sample multi-track audio data and the labeled audio feature information corresponding to N audio segments respectively; the sample multi-track audio data includes the N audio segments generated by at least two performing musical instruments; N is an integer greater than or equal to 1;

[0116] Determine the predicted audio feature information of the audio clip N1 according to the labeled audio feature information of the audio clip N1; the audio clip N1 is the audio clip with the earliest playing time among the N audio clips;

[0117] Use the initial audio generation model to predict the audio clip N according to the labeled audio feature information of the audio clips in the audio clip set i 's predicted audio feature information; the audio clip N i belongs to the audio clips among the N audio clips other than the audio clip N1, and i is a positive integer greater than 1 and less than or equal to N; the audio clip set includes all the audio clips among the N audio clips whose playing time is before the audio clip N i ;

[0118] If the predicted audio feature information corresponding to the N audio clips is obtained, then adjust the initial audio generation model according to the labeled audio feature information corresponding to the N audio clips and the predicted audio feature information corresponding to the N audio clips, and determine the adjusted initial audio generation model as the target audio generation model for generating the target multi-track audio data.

[0119] In one embodiment, the processor 1001 may be used to call the device control application program stored in the memory 1005 to obtain the labeled audio feature information corresponding to the N audio clips, including:

[0120] Perform beat detection on the sample multi-track audio data to obtain M audio bars of the sample multi-track audio data; M is an integer greater than or equal to 1;

[0121] Perform note recognition on the audio bar M j to obtain the audio clip corresponding to the audio bar M j and the basic audio attributes of the audio clip corresponding to the audio bar M j ; j is a positive integer less than or equal to M, and one note in the audio bar M j corresponds to one audio clip, and the number of audio clips corresponding to the M audio bars is N;

[0122] According to the basic audio attributes of the audio clip corresponding to the audio bar M j determine the labeled audio feature information of the audio clip corresponding to the audio bar M j ;

[0123] In one embodiment, the audio bar M jThe basic audio attributes of the corresponding target audio segment include the note type, pronunciation intensity, pronunciation duration, timbre, and audio beat of the target audio segment; the target audio segment is the audio measure M j Any audio segment in the corresponding audio segment;

[0124] The processor 1001 can be used to call the device control application stored in the memory 1005 to implement according to the audio measure M j The basic audio attributes of the corresponding audio segment to determine the audio measure M j The labeled audio feature information of the corresponding audio segment, including:

[0125] For the audio measure M j Perform distribution detection on the pronunciation intensity of all audio segments corresponding to, and determine the audio measure M j The corresponding pronunciation intensity distribution feature;

[0126] According to the pronunciation intensity distribution feature of the audio measure M j Determine the audio beat category of the target audio segment;

[0127] According to the audio measure M j The note types of all audio segments corresponding to, determine the chord feature of the target audio segment; the audio measure M j The chord features between different corresponding audio segments are the same;

[0128] According to the timbre of the target audio segment, determine the playing instrument category corresponding to the target audio segment;

[0129] According to the audio measure M j The pronunciation durations of all audio segments corresponding to, and the audio measure M j The audio beats of all audio segments corresponding to, determine the pronunciation speed of the target audio segment; the audio measure M j The pronunciation speeds between different corresponding audio segments are the same;

[0130] Determine the audio beat category, chord feature, playing instrument category, note type, pronunciation intensity, pronunciation duration, and pronunciation speed corresponding to the target audio segment as the labeled audio feature information of the target audio segment.

[0131] In one embodiment, the processor 1001 can be used to call the device control application stored in the memory 1005 to implement determining the playing instrument category corresponding to the target audio segment according to the timbre of the target audio segment, including:

[0132] Determine the initial performance instrument corresponding to the target audio segment according to the timbre of the target audio segment;

[0133] Query the performance instrument category to which the initial performance instrument belongs from the instrument category mapping table, and determine the performance instrument category to which the initial performance instrument belongs as the performance instrument category corresponding to the target audio segment.

[0134] In one embodiment, the processor 1001 may be used to call the device control application program stored in the memory 1005 to implement according to the pronunciation duration of all audio segments corresponding to the audio measure M j and the audio beats of all audio segments corresponding to the audio measure M j to determine the pronunciation speed of the target audio segment, including:

[0135] According to the audio beats of all audio segments corresponding to the audio measure M j statistically count the total number of audio beats within the audio measure M j ;

[0136] Calculate the total pronunciation duration corresponding to the audio measure Mj according to the pronunciation duration of all audio segments corresponding to the audio measure Mj;

[0137] Determine the pronunciation speed of the target audio segment as the ratio between the total pronunciation duration and the total number of audio beats.

[0138] In one embodiment, the processor 1001 may be used to call the device control application program stored in the memory 1005 to implement:

[0139] Obtain the initial pronunciation duration of the target audio segment;

[0140] Adjust the initial pronunciation duration of the target audio segment according to the pronunciation speed of the target audio segment to obtain the pronunciation duration of the target audio segment.

[0141] In one embodiment, the processor 1001 may be used to call the device control application program stored in the memory 1005 to implement adjusting the initial audio generation model according to the labeled audio feature information and the predicted audio feature information respectively corresponding to the N audio segments, including:

[0142] Determine the audio feature prediction error of the initial audio generation model according to the labeled audio feature information and the predicted audio feature information respectively corresponding to the N audio segments;

[0143] If the audio feature prediction error is not in a converged state, the initial audio generation model is adjusted according to the audio feature prediction error to obtain an adjusted initial audio generation model.

[0144] In one embodiment, the processor 1001 may be configured to call the device control application stored in the memory 1005 to implement:

[0145] Obtain the audio feature information of the reference audio segment;

[0146] Use the target audio generation model to identify the audio feature information of the reference audio segment to obtain the audio feature information corresponding to K audio segments;

[0147] Fuse the audio feature information of the reference audio segment and the audio feature information corresponding to the K audio segments to obtain fused audio feature information;

[0148] Generate target multi-track audio data according to the fused audio feature information.

[0149] It should be understood that the computer device 1000 described in the embodiments of the present application may execute the descriptions of the audio data processing methods in the corresponding embodiments mentioned above, and may also execute the descriptions of the audio data processing devices in the corresponding embodiments mentioned above, which will not be elaborated here. In addition, the beneficial effects of using the same method will not be elaborated either. Figure 2 Figure 5

[0150] Figure 2 Figure 4

[0151]

[0152] In addition, it should be noted here that: the embodiments of the present application also provide a computer-readable storage medium, and the computer program executed by the above-mentioned audio data processing device is stored in the above-mentioned computer-readable storage medium, and the above-mentioned computer program includes program instructions. When the above-mentioned processor executes the above-mentioned program instructions, it can execute the descriptions of the above-mentioned audio data processing methods in the corresponding embodiments mentioned above and the corresponding embodiments mentioned above, and therefore, it will not be elaborated here. In addition, the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the descriptions of the method embodiments of the present application.

[0151] As an example, the above-mentioned program instructions may be deployed to be executed on a computer device, or may be deployed to be executed on at least two computer devices at one location. Or, they may be executed on at least two computer devices distributed at at least two locations and interconnected through a communication network. The at least two computer devices distributed at at least two locations and interconnected through a communication network may form a blockchain network.

[0152] The above computer-readable storage medium may be the audio data processing device provided in any of the foregoing embodiments or the middle storage unit of the above computer device, such as the hard disk or the middle memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the middle storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.

[0153] The terms "first", "second", etc. in the description, claims and drawings of the embodiments of the present application are used to distinguish the content in different media, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but may optionally further include steps or modules not listed, or may optionally further include other step units inherent to these processes, methods, devices, products or equipment.

[0154] The embodiments of the present application also provide a computer program product, including computer programs / instructions, which when executed by a processor implement the descriptions of the above audio data processing method in the corresponding embodiments as described above. Figure 4 and Figure 2 Therefore, details will not be described herein again. In addition, the description of the beneficial effects of using the same method will not be repeated. For the technical details not disclosed in the embodiments of the computer program product involved in the present application, please refer to the description of the method embodiments of the present application.

[0155] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0156] The method and related devices provided by the embodiments of the present application are described with reference to the method flowcharts and / or structural schematic diagrams provided by the embodiments of the present application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable network-connected devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable network-connected devices generate a device for implementing the function specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable network-connected devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the function specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable network-connected devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for implementing the function specified in Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.

[0157] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. An audio data processing method, characterized in that, Including: Obtaining sample multi-track audio data and the labeled audio feature information corresponding to N audio segments respectively; The sample multi-track audio data includes the N audio segments generated by at least two playing instruments; N is an integer greater than or equal to 1; Determining the labeled audio feature information of audio segment N1 as the predicted audio feature information of audio segment N1; the audio segment N1 is the audio segment with the earliest playing time among the N audio segments; Using the initial audio generation model to predict the audio feature information of audio segment N according to the labeled audio feature information of i-1 audio segments in the audio segment set i The predicted audio feature information of; the audio segment N i Belongs to the audio segments other than the audio segment N1 among the N audio segments, and i is a positive integer greater than 1 and less than or equal to N; the audio segment set includes all the audio segments among the N audio segments whose playback time is before the audio segment N i Previously; the i-1 audio segments are all the audio segments in the audio segment set whose playback time is before the playback time of the audio segment N i Previously; If the predicted audio feature information corresponding to the N audio segments is obtained, then according to the labeled audio feature information corresponding to the N audio segments and the predicted audio feature information corresponding to the N audio segments, adjust the initial audio generation model, and determine the adjusted initial audio generation model as the target audio generation model for generating target multi-track audio data.

2. The method according to claim 1, wherein The obtaining the labeled audio feature information corresponding to the N audio segments respectively includes: Performing beat detection on the sample multi-track audio data to obtain M audio bars of the sample multi-track audio data; M is an integer greater than or equal to 1; Perform note recognition on the audio measure M j to obtain the audio segment corresponding to the audio measure M j and the basic audio attributes of the audio segment corresponding to the audio measure M j ; j is a positive integer less than or equal to M, and one note within the audio measure M j corresponds to one audio segment, and the number of audio segments corresponding to the M audio measures is N; Based on the basic audio attributes of the corresponding audio segment of the audio section M j determine the labeled audio feature information of the corresponding audio segment of the audio section M j of the corresponding audio segment.

3. The method according to claim 2, wherein the audio segment M j The basic audio attributes of the corresponding target audio segment include the note type, pronunciation intensity, pronunciation duration, timbre, and audio beat of the target audio segment; the target audio segment is the audio segment M j Any audio segment in the corresponding audio segment; The audio segment M corresponding to the audio subsection M j Determine the audio subsection M based on the basic audio attributes of the corresponding audio segment j The labeled audio feature information of the corresponding audio segment, including: Perform distribution detection on the pronunciation intensities of all audio segments corresponding to the audio section M j to determine the pronunciation intensity distribution characteristics corresponding to the audio section M j ; Based on the pronunciation intensity distribution feature corresponding to the audio segment M j determine the audio beat category of the target audio segment; According to the note types of all audio segments corresponding to the audio measure M j determine the chord feature of the target audio segment; the chord features of different audio segments corresponding to the audio measure M j are the same; Determining the playing instrument category corresponding to the target audio segment according to the timbre of the target audio segment; According to the pronunciation duration of all audio segments corresponding to the audio section M j and the audio beats of all audio segments corresponding to the audio section M j to determine the pronunciation speed of the target audio segment; the pronunciation speeds between different audio segments corresponding to the audio section M j are the same; Determining the audio beat category, chord feature, playing instrument category, note type, pronunciation intensity, pronunciation duration, and pronunciation speed corresponding to the target audio segment as the labeled audio feature information of the target audio segment.

4. The method according to claim 3, wherein The determining the playing instrument category corresponding to the target audio segment according to the timbre of the target audio segment includes: Determining the initial playing instrument corresponding to the target audio segment according to the timbre of the target audio segment; Querying the playing instrument category to which the initial playing instrument belongs from the instrument category mapping table, and determining the playing instrument category to which the initial playing instrument belongs as the playing instrument category corresponding to the target audio segment.

5. The method according to claim 3, wherein Based on the audio segment M j the pronunciation durations of all corresponding audio clips, and the audio beats of all audio clips corresponding to the audio segment M j corresponding to determine the pronunciation speed of the target audio clip, including: According to the audio measure M j Based on the audio beats of all corresponding audio segments, count the total number of audio beats within the audio measure M j ; Calculating the total pronunciation duration corresponding to the audio bar Mj according to the pronunciation durations of all audio segments corresponding to the audio bar Mj; Determining the pronunciation speed of the target audio segment as the ratio between the total pronunciation duration and the total number of audio beats.

6. The method according to claim 5, characterized in that The method further includes: Obtaining the initial pronunciation duration of the target audio segment; Adjusting the initial pronunciation duration of the target audio segment according to the pronunciation speed of the target audio segment to obtain the pronunciation duration of the target audio segment.

7. The method according to claim 1, characterized in that The adjusting the initial audio generation model according to the labeled audio feature information corresponding to the N audio segments and the predicted audio feature information corresponding to the N audio segments includes: Determining the audio feature prediction error of the initial audio generation model according to the labeled audio feature information corresponding to the N audio segments and the predicted audio feature information corresponding to the N audio segments; If the audio feature prediction error is not in a convergent state, then adjusting the initial audio generation model according to the audio feature prediction error to obtain the adjusted initial audio generation model.

8. The method according to claim 1, characterized in that The method further includes: Obtaining the audio feature information of the reference audio segment; Use the target audio generation model to identify the audio feature information of the reference audio segment, and obtain the audio feature information corresponding to K audio segments; Fuse the audio feature information of the reference audio segment and the audio feature information corresponding to the K audio segments to obtain fused audio feature information; Generate target multi-track audio data according to the fused audio feature information.

9. An audio data processing device, characterized in that, It includes: An acquisition module for acquiring sample multi-track audio data and the labeled audio feature information respectively corresponding to N audio segments; The sample multi-track audio data includes the N audio segments generated by at least two playing musical instruments; N is an integer greater than or equal to 1; A determination module for determining the labeled audio feature information of the audio segment N1 as the predicted audio feature information of the audio segment N1; the audio segment N1 is the audio segment with the earliest playback time among the N audio segments; A prediction module, configured to use an initial audio generation model to predict the predicted audio feature information of audio segment N according to the labeled audio feature information of i-1 audio segments in the audio segment set; the audio segment N i is an audio segment other than the audio segment N1 among the N audio segments, and i is a positive integer greater than 1 and less than or equal to N; the audio segment set includes all audio segments among the N audio segments whose playback time is before the audio segment N i ; the i-1 audio segments are all audio segments in the audio segment set whose playback time is before the playback time of the audio segment N i ; i ​ An adjustment module for, if the predicted audio feature information respectively corresponding to the N audio segments is obtained, adjusting the initial audio generation model according to the labeled audio feature information respectively corresponding to the N audio segments and the predicted audio feature information respectively corresponding to the N audio segments, and determining the adjusted initial audio generation model as the target audio generation model for generating target multi-track audio data.

10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Information processing method and device, electronic equipment and storage medium

    CN114005424A

  • Automatic arrangement method

    JP2019200427A