Information processing device, information processing method, and information processing program

The information processing device synthesizes high-quality audio using One Shot samples and a deep generative model to address the scarcity of training data and timbre limitations in MIDI-to-Audio technologies, enhancing music production with realistic and detailed timbre representation.

JP2026061418APending Publication Date: 2026-04-09SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional MIDI-to-Audio technologies require training data of audio and MIDI pairs, which are often scarce or unavailable, leading to difficulties in generating high-quality audio, especially for diverse musical pieces, and lack detailed timbre specification and synchronization, limiting their applicability and effectiveness.

Method used

An information processing device that uses pre-prepared One Shot audio samples and a deep generative model to synthesize high-quality audio by pasting and processing MIDI data, incorporating noise removal and denoising techniques to enhance realism and timbre accuracy.

Benefits of technology

Enables the generation of high-quality audio that closely resembles live performances, allowing detailed timbre specification and synchronization, even in domains with limited training data, thereby improving music production capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026061418000001_ABST
    Figure 2026061418000001_ABST
Patent Text Reader

Abstract

To propose an information processing device, an information processing method, and an information processing program that enable the automatic generation of high-quality audio for diverse musical genres. [Solution] The information processing device according to this disclosure includes an acquisition unit that acquires first audio data to be processed, and a generation unit that generates second audio data of higher quality than the first audio data by inputting the first audio data to a pre-prepared deep generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and an information processing program.

Background Art

[0002] In recent years, technologies for automatically generating audio such as music have been attracting attention. For example, a technology for generating audio from MIDI data that records performance information such as which pitch, with what intensity, and for how long to play, called MIDI (Musical Instruments Digital Interface), has been attracting attention.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Non-Patent Document 4

[0004] However, conventional technology requires training data of MIDI data for audio generation, and if such training data is scarce or unavailable, it is not possible to generate high-quality audio.

[0005] Therefore, this disclosure proposes an information processing device, an information processing method, and an information processing program that enable the automatic generation of high-quality audio for a variety of music genres. [Means for solving the problem]

[0006] To solve the above problems, one form of information processing apparatus according to the present disclosure includes an acquisition unit that acquires first audio data to be processed, and a generation unit that generates second audio data of higher quality than the first audio data by inputting the first audio data to a pre-prepared deep generation model. [Brief explanation of the drawing]

[0007] [Figure 1] This diagram shows an overview of the information processing system according to the embodiment. [Figure 2] This figure shows an example of MIDI data according to the embodiment. [Figure 3] This figure shows an example of paste synthesis according to the embodiment. [Figure 4A] This is a sequence diagram (1) showing an overview of the information processing system according to the embodiment. [Figure 4B] This is a sequence diagram (2) showing an overview of the information processing system according to the embodiment. [Figure 4C] This is a sequence diagram (3) showing an overview of the information processing system according to the embodiment. [Figure 5] This figure shows an example configuration of a user terminal according to the embodiment. [Figure 6] This figure shows an example configuration of an information processing device according to the embodiment. [Figure 7] This figure shows an example of a One Shot sample storage unit according to the embodiment. [Figure 8] This figure shows an example of a model storage unit according to the embodiment. [Figure 9] Figure (1) shows an example of a UI screen according to the embodiment. [Figure 10] Figure (2) shows an example of a UI screen according to the embodiment. [Figure 11] This is a flowchart (1) showing the flow of the generation process (first process) in the control unit. [Figure 12] This is a flowchart (2) showing the flow of the generation process (first process) in the control unit. [Figure 13] This is a flowchart showing the flow of the generation process (second process) in the control unit. [Figure 14] This is a hardware configuration diagram showing an example of a computer that implements the functions of an information processing device. [Modes for carrying out the invention]

[0008] Hereinafter, embodiments of the present disclosure will be described in detail based on the drawings. In each of the following embodiments, the same parts are denoted by the same reference numerals, and redundant descriptions are omitted.

[0009] The present disclosure will be described in accordance with the item order shown below. 1. Embodiment 1-1. Outline of the information processing apparatus according to the embodiment 1-2. Outline of the information processing system 1 according to the embodiment 1-3. Outline of the information processing system 2 according to the embodiment 1-4. Configuration of the information processing apparatus according to the embodiment 2. Other embodiments 3. Effects of the information processing apparatus according to the present disclosure 4. Hardware configuration

[0010] (1. Embodiment) (1-1. Outline of the information processing apparatus according to the embodiment) In recent years, technology for automatically generating MIDI data from audio has been attracting attention. Such technology is also called Automatic Music Transcription (AMT) and is an important technology in music production and the like. As a technology related to AMT, there is known a technology that enables automatic generation of MIDI data by learning annotations for audio as teacher data.

[0011] Thus, while technology for generating MIDI data from audio is attracting attention, technology for generating audio from MIDI data is also attracting attention. Such technology for generating audio from MIDI data is also called MIDI-to-Audio and, like AMT, is an important technology in music production and the like.

[0012] A known conventional technique for MIDI-to-Audio is one that enables automatic audio generation by training the system with MIDI data as training data for audio. Thus, conventional MIDI-to-Audio techniques require training with pairs of audio and MIDI data as training data. For example, actual performance audio is needed for MIDI data, and this serves as the training data.

[0013] Therefore, performance audio is required that matches the timeline of the MIDI data, with the MIDI data and the performance time being synchronized. However, in the case of human performance, even if the performer is looking at the sheet music and playing along, it is easy to imagine that some discrepancies will occur. For example, it is easy to imagine that slight discrepancies will occur if the release of the keys is too early or too late when striking them. Considering that this is training data for machine learning, such discrepancies are undesirable, and there is a problem that machine learning cannot be performed properly unless the performance audio matches the MIDI data and the performance time. Furthermore, it is difficult to mechanically edit (correct or process) such discrepancies.

[0014] In addition to the issue of discrepancies that can prevent machine learning from functioning properly, there are other challenges as well. For example, a creator might have created MIDI data for a piece of music but hasn't made it public. This means that MIDI data for music isn't always available, and pairs of audio and MIDI data aren't easily obtained.

[0015] Thus, in addition to the difficulty of obtaining a large number of audio-MIDI data pairs, it is easy to imagine that obtaining a large number of audio-MIDI data pairs that are synchronized with the performance and without any discrepancies is also difficult.

[0016] Therefore, while conventional MIDI-to-Audio technologies can sometimes obtain audio and MIDI data pairs for a limited number of musical pieces, enabling learning, it is difficult to properly train a system on a diverse range of complex musical pieces. For example, some pianos have features (such as sensors) that automatically obtain MIDI data when played, allowing for some degree of learning for piano music. However, not many instruments are equipped with such features, and because there are countless musical combinations, it is difficult to properly train a system on all kinds of music. Furthermore, it is easy to imagine that it is also difficult to actually listen to the sounds being played and transcribe them into musical notation.

[0017] In addition, conventional MIDI-to-Audio technologies have limitations, such as the fact that timbres can only be specified at the category level. For example, conventional technologies only allow timbre specification at the instrument level, such as piano or violin. In music production, subtle differences in timbre are important. For example, even when we say "piano," there are various types such as grand pianos, upright pianos, and digital pianos (electronic pianos), and these subtle differences in timbre are important in music production. Creators, for example, are particular about subtle timbres and carefully select timbres when creating music. With conventional technologies, it was only possible to specify timbres at broad levels such as categories, making practical use in music production difficult.

[0018] The timbre of the embodiment will be further explained. Timbre is a difference in musical expression, and is not limited to differences in instruments, but also includes differences in expression by performers. For example, it is easy to imagine that even when using the same sheet music and the same instruments, the musical expression will differ depending on the performer. For this reason, timbre can also be understood as the expression added to music. In MIDI-to-Audio, for example, if you specify the voice of a certain artist, the song will be sung in that artist's voice. By being able to specify timbre separately, creators can produce music while considering note information and timbre information separately.

[0019] Furthermore, AI (Artificial Intelligence)-based music generation technology has been developing rapidly in recent years. For example, it is possible to generate a final audio waveform simply by inputting text. For instance, simply inputting "a song played on the piano" can generate an audio waveform of a piano performance. In addition, by inputting a music genre such as "classical" or "jazz," it is possible to generate an audio waveform of a piano performance in accordance with that genre. Moreover, by inputting lyrics, it is possible to generate an audio waveform that sounds as if a person is singing along to those lyrics.

[0020] However, because such technologies do not understand information such as MIDI data or timbre, they could not adequately respond to requests such as changing certain notes. In this way, they could not introduce important parameters for creators that would allow them to control the content of the music.

[0021] Therefore, this disclosure aims to enable the automatic generation of high-quality audio for a wide variety of music, including general music. For example, it aims to provide MIDI-to-Audio technology applicable to a wide variety of music. Furthermore, it aims to provide MIDI-to-Audio technology that allows for detailed timbre specification, not just broad units such as categories.

[0022] Furthermore, this disclosure aims to improve the performance of MIDI-to-Audio even in domains where audio and MIDI data pairs are absent or scarce (e.g., instruments, timbres, artists such as singers, genres, languages, etc.). It also aims to enable more detailed synthetic timbre design, for example, while paying attention to the subtle timbres of each individual note. Finally, it aims to provide a user-friendly music production system that facilitates collaboration between AI and humans.

[0023] The following describes the outline of the information processing device 100 according to the embodiment. The information processing device 100 is an information processing device aimed at enabling the automatic generation of high-quality audio for various types of music, and can be any device as long as it can realize the processing in the embodiment. The information processing device 100 is realized by, for example, a server device or a cloud system, and executes the information processing according to the embodiment. For example, the information processing device 100 is realized by a server device or a cloud system that provides or manages a specific service that enables creators to freely produce music, etc. For example, the information processing device 100 is realized by a server device or a cloud system that provides or manages a specific service that supports creators to freely produce music, etc. (for example, by making various tools for music production etc. available).

[0024] As an example, the information processing device 100 uses audio sample data called "One Shot" in the MIDI standard. Hereafter, such audio sample data called "One Shot" will be referred to as "One Shot samples" as appropriate. A One Shot sample is audio data in which sounds are pre-registered, and is audio data of sounds divided into a certain unit (in this case, pitch), such as the note "C" or the note "D". For example, a One Shot sample is audio data of each individual note, such as the note "C" or the note "D". Note that a One Shot sample is not limited to pitch, and can be audio data of sounds divided into any unit. Also, a One Shot sample may be audio data of actual sounds pre-recorded in real space.

[0025] For example, the information processing device 100 generates an audio waveform by pasting and synthesizing One Shot samples according to the MIDI data, and converts it into high-quality audio via a deep generative model (for example, an extended model). For example, the information processing device 100 performs various data processing on the MIDI data transmitted from an external information processing device (such as the user terminal 10 described later) and converts it into high-quality audio desired by the user.

[0026] In the embodiments described below, converting to high-quality audio means, for example, converting to realistic audio that closely resembles a live performance. For example, audio synthesized by pasting together one-shot samples may not be smooth at the transitions between one-shot samples, resulting in unnatural and mechanical-sounding audio. Therefore, this refers to converting the audio to be smooth, including the transitions between one-shot samples. For example, this refers to converting audio that sounds pasted together and lacks dynamics (such as volume and expression) and timbre into audio that closely resembles a performance by a professional musician. For example, with a piano, the dynamics and timbre can be varied in many ways depending on how the hands and arms are moved, and how the whole body is moved, so this refers to converting to audio that closely resembles such a performance.

[0027] These examples are just a few examples and do not necessarily have to be limiting. In other words, high-quality audio can be any audio that is realistic and closer to a live performance than audio that sounds mechanically pasted on.

[0028] Furthermore, in the embodiments described below, the One Shot sample may be, for example, audio data pre-prepared for a specific service, or audio data input by the user each time audio is generated. When pre-prepared audio data for a specific service is used as the One Shot sample, the user can generate audio by preparing MIDI data. When audio data input by the user is used as the One Shot sample, the user will prepare the One Shot sample in addition to the MIDI data. The embodiments described below can be implemented in any form, provided that the user prepares the MIDI data. A variation will be described later, in which an audio waveform synthesized by a predetermined electronic engineering method (such as a synthesizer) is directly used without the user preparing MIDI data.

[0029] Furthermore, in the embodiments described below, the MIDI data may be data created by, for example, a user inputting a melody they have come up with using a keyboard or similar device. Alternatively, the MIDI data may be data obtained from, for example, a specific service that publishes MIDI data of existing songs, and edited by the user.

[0030] (1-2. Overview of the Information Processing System According to the Embodiment 1) Next, an overview of the information processing system 1 used to explain the technology of this disclosure will be described using Figure 1. Figure 1 is a diagram showing an overview of the information processing system 1 according to an embodiment.

[0031] In Figure 1, the information processing device 100 generates an audio waveform OW1 by pasting and synthesizing the audio data OD1 to OD4, which are One Shot samples, along the MIDI data MD1. Then, the information processing device 100 generates an audio waveform OW2 by performing a noise process P1 on audio waveform OW1 to eliminate unrealistic parts (low-quality parts). Finally, the information processing device 100 generates an audio waveform OW3 by performing a noise removal process P2 on audio waveform OW2. In Figure 1, this audio waveform OW3 is the target audio. That is, it is higher quality audio compared to audio waveform OW1. Note that processes P1 and P2 may be performed as a single process. The details of each process shown in Figure 1 will be explained below.

[0032] The information processing device 100 performs two main processes. One is the process of generating audio waveform OW1, and the other is the process of generating audio waveform OW3. Audio waveform OW2 is an audio waveform obtained during the process of generating audio waveform OW3, so it is included in the second process here. The first process, generating audio waveform OW1, is a process to obtain the input information necessary for the second process, generating audio waveform OW3. First, the process of generating audio waveform OW1 will be explained, but there are several variations to the first process of generating audio waveform OW1. Below, the process of generating audio waveform OW1 using MIDI data and One Shot samples will be explained as an example, and the other variations will be explained at the end. Hereafter, the first process will be referred to as "Process 1" and the second process as "Process 2" as appropriate.

[0033] MIDI data will be explained using Figure 2. Figure 2 is a diagram showing an example of MIDI data according to the embodiment. MIDI data is data similar to digital musical notation, where the horizontal axis represents time and the vertical axis represents pitch. The pitch corresponds to, for example, the keys on a piano.

[0034] In Figure 2, the alphanumeric characters on the vertical axis, such as "F3," "B3," "D4," and "F4," represent information about musical pitch. Each rectangle represents a musical note. For example, the rectangle corresponding to "F3" on the vertical axis indicates that the note has an pitch of "F3." The length of the rectangle also indicates the duration for which the note is played. For example, the length of the rectangle for note N1 is approximately 2.3 seconds to 3.4 seconds on the horizontal axis, indicating that the note should be played for approximately 2.3 seconds to 3.4 seconds.

[0035] Furthermore, the horizontal axis shows notes N1 to N4 in parallel, indicating that notes of multiple pitches are played simultaneously. Specifically, it indicates that the pitches "G3#", "B3", "D4", and "D5", corresponding to notes N1 to N4, are played simultaneously. For example, on a piano, it indicates that the keys corresponding to "G3#", "B3", "D4", and "D5" are played simultaneously. For instance, the notes of "G3#", "B3", "D4", and "D5" begin playing around 2.3 seconds into the piece, and note N4 finishes playing around 3 seconds before that. Note that in Figure 2, only the first four rectangles are labeled; this is for explanatory purposes only, and other rectangles may also be labeled as appropriate.

[0036] Next, we will explain the pasting and synthesis of One Shot samples using Figure 3. Figure 3 is a diagram showing an example of pasting and synthesis according to the embodiment. In Figure 3, we will explain that by specifying a certain timbre category, One Shot samples of various pitches of that timbre can be obtained. That is, we will explain that a group of One Shot samples of a single timbre can be obtained. For example, if the specified timbre category is grand piano, a group of One Shot samples of pitches such as "F3" to "G5#" corresponding to a grand piano can be obtained, and if the specified timbre category is electronic piano, a group of One Shot samples of pitches such as "F3" to "G5#" corresponding to an electronic piano can be obtained.

[0037] MIDI data MD1 indicates that note N11 is played at times t1-t2, note N12 at times t2-t3, note N13 at times t4-t6, and note N14 at times t5-t7. Note that times t1-t6 are the times arranged in timeline order, and of these, times t3-t4 are silent.

[0038] For note N11, based on the pitch of note N11, audio data OD1 is selected from the One Shot sample group, and a portion of it (a portion of audio data OD1) corresponding to time t1 to t2 is pasted. Similarly, for note N12, audio data OD2 is selected, and a portion of it (a portion of audio data OD2) corresponding to time t2 to t3 is pasted.

[0039] Similarly, at note N13, audio data OD3 is selected, and a portion of it (a portion of audio data OD3) corresponding to time t4~t6 is pasted. Similarly, at note N14, audio data OD4 is selected, and a portion of it (a portion of audio data OD4) corresponding to time t5~t7 is pasted. Note that at time t5~t6, a portion of audio data OD3 and a portion of audio data OD4 will overlap when pasted.

[0040] Furthermore, while the cutting always involves the beginning portion of the audio data, the portion to be cut is not necessarily limited. For example, in the case of note N11, the beginning portion of audio data OD1 is cut and pasted, but the middle portion may also be cut and pasted, or the tail portion may be cut and pasted, and the example is not limited to this. Also, for example, if there is an important part in the audio data and that part is specified in advance, that part may be preferentially cut and pasted. Also, for example, if there is an important part in the audio data but that part is not specified in advance, that part may be estimated (for example, the part with the most dynamics from the audio waveform may be estimated as the important part), and the estimated part may be preferentially cut and pasted.

[0041] Furthermore, in MIDI data MD1, time t3~t4 is a silent section, and therefore the pasted and synthesized audio waveform will also be silent. Also, in MIDI data MD1, time t5~t6 is a section where sounds overlap, and therefore the pasted and synthesized audio waveform will also be a section where sounds overlap. Specifically, note N14 will be played while note N13 is still playing.

[0042] In Figure 3, for the sake of explanation, it is shown that specifying a certain timbre category yields One Shot samples of various pitches for that timbre. However, it is not limited to specifying only one timbre category; multiple timbre categories may be specified. For example, audio data OD1 may be piano and audio data OD2 may be guitar, so audio waveform OW1 may start as piano and then change to guitar midway through. Thus, it is not limited to synthesizing audio waveforms using One Shot samples from a single timbre category; audio waveforms may be synthesized using One Shot samples from multiple timbre categories. For example, one could select One Shot samples corresponding to each note from among the various pitches of One Shot samples corresponding to each of the multiple timbre categories and synthesize an audio waveform.

[0043] Furthermore, by specifying multiple timbre categories, audio waveforms from multiple timbre categories may be synthesized, or audio waveforms obtained by specifying one timbre category may be combined to ultimately synthesize audio waveforms from multiple timbre categories. In this way, audio waveforms from multiple timbre categories may be synthesized by combining the audio waveforms themselves obtained by specifying timbre categories. By using One Shot samples, the editability of the synthesis is increased, making it applicable to a wide range of audio generation.

[0044] Furthermore, the information processing device 100 may present the generated audio waveform OW1 to the user, enabling editing of MIDI data MD1 and One Shot samples such as audio data OD1 to OD4. For example, the information processing device 100 may enable editing of One Shot samples using ADSR (Attack Decay Sustain Release) or filters. This allows for editing of dynamics, timbre, etc. For example, the user can listen to the pasted and synthesized audio waveform and make changes to the parts they want to modify. For example, if the user listens and decides that a certain note is unnecessary, they can delete the note information from MIDI data MD1, and if they decide that a certain note is necessary, they can add the note information to MIDI data MD1.

[0045] The above describes how the information processing device 100 generates the audio waveform OW1 by pasting and combining audio data OD1 to OD4 according to MIDI data MD1. The following describes how the information processing device 100 converts the audio waveform OW1 into a high-quality audio waveform.

[0046] The audio waveform OW1 implicitly contains MIDI data MD1 and timbre information. The information processing device 100 converts the audio waveform OW1 into a realistic audio waveform while preserving the MIDI data MD1 and timbre information.

[0047] One example of a method for converting to a realistic audio waveform is the noise-denoising approach. In the noise-denoising approach, noise is added to eliminate unrealistic parts. However, if too much noise is added, important parts of the MIDI data MD1 and timbre information will also be erased, so an appropriate level of noise is added. Therefore, it is necessary to adjust the added noise to an appropriate level. This corresponds to the process P1 shown in Figure 1. The information processing device 100 generates the audio waveform OW2 by adding noise adjusted to an appropriate level to the audio waveform OW1.

[0048] Furthermore, the information processing device 100 generates an audio waveform OW3 using a deep generative model that identifies the noisy portion from the noisy audio waveform OW2 and restores it to noise-free audio. This corresponds to process P2 shown in Figure 1. As will be described in detail later, the information processing device 100 generates the audio waveform OW3 by removing noise to create realistic audio. In this case, the information processing device 100 may accept conditions on the direction in which the noise should be removed and generate the audio waveform OW3 based on those conditions. For example, if the information processing device 100 accepts a condition such as "realistic performance," it may generate the audio waveform OW3 by removing noise to create audio of a realistic performance.

[0049] Here, we will explain the details of processes P1 and P2 shown in Figure 1. The information processing device 100 generates an audio waveform OW3 using a pre-trained deep generative model.

[0050] The information processing device 100 adds a noise level to the audio waveform OW1 at the time intervals "T, T-1, ..., 1, 0". The noise level changes stepwise over the intervals "T, T-1, ..., 1, 0". Then, the noise level at a certain time interval is added to the audio waveform OW1. For example, the information processing device 100 adds a noise level at a time "t" selected from the time intervals "T" where the noise level is highest to the time interval "0" where the noise level is lowest. For example, the information processing device 100 adds a noise level at a time "t" that is approximately midway between time "T" and time "0". Alternatively, for example, the information processing device 100 allows the user to select a time "t" (or select information corresponding to time "t") from the intervals "T" to time "0", and adds the noise level at the selected time "t". The user's selection of the noise level can be made using any method. For example, selection based on UI (User Interface) operations such as a slider can be used.

[0051] When allowing the user to select a noise level, the information processing device 100 may allow the user to select the noise level at any stage. For example, the information processing device 100 may allow the user to select the noise level in advance, such as when MIDI data MD1 is input, before the generation of the audio waveform OW1, or it may allow the user to select the noise level when generating the audio waveform OW1 (e.g., immediately after generation). For example, when generating the audio waveform OW1, the information processing device 100 may allow the user to select the noise level at the same time as presenting the audio waveform OW1 to the user and enabling editing of the MIDI data MD1 or One Shot sample. For example, the information processing device 100 may allow the user to select the noise level on the same UI screen when enabling editing of the MIDI data MD1 or One Shot sample. Alternatively, for example, the information processing device 100 may allow the user to select the noise level on a different UI screen than the one used for presenting the audio waveform OW1 to the user or for editing the MIDI data MD1 or One Shot sample.

[0052] The process described above involves adding noise, which results in some information being lost. However, by adding denoising to compensate for this loss, it's possible to generate a realistic audio waveform OW3. By allowing the user to select the noise level, it becomes possible to generate audio waveforms that take into account the user's intentions and requests, such as how much they want to change or how much change they want to tolerate. By allowing users to freely select such parameters, it's possible to provide a music production system with excellent usability.

[0053] The information processing device 100 generates an audio waveform OW3 using, for example, a deep generative model such as the one disclosed in Non-Patent Document 1. For example, the information processing device 100 generates an audio waveform OW3 using a Text-to-Music deep generative model that outputs audio based on text conditions. For example, when a user inputs information that aligns with their purpose (information such as what kind of realism they desire) as text, the information processing device 100 generates an audio waveform OW3 that takes such a purpose into account. In other words, the information processing device 100 generates an audio waveform OW3 by removing noise to achieve a "high-quality violin performance."

[0054] The information processing device 100 may generate the audio waveform OW3 using other noise denoising approaches (e.g., a denoising autoencoder (DAE)), not limited to such deep generative models. For example, the information processing device 100 may generate the audio waveform OW3 using a method that applies a noise denoising approach while preserving MIDI data MD1 and timbre information (it does not have to be a method that explicitly preserves MIDI data MD1 and timbre information). For example, the information processing device 100 may generate the audio waveform OW3 using a deep generative model that is conditioned on MIDI data MD1 and timbre information. Furthermore, the information processing device 100 may generate the audio waveform OW3 using noise denoising approaches with the same deep generative model but different techniques.

[0055] In addition, in the method disclosed in Non-Patent Document 1, noise is added randomly, so noise is not added to all parts that lack realism. However, the information processing device 100 may perform processing to add noise to as many parts that lack realism as possible. The information processing device 100 may perform processing to add noise while adjusting so that noise is added to as many parts that lack realism as possible.

[0056] Note that in the audio waveform OW2 in Figure 1, a black line has been added that is not present in the audio waveform OW1. For the sake of explanation, this black line represents noise added to the audio waveform OW1. Also, in the audio waveform OW3 in Figure 1, it is represented by a single, clean timbre. For the sake of explanation, this indicates that the audio waveform OW3 is smooth and of high quality. In reality, just like the audio waveforms OW1 and OW2, the pitch of the audio waveform OW3 may change midway through.

[0057] The above describes the second process, which converts the data into a high-quality audio waveform. Below, variations of the first process, the generation of the audio waveform OW1, will be explained. In the above embodiment, the process of generating the audio waveform OW1 using MIDI data and a One Shot sample was described as an example, but other variations will be described. Here, two variations will be given as examples, but the method is not limited to these examples, and the audio waveform OW1 may be generated (or acquired) using any method. In all variations, the audio waveform OW1 becomes the input information for the second process.

[0058] (Information processing variation 1: Processing using MIDI data and timbre categories) In the above embodiment, the information processing device 100 may use timbre categories to generate the audio waveform OW1. The information processing device 100 may generate the audio waveform OW1 using MIDI data and timbre categories. For example, if a user specifies MIDI data and a timbre category, the information processing device 100 may select a One Shot sample based on the specified timbre category. For example, the information processing device 100 may randomly select a One Shot sample, or if the optimal One Shot sample is pre-registered for each timbre category, it may select one of the registered One Shot samples.

[0059] Furthermore, when a user specifies a sound category, the information processing device 100 may present multiple candidates and then select a One Shot sample from among those candidates based on the specified sound category. For example, if the sound category specified by the user is piano, the information processing device 100 may present a list of candidates such as "grand piano," "upright piano," and "digital piano" along with a comment such as, "Generally, there are several categories of pianos, but which category would you prefer?" and allow the user to make a selection. In this case, the information processing device 100 may, for example, narrow down the candidates to some extent according to user information (for example, the user's preferences regarding music, sound, instruments, etc., and their history of specifying sound categories) before presenting the list of candidates, or it may present the list of candidates so that certain candidates are displayed preferentially at the top according to user information.

[0060] (Information processing variation 2: Processing using synthesizers, etc.) In the above embodiment, the information processing device 100 may directly acquire an audio waveform that has been pre-synthesized by an electronic engineering method such as a synthesizer and use it as the audio waveform OW1. For example, the user may specify an audio waveform synthesized by a synthesizer, and the information processing device 100 may acquire the audio waveform specified by the user. In this case, the user may provide the audio waveform synthesized by a synthesizer to the information processing device 100 when specifying the audio waveform. The information processing device 100 may then use the audio waveform provided by the user as the audio waveform OW1.

[0061] Thus, the information processing device 100 may use an audio waveform synthesized by a synthesizer or the like as the audio waveform OW1. Furthermore, the information processing device 100 is not limited to a synthesizer, but may use an audio waveform synthesized by other electronic engineering methods as the audio waveform OW1. In this way, the audio waveform OW1 may be an audio waveform generated by a synthesizer or other electronic engineering methods.

[0062] The variations in the audio waveform OW1 generation process have been described above. Various variations according to the embodiment will be described below.

[0063] (Variations in Information Processing 3: Scope of Application of Information Processing) The information processing according to the above embodiment may be implemented, for example, in a web application, application software, or a plugin for a music production tool. Furthermore, the information processing according to the above embodiment may be implemented, for example, in a server device, a cloud system, or the user terminal 10 described later. Also, the information processing according to the above embodiment may be implemented, for example, in a device, method, program, or system. Thus, the information processing according to the above embodiment may be implemented in any manner. For example, the information processing according to the above embodiment is not limited to the case where the information processing device 100 and the user terminal 10 are separate devices, but may also be implemented where the information processing device 100 and the user terminal 10 are integrated.

[0064] (Information processing variation 4: One-shot sample generation process) In the above embodiment, a One Shot sample may consist of multiple types of audio data. For example, a One Shot sample may include multiple types of audio data, such as main audio (main melody, etc.) and sub-audio (secondary melody, etc.). For example, each of the audio data OD1 to OD4 may contain multiple types of audio data, such as the corresponding main audio and sub-audio.

[0065] Furthermore, a One Shot sample may contain, for example, only the main audio, only the sub-audio, or audio data of a different type than the main or sub-audio. For example, a One Shot sample may contain audio data of a sub-audio type that is even more sub-audio than the sub-audio.

[0066] The information processing device 100 may generate a One Shot sample by combining these multiple types of audio data. For example, the information processing device 100 may generate a One Shot sample by combining multiple types of audio data, such as main audio and sub-audio.

[0067] The information processing device 100 may assign weights to the main audio and the sub-audio, and generate One Shot samples according to these weights. For example, the note "C" on a piano and the note "C" on a guitar are different, even though they are both "C" notes. The information processing device 100 may assign weights to each to generate a new One Shot sample of the "C" note. This allows the information processing device 100 to generate a new One Shot sample that has elements of both piano and guitar. As a result, it becomes possible to synthesize audio waveforms with diverse timbres.

[0068] The information processing device 100 may generate One Shot samples by changing the weighting so that the timbre changes on the timeline. For example, the information processing device 100 may generate One Shot samples by changing the weighting for each instrument so that the timbre changes on the timeline.

[0069] If there are multiple types of main audio and sub-audio, the information processing device 100 may generate a One Shot sample by appropriately combining audio data of these types. For example, the information processing device 100 may generate a One Shot sample by appropriately combining one main audio and two sub-audio.

[0070] The information processing device 100 may generate a One Shot sample by combining audio data of the same instrument but of different types, or by combining audio data of different instruments but of different types. For example, the information processing device 100 may combine two types of piano sounds (for example, a grand piano sound and a jazz piano sound), a piano sound and a guitar sound, or a piano sound and a singing voice.

[0071] The information processing device 100 may generate an audio waveform by pasting and synthesizing the One Shot samples generated in this way according to the MIDI data MD1. For example, the information processing device 100 may generate an audio waveform based on the information contained in the One Shot samples and the information contained in the MIDI data MD1. For example, if the information processing device 100 obtains information from the MIDI data MD1 such as the first note being a "C" on a piano and the duration being 1 second from "0:01 to 0:02" (the time from pressing the key to release being 1 second), it may identify the corresponding One Shot sample from the One Shot sample group and paste and synthesize it.

[0072] (Information processing variation 5: Audio waveform editing processing) In the above embodiment, the information processing device 100 may perform editing (pre-processing for the second processing) on ​​the audio waveform OW1. For example, the information processing device 100 may perform editing such as compression or limiting.

[0073] The information processing device 100 may perform editing such as compressing the dynamics so that the volume of the audio when a single note is played and the audio when multiple notes are played within the audio waveform OW1 are kept relatively constant. The information processing device 100 may also perform editing such as applying threshold-based limiting (setting an upper limit on the amplitude) so that the volume is kept relatively constant.

[0074] (Information Processing Variation 6: MIDI Data Variations) In the above embodiment, the MIDI data to be processed may be, for example, the data of an entire song (a collection of songs), a part or section of a song, or a part of an instrument within a song, and is not particularly limited to these.

[0075] Furthermore, the MIDI data to be processed may include information about the release time. Release time refers to the duration of the reverberation after releasing the keys, for example, in the case of a piano. For instance, the time it takes for the sound to decay after the offset ends varies depending on the instrument.

[0076] The information processing device 100 may perform paste synthesis while taking release time into consideration. For example, the information processing device 100 may perform paste synthesis by determining the audio release time and applying it to the One Shot sample. For example, the information processing device 100 may determine a random release time within a certain range based on MIDI data and perform paste synthesis so that the One Shot sample fades out according to that release time.

[0077] (1-3. Overview of the Information Processing System According to the Embodiment 2) Next, the relationship between the information processing of the information processing device 100 and the user terminal 10 will be explained using Figures 4A to 4C. Figures 4A to 4C are sequence diagrams (1) to (3) showing an overview of the information processing system 1 according to the embodiment. Figures 4A to 4C differ in that the acquired data is MIDI data / One Shot sample (Figure 4A), MIDI data / timbre category (Figure 4B), and synthesized audio waveform (Figure 4C), but are similar in other respects, so they will be explained together for the sake of explanation.

[0078] Before explaining Figures 4A-4C, let's first describe the user terminal 10. The user terminal 10 is an information processing device used by a user who desires audio data corresponding to predetermined MIDI data. This user is, for example, a user who creates music while paying attention to the subtle nuances of each individual note, and who desires high-quality audio data corresponding to predetermined MIDI data.

[0079] The user provides predetermined MIDI data and, through information processing according to the embodiment, obtains audio data corresponding to that MIDI data. In this process, the user specifies a timbre category and obtains audio data corresponding to the MIDI data that was played using that timbre category.

[0080] The user terminal 10 can be any device as long as it can perform the processing described in the embodiment. The user terminal 10 may also be a smartphone, tablet, notebook PC, desktop PC, mobile phone, PDA, or other device. Figures 4A to 4C show the case where the user terminal 10 is a smartphone.

[0081] The user terminal 10 is, for example, a smart device such as a smartphone or tablet, and is a portable terminal device that can communicate with any server device via a wireless communication network such as 4G-5G (Generation) or LTE (Long Term Evolution). The user terminal 10 also has a screen such as an LCD display that has touch panel functionality, and may accept various operations on displayed data such as content from the user, such as tapping, sliding, and scrolling, using a finger or stylus. In Figures 4A-4C, the user terminal 10 is used by user U1.

[0082] In Figures 4A to 4C, the information processing device 100 acquires the MIDI data to be processed. For example, the information processing device 100 acquires the MIDI data provided by user U1 via user terminal 10 as the MIDI data to be processed.

[0083] Furthermore, for example, the information processing device 100 may acquire MIDI data to be processed by user U1 via user terminal 10, through a predetermined application or on the web.

[0084] Furthermore, for example, the information processing device 100 may acquire MIDI data created by a user U1 or the like inputting a melody or other idea using a keyboard, as the MIDI data to be processed.

[0085] In Figure 4A, the information processing device 100 acquires a One Shot sample along with the MIDI data to be processed (step S1). For example, the information processing device 100 may acquire a One Shot sample that has been prepared in advance for a specific service, or it may acquire a One Shot sample provided by user U1 via user terminal 10.

[0086] In Figure 4B, the information processing device 100 acquires the timbre category along with the MIDI data to be processed (step S2). For example, the information processing device 100 may acquire the timbre category specified by user U1 via user terminal 10.

[0087] In step S2, the information processing device 100 selects a One Shot sample based on the timbre category (step S3). For example, the information processing device 100 may select a One Shot sample that is registered as the optimal One Shot sample for the timbre category.

[0088] In steps S1 and S3, the information processing device 100 generates an audio waveform by pasting and synthesizing the acquired or selected One Shot samples along with the MIDI data to be processed (step S4).

[0089] Then, the information processing device 100 presents the generated audio waveform to the user U1 (step S5) and accepts edits from the user U1 (step S6).

[0090] In Figure 4C, the information processing device 100 acquires an audio waveform that has been pre-synthesized by a synthesizer or the like (step S7). For example, the information processing device 100 may acquire a synthesized audio waveform provided by user U1 via user terminal 10, which is synthesized by a synthesizer or the like.

[0091] In steps S6 and S7, the information processing device 100 inputs the audio waveform into a pre-prepared deep generation model and performs noise / denosification to convert it into a high-quality audio waveform (step S8). The information processing device 100 then provides the converted audio data waveform to the user U1 (step S9).

[0092] By using the information processing according to the above embodiment, it is possible to automatically generate high-quality audio even in domains where audio and MIDI data pairs are absent or scarce. In other words, by using the information processing according to the above embodiment, annotation-free MIDI-to-Audio can be realized.

[0093] In the above embodiment, using One Shot samples enables the generation of audio waveforms that are highly editable and have a wide range of applications. Because the use of One Shot samples improves the editability of the synthesis, it becomes applicable to a variety of audio generation processes.

[0094] Furthermore, using One Shot samples makes it possible to create music that matches the music of the DTM (Desktop Music) era. Since DTM is a form of music that is commonly heard in everyday life, matching it to DTM makes it possible to create music that is familiar to listeners.

[0095] (1-4. Configuration of the user terminal according to the embodiment) Next, the configuration of the user terminal 10 according to the embodiment will be described using Figure 5. Figure 5 is a diagram showing an example of the configuration of the user terminal 10 according to the embodiment. As shown in Figure 5, the user terminal has a communication unit 11, an input unit 12, an output unit 13, and a control unit 14.

[0096] The communication unit 11 is implemented by, for example, a NIC (Network Interface Card) or a Network Interface Controller. The communication unit 11 is connected to the network N by wire or wireless connection and transmits and receives information with the information processing device 100, etc., via the network N. The network N is implemented by, for example, a wireless communication standard or method such as Bluetooth®, the Internet, Wi-Fi®, UWB (Ultra Wide Band), or LPWA (Low Power Wide Area).

[0097] The input unit 12 accepts various operations from the user. In Figures 4A to 4C, it accepts various operations from user U1. For example, the input unit 12 may accept various operations from the user via the display surface using a touch panel function. Alternatively, the input unit 12 may accept various operations from buttons provided on the user terminal 10, or from a keyboard or mouse connected to the user terminal 10.

[0098] The output unit 13 is a display screen for a tablet terminal, for example, which is implemented using a liquid crystal display or an organic EL (Electro-Luminescence) display, and is a display device for displaying various types of information. For example, the output unit 13 displays information transmitted from the information processing device 100.

[0099] The control unit 14 is implemented, for example, by a CPU (Central Processing Unit), MPU (Micro Processing Unit), GPU (Graphics Processing Unit), etc., which executes a program stored inside the user terminal 10 using RAM (Random Access Memory) or the like as a working area. The control unit 14 is also a controller and may be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gate Array), or MCU (Micro Controller Unit).

[0100] As shown in Figure 5, the control unit 14 has a receiving unit 141 and a transmitting unit 142, and realizes or executes the information processing operations described below.

[0101] The receiving unit 141 receives various types of information from other information processing devices, such as the information processing device 100. For example, the receiving unit 141 receives audio data generated in response to MIDI data input, specified, or selected by a user. For example, the receiving unit 141 receives audio data generated in response to MIDI data created by a user inputting a melody or other idea using a keyboard.

[0102] The transmitting unit 142 transmits various types of information to other information processing devices, such as the information processing device 100. For example, the transmitting unit 142 transmits MIDI data that a user has input, specified, or selected. For example, the transmitting unit 142 transmits MIDI data created by a user inputting a melody or other idea using a keyboard.

[0103] Furthermore, for example, the transmission unit 142 transmits One Shot sample information that is input, specified, or selected by the user. For example, the transmission unit 142 transmits One Shot sample information that is input, specified, or selected by the user, along with MIDI data that is input, specified, or selected by the user.

[0104] Furthermore, for example, the transmission unit 142 transmits timbre category information that is input, specified, or selected by the user. For example, the transmission unit 142 transmits timbre category information that is input, specified, or selected by the user, along with MIDI data that is input, specified, or selected by the user.

[0105] Furthermore, for example, the transmission unit 142 transmits synthesized audio data that is input, specified, or selected by the user. For example, the transmission unit 142 transmits audio data that has been pre-synthesized by a synthesizer or other electronic engineering techniques.

[0106] (1-5. Configuration of the information processing device according to the embodiment) Next, the configuration of the information processing device 100 according to the embodiment will be described using Figure 6. Figure 6 is a diagram showing an example of the configuration of the information processing device 100 according to the embodiment. As shown in Figure 6, the information processing device 100 has a communication unit 110, a storage unit 120, and a control unit 130. The information processing device 100 may also have an input unit (for example, a keyboard or mouse) that receives various operations from the administrator of the information processing device 100, and a display unit (for example, a liquid crystal display) for displaying various information.

[0107] The communication unit 110 is implemented, for example, by a NIC or a network interface controller. The communication unit 110 is connected to the network N by wire or wireless connection and transmits and receives information with the user terminal 10, etc., via the network N. The network N is implemented by wireless communication standards or methods such as Bluetooth, the Internet, Wi-Fi, UWB, LPWA, etc.

[0108] The storage unit 120 is implemented by, for example, semiconductor memory elements such as RAM and flash memory, or storage devices such as hard disks, SSDs (Solid State Drives), and optical discs. As shown in Figure 6, the storage unit 120 has a One Shot sample storage unit 121 and a model storage unit 122.

[0109] The One Shot sample storage unit 121 stores information about One Shot samples (corresponding to audio data OD1 to OD4 in Figure 1). Here, Figure 7 shows an example of the One Shot sample storage unit 121 according to this embodiment. As shown in Figure 7, the One Shot sample storage unit 121 has items such as "sample ID", "timbre", "pitch", and "One Shot sample".

[0110] "Sample ID" indicates identification information for identifying a One Shot sample. "Timber" indicates the timbre. For example, "Timber" stores information about the timbre of the instrument being played. For example, "Timber" stores information such as piano or violin. "Pitch" indicates the pitch. For example, "Pitch" stores information about the pitch being played. For example, "Pitch" stores information such as G3# or B3. "One Shot Sample" indicates the audio data of the One Shot sample. In the example shown in Figure 7, conceptual information such as "One Shot Sample #1" and "One Shot Sample #2" is shown as being stored in "One Shot Sample," but in reality, audio waveforms (for example, data with amplitude and time as axes) are stored.

[0111] The model storage unit 122 stores information about the deep generative model to be applied in the second processing. Here, Figure 8 shows an example of the model storage unit 122 according to the embodiment. As shown in Figure 8, the model storage unit 122 has items such as "Model ID" and "Model".

[0112] The "Model ID" indicates identification information for identifying a deep generative model. The "Model" indicates the deep generative model. In the example shown in Figure 8, conceptual information such as "Model #1" and "Model #2" is stored in "Model," but in reality, parameters of the deep generative model are stored there. In addition, information about the noise level to be applied may also be stored in "Model."

[0113] The control unit 130 is implemented, for example, by a CPU, MPU, GPU, etc., which executes a program (for example, the information processing program according to this disclosure) stored inside the information processing device 100 using RAM or the like as a working area. The control unit 130 is also a controller and may be implemented by an integrated circuit such as an ASIC, FPGA, or MCU.

[0114] As shown in Figure 6, the control unit 130 includes an acquisition unit 131, a first generation unit 132, an editing unit 133, a second generation unit 134, and a providing unit 135, and realizes or executes the information processing operations described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in Figure 6, and other configurations are also acceptable as long as they perform the information processing described later.

[0115] The acquisition unit 131 acquires various types of information. For example, the acquisition unit 131 acquires information transmitted from the user terminal 10 or from external information processing devices such as server devices or cloud systems that provide or manage specific services.

[0116] In the first processing step, the acquisition unit 131 acquires, for example, MIDI data (corresponding to MIDI data MD1 in Figure 1). For example, the acquisition unit 131 acquires MIDI data that is input, specified, or selected by a user. For example, the acquisition unit 131 acquires MIDI data created by a user inputting a melody or other tune using a keyboard.

[0117] In the first processing step, the acquisition unit 131 acquires, for example, One Shot sample information (corresponding to audio data OD1 to OD4 in Figure 1). For example, the acquisition unit 131 acquires One Shot sample information that is input, specified, or selected by the user. For example, the acquisition unit 131 acquires One Shot sample information that is input, specified, or selected by the user, along with MIDI data that is input, specified, or selected by the user.

[0118] In the first processing, the acquisition unit 131 acquires, for example, timbre category information. For example, the acquisition unit 131 acquires timbre category information that is input, specified, or selected by the user. For example, the acquisition unit 131 acquires timbre category information that is input, specified, or selected by the user, along with MIDI data that is input, specified, or selected by the user.

[0119] In the first processing, the acquisition unit 131 acquires, for example, synthesized audio data. For example, the acquisition unit 131 acquires synthesized audio data that has been input, specified, or selected by a user or the like. For example, the acquisition unit 131 acquires audio data that has been pre-synthesized by a synthesizer or other electronic engineering techniques.

[0120] In the second process, the acquisition unit 131 acquires, for example, the audio data to be processed (hereinafter referred to as "first audio data" as appropriate). The first audio data is, for example, audio data generated based on MIDI data. For example, the first audio data is audio data generated by pasting and synthesizing One Shot samples along with MIDI data.

[0121] This One Shot sample is, for example, audio data that has been pre-prepared as a One Shot sample to be used for pasting and compositing. For example, this One Shot sample is audio data that has been pre-prepared as a One Shot sample to be used for pasting and compositing, regardless of whether it is MIDI data or not. For example, it is audio data that has been pre-prepared for a specific service, etc., as freely obtainable audio data.

[0122] Furthermore, this One Shot sample is, for example, audio data selected from pre-prepared audio data according to the sound category specified by the user. For example, this One Shot sample is audio data selected from pre-prepared audio data according to the sound category selected by the user from among the sound category candidates presented to the user according to the sound category specified by the user.

[0123] For example, if a user specifies a sound category using a broad unit such as "piano," the system will present the user with a comment such as, "Generally, there are several categories for pianos, but which category would you prefer?", along with a list of sound category options such as "grand piano," "upright piano," and "digital piano." If the user selects the "digital piano" sound category from these options, the system will then select the corresponding One Shot sample from the One Shot sample group corresponding to the "digital piano" sound category.

[0124] Furthermore, the first audio data is, for example, edited audio data that has been edited via a screen presented to the user to enable editing of the first audio data. For example, the first audio data is newly generated audio data that is created by the user editing at least one of the MIDI data and the One Shot sample data, and then pasting and synthesizing the edited data again.

[0125] The first generation unit 132 generates first audio data, which will be the audio data to be processed in the second process, based on MIDI data acquired by the acquisition unit 131, for example. For example, the first generation unit 132 generates first audio data based on MIDI data and a One Shot sample. Alternatively, for example, the first generation unit 132 generates first audio data based on MIDI data and a timbre category. For example, the first generation unit 132 generates first audio data based on MIDI data and a One Shot sample selected according to a timbre category specified by the user.

[0126] Furthermore, the first generation unit 132 generates first audio data based on edited data (such as edited MIDI data or One Shot samples) edited by the editing unit 133, which will be described later.

[0127] The editing unit 133 performs processing to enable editing of the first audio data generated by the first generation unit 132. For example, the editing unit 133 displays a screen that enables playback of the generated first audio data. Alternatively, for example, the editing unit 133 displays a screen that enables editing of MIDI data or One Shot samples in order to edit the first audio data. In this case, for example, the screen that enables playback of the first audio data may include operation buttons that enable editing of MIDI data or One Shot samples. By operating these operation buttons (for example, by clicking or tapping), the user may transition to a screen where the MIDI data or One Shot samples can be edited.

[0128] Figure 9 is Figure (1) showing an example of a UI screen according to the embodiment. Here, an example of a UI screen for editing according to the embodiment is shown. Screen G1 is a screen displayed on the user terminal 10. Screen G1 includes, for example, an operation button B1 for playing the first audio data generated by the first generation unit 132. By operating this operation button B1, the first audio data is played. In other words, the user can check what kind of music the generated first audio data is actually. If there is image data (including still images and moving images) corresponding to the first audio data, the image data may also be played. For example, this may be the case when the user wants to create music that matches the image data.

[0129] Furthermore, screen G1 includes an operation button B2 for editing the MIDI data used to synthesize the first audio data. By operating this operation button B2, a screen where the MIDI data can be edited is displayed. The user can, for example, listen to what the generated first audio data actually sounds like and make edits such as deleting unnecessary notes and adding necessary notes.

[0130] Furthermore, screen G1 includes an operation button B3 for editing, for example, the One Shot samples used to synthesize the first audio data. By operating this operation button B3, a screen where the One Shot samples can be edited is displayed. The user can then edit the generated first audio data, for example, by listening to what the music actually sounds like and replacing some of the One Shot samples with One Shot samples of other timbres, or by changing the dynamics of some of the One Shot samples.

[0131] Although not shown in Figure 9, screen G1 may include operation buttons to enable changing the sound category. By operating these operation buttons, the user may be redirected to a screen where the sound category can be changed.

[0132] Furthermore, screen G1 is an example of a UI screen according to the embodiment, and the UI screen for performing editing according to the embodiment is not limited to this example.

[0133] The second generation unit 134 generates audio data (hereinafter referred to as "second audio data") that has improved quality compared to the first audio data, based on the first audio data generated by the first generation unit 132 (which may also be the edited first audio data edited by the editing unit 133). For example, the second generation unit 134 generates second audio data that mimics a real performance in a real space compared to the first audio data (for example, second audio data whose quality has been improved so that it is closer to a real performance in a real space than a mechanical performance). In other words, the second generation unit 134 converts the first audio data into second audio data.

[0134] The second generation unit 134 generates second audio data by inputting the first audio data to, for example, a pre-prepared deep generative model. For example, the second generation unit 134 generates second audio data using a pre-trained deep generative model that has been trained to improve the quality of audio data by adding noise to remove some information and then performing denoising to fill in the missing information.

[0135] The second generation unit 134 generates second audio data by, for example, adding noise of a randomly selected noise level to the first audio data to cause some information to be lost, and then performing denoising to compensate for the lost information.

[0136] The second generation unit 134 generates second audio data by, for example, adding noise at a noise level selected based on user operation to the first audio data, causing some information to be lost, and then performing denoising to fill in the lost information. For example, the second generation unit 134 generates second audio data by adding noise at a noise level selected via a screen presented to the user to enable editing of the first audio data, causing some information to be lost, and then performing denoising to fill in the lost information.

[0137] Figure 10 is Figure (2) showing an example of a UI screen according to the embodiment. Here, an example of a UI screen for adjusting the noise level according to the embodiment is shown. Screen G2 is a screen displayed on the user terminal 10. Screen G2 includes, for example, an operation button B11 for adjusting the noise level. This operation button B11 is a slider and is operated by the user sliding the operation button B11 left or right. By operating this operation button B11, the noise level to be applied is determined. In other words, the user can freely determine parameters such as how much level of noise to add, how much to change, or how much change to allow. In Figure 10, the user can freely determine the noise level from 0 to 100 by operating the operation button B11. In Figure 10, the user has adjusted and set the noise level to "40". Also, when the operation button B12 included in screen G2 is operated, a process is executed to generate second audio data with the noise level adjusted by the user with operation button B11.

[0138] The providing unit 135 provides, for example, second audio data generated by the second generation unit 134. For example, the providing unit 135 provides information about the second audio data to a user who has input, specified, or selected MIDI data. Also, for example, the providing unit 135 provides information about the second audio data to a user who has input, specified, or selected a One Shot sample. Also, for example, the providing unit 135 provides information about the second audio data to a user who has input, specified, or selected a timbre category. Also, for example, the providing unit 135 provides information about the second audio data to a user who has input, specified, or selected synthesized audio data.

[0139] The providing unit 135 may provide the second audio data in any manner that allows the user to obtain it. For example, the providing unit 135 may directly transmit the second audio data to the user terminal 10, or it may transmit the second audio data to a server device or cloud system of a specific service that allows the user to obtain it.

[0140] Next, the processing of each part constituting the information processing device 100 described above will be explained in detail using Figures 11 to 13, following the flow. Figure 11 is a flowchart (1) showing the flow of the generation process (first process) in the control unit 130.

[0141] When the information processing device 100 receives MIDI data from the user, it determines whether or not it has received a specification of a timbre category from the user (step S11).

[0142] If the information processing device 100 determines that it has received a specification of a tone category from the user (step S11; YES), it selects a One Shot sample corresponding to the tone category (step S12).

[0143] On the other hand, if the information processing device 100 determines that it has not received a specification of a tone category from the user (step S11; NO), it obtains a pre-prepared One Shot sample regardless of the tone category (step S13).

[0144] The information processing device 100 generates an audio waveform by pasting and synthesizing One Shot samples according to the MIDI data (step S14).

[0145] Figure 12 shows variations of the generation process in the control unit 130. Figure 12 is a flowchart (2) showing the flow of the generation process (first process) in the control unit 130.

[0146] When the information processing device 100 receives MIDI data from the user, it determines whether or not it has received a One Shot sample specification from the user (step S21).

[0147] If the information processing device 100 determines that it has received a request for a One Shot sample from the user (step S21; YES), it obtains the One Shot sample received from the user (step S22).

[0148] On the other hand, if the information processing device 100 determines that it has not received a One Shot sample specification from the user (step S21; NO), it determines whether or not it has received a timbre category specification from the user (step S23).

[0149] If the information processing device 100 determines that it has received a specification of a tone category from the user (step S23; YES), it selects a One Shot sample corresponding to the tone category (step S24).

[0150] On the other hand, if the information processing device 100 determines that it has not received a specification of a tone category from the user (step S23; NO), it obtains a pre-prepared One Shot sample regardless of the tone category (step S25).

[0151] The information processing device 100 generates an audio waveform by pasting and synthesizing One Shot samples according to the MIDI data (step S26).

[0152] Next, the information processing device 100 performs generation processing using a deep generative model or the like. This processing will be explained using Figure 13. Figure 13 is a flowchart showing the flow of generation processing (second processing) in the control unit 130.

[0153] When the information processing device 100 acquires the first audio data to be processed, it determines whether or not it has received a specification from the user regarding the noise level to be applied to the acquired first audio data (step S31).

[0154] If the information processing device 100 determines that it has received a noise level specification from the user (step S31; YES), it decides to apply the noise level received from the user (step S32).

[0155] On the other hand, if the information processing device 100 determines that it has not received a noise level specification from the user (step S31; NO), it decides to randomly select and apply a noise level (step S33).

[0156] The information processing device 100 generates second audio data by adding noise at a predetermined noise level to the acquired first audio data and then performing denoising (step S34).

[0157] The information processing device 100 then provides the generated second audio data (step S35).

[0158] (2. Other Embodiments) The processes described in each embodiment above may be carried out in various other forms besides those described above.

[0159] Of the processes described in the embodiments of this disclosure described above, all or part of the processes described as being performed automatically may be performed manually, or all or part of the processes described as being performed manually may be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings may be changed at will unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.

[0160] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.

[0161] Furthermore, the embodiments of this disclosure described above can be combined as appropriate in areas that do not contradict the processing content. Also, the steps shown in the sequence diagram or flowchart of this embodiment can be changed in order as appropriate. For example, each step may be processed chronologically, repeatedly, or partially in parallel.

[0162] Furthermore, the effects described herein are merely illustrative and not limiting; other effects may also occur.

[0163] (3. Effects of the information processing device related to this disclosure) As described above, the information processing device according to this disclosure (information processing device 100 in the embodiment) comprises an acquisition unit (acquisition unit 131 in the embodiment) and a generation unit (second generation unit 134 in the embodiment). The acquisition unit acquires first audio data to be processed. The generation unit inputs the first audio data to a pre-prepared deep generation model to generate second audio data of higher quality than the first audio data.

[0164] Thus, the information processing device described herein can enable the automatic generation of high-quality audio for a wide variety of complex music, including general music. For example, the information processing device can enable the automatic generation of high-quality audio even for music where there is little or no training data for MIDI data for the audio.

[0165] Furthermore, the generation unit generates second audio data that mimics an actual performance in a real space.

[0166] In this way, the information processing device can enable the automatic generation of high-quality audio that mimics actual performances in real space.

[0167] Furthermore, the generation unit generates second audio data using a deep generative model that has been trained to improve the quality of audio data by adding noise to remove some information and then performing denoising to fill in the missing information.

[0168] In this way, the information processing device can add noise to eliminate unrealistic, pasted-on sections and unnatural transitions in sound, enabling the generation of high-quality audio that is highly realistic and close to actual performance.

[0169] Furthermore, the generation unit uses a deep generative model to add noise of a randomly selected noise level to the first audio data, causing some information to be lost. Denoising is then performed to compensate for the lost information, thereby generating the second audio data.

[0170] In this way, information processing devices can be made easy for users to operate, thus simplifying the automation of high-quality audio generation that is highly realistic and closely resembles actual performances.

[0171] Furthermore, the generation unit uses a deep generation model to add noise at a noise level selected based on user input to the first audio data, causing some information to be lost. Denoising is then performed to compensate for the lost information, thereby generating the second audio data.

[0172] In this way, by allowing users to freely select the noise level, the information processing device can generate high-quality audio that is highly realistic and close to actual performance, taking into account the user's intentions and requests.

[0173] Furthermore, when generating the first audio data, the generation unit generates the second audio data using noise at a noise level selected via a screen presented to the user to enable editing of the first audio data.

[0174] In this way, the information processing device can provide a highly user-friendly music production system that facilitates collaboration between AI and humans by allowing users to actually listen to audio and select noise levels on the UI screen.

[0175] Furthermore, the acquisition unit acquires the first audio data generated based on the MIDI data.

[0176] In this way, the information processing device can enable the automatic generation of high-quality audio that conforms to MIDI data.

[0177] Furthermore, the acquisition unit acquires the first audio data, which is generated by pasting and synthesizing One Shot samples according to the MIDI data.

[0178] In this way, the information processing device can enable the automatic generation of high-quality audio based on MIDI data / One Shot samples.

[0179] Furthermore, the acquisition unit acquires the first audio data using a One Shot sample that has been prepared in advance as a One Shot sample to be used for pasting and compositing.

[0180] In this way, the information processing device can automatically generate high-quality audio that conforms to pre-prepared One Shot samples in MIDI data or specific services.

[0181] Furthermore, the acquisition unit acquires the first audio data using a One Shot sample selected according to the timbre category specified by the user.

[0182] In this way, the information processing device can enable the automatic generation of high-quality audio that conforms to MIDI data / timbre categories.

[0183] Furthermore, the acquisition unit acquires the first audio data using a One Shot sample selected by the user from among the candidate timbre categories presented to the user according to the timbre category.

[0184] In this way, the information processing device can enable the automatic generation of high-quality audio based on One Shot samples selected according to the MIDI data / timbre category.

[0185] Furthermore, the acquisition unit acquires the edited audio data, which has been edited via a screen presented to the user to enable editing of the first audio data, as the first audio data when the first audio data is generated.

[0186] In this way, by enabling users to actually listen to audio and edit it on the UI screen, the information processing device can provide a music production system with excellent usability that facilitates collaboration between AI and humans.

[0187] Furthermore, the acquisition unit acquires audio data based on the edited MIDI data edited by the user via the screen as the first audio data.

[0188] In this way, the information processing device enables the automatic generation of high-quality audio that conforms to the edited MIDI data by allowing users to actually listen to the audio and edit the MIDI data on the UI screen.

[0189] Furthermore, the acquisition unit acquires audio data that has been pre-synthesized using a predetermined electronic engineering method as first audio data.

[0190] In this way, information processing devices can simplify the automation of high-quality audio generation by using audio data that has been pre-synthesized using electronic engineering techniques such as synthesizers.

[0191] (4. Hardware Configuration) The information processing device 100 and the like according to the embodiments of this disclosure described above are realized by a computer 1000 having a configuration such as that shown in Figure 14. The information processing device 100 will be explained as an example. Figure 14 is a hardware configuration diagram showing an example of a computer 1000 that realizes the functions of the information processing device 100. The computer 1000 has a processing circuitry 1100, RAM 1200, ROM 1300, secondary storage device 1400, communication interface 1500, input / output interface 1600, display unit 1700, camera unit 1800, microphone 1900, and speaker 2000. The parts of the computer 1000 are connected by a bus 1050.

[0192] The processing circuit 1100 operates based on a program stored in the ROM 1300 or secondary storage device 1400, and controls each part. For example, the processing circuit 1100 loads the program stored in the ROM 1300 or secondary storage device 1400 into the RAM 1200 and executes processing corresponding to various programs.

[0193] ROM1300 stores boot programs such as the BIOS (Basic Input Output System) that are executed by the processing circuit 1100 when the computer 1000 starts up, as well as programs that depend on the computer 1000's hardware.

[0194] The secondary storage device 1400 is a computer-readable recording medium that non-temporarily records programs executed by the processing circuit 1100 and data used by such programs. Specifically, the secondary storage device 1400 is a recording medium that records programs for each process of the information processing device 100 according to the embodiment of this disclosure, which is an example of program data 1450.

[0195] The communication interface 1500 is an interface for the computer 1000 to connect to the external network 1550. The communication interface 1500 corresponds to the communication unit 110 of the information processing device 100. For example, the processing circuit 1100 receives data from other devices or transmits data generated by the processing circuit 1100 to other devices via the communication interface 1500.

[0196] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the processing circuit 1100 receives data from input devices such as a microphone 1900 or a touch panel via the input / output interface 1600. The processing circuit 1100 also transmits data to output devices such as a display unit 1700 or a speaker 2000 via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium (media). Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical Disks), tape media, magnetic recording media, or semiconductor memory.

[0197] The display unit 1700 is an interface for displaying information processed by the computer 1000. The display unit 1700 is, for example, a liquid crystal display or an organic electroluminescent display (Organic Electro Luminescence Display). Alternatively, the display unit 1700 may be a touch panel display device or an image projection device.

[0198] The camera unit 1800 is an interface for the computer 1000 to capture images. The microphone 1900 is an interface for the computer 1000 to capture audio. The speaker 2000 is an interface for the computer 1000 to output the processed audio. The various parts of the computer 1000 are connected by the bus 1050. Each interface does not necessarily have to be located inside the computer 1000; it may be located outside the computer 1000 via a network or the like. Furthermore, each part of the computer 1000 may be controlled by a circuit different from the processing circuit 1100. For example, the display unit 1700 may be controlled not by the processing circuit 1100, but by a circuit dedicated to display processing that is provided within the display unit 1700.

[0199] For example, when computer 1000 functions as an information processing device 100 according to an embodiment of this disclosure, the processing circuit 1100 of computer 1000 functions as a control unit 130 by executing a program loaded onto RAM 1200. The secondary storage device 1400 stores the information processing program according to this disclosure and various data stored by the storage unit 120. The processing circuit 1100 reads and executes program data 1450 from the secondary storage device 1400, but as another example, these programs may be obtained from other devices via an external network 1550. In other words, the secondary storage device 1400 is not limited to being inside computer 1000, but may be located outside computer 1000. The processing circuit 1100 is an example of an integrated circuit, and CPU, MPU, GPU, APU, ASIC, and FPGA can all be considered integrated circuits.

[0200] Furthermore, this technology can also be configured as follows. (1) An acquisition unit that acquires the first audio data to be processed, A generation unit that generates a second audio data of higher quality than the first audio data by inputting the first audio data into a pre-prepared deep generation model, An information processing device equipped with the following features. (2) The generating unit is Generates the preceding second audio data that mimics a real-world performance. The information processing device described in (1) above. (3) The generating unit is The second audio data is generated using the deep generative model, which has been trained to improve the quality of audio data by adding noise to remove some information and then performing denoising to compensate for the missing information. The information processing device described in (1) or (2) above. (4) The generating unit is Using the deep generation model, the second audio data is generated by adding noise of a randomly selected noise level to the first audio data to remove some information, and then performing denoising to restore the missing information. An information processing device as described in any one of (1) to (3) above. (5) The generating unit is Using the deep generation model, the second audio data is generated by adding noise at a noise level selected based on user input to the first audio data, causing some information to be lost, and then performing denoising to compensate for the lost information. An information processing device as described in any one of (1) to (4) above. (6) The generating unit is When generating the first audio data, the second audio data is generated using noise at the noise level selected via a screen presented to the user to enable editing of the first audio data. The information processing device described in (5) above. (7) The acquisition unit is, Obtain the first audio data generated based on the MIDI data. An information processing device as described in any one of (1) to (6) above. (8) The acquisition unit is, The first audio data is obtained by pasting and combining One Shot samples according to the aforementioned MIDI data. The information processing device described in (7) above. (9) The acquisition unit is, The first audio data is acquired using the One Shot sample, which has been prepared in advance as the One Shot sample used for the paste synthesis. The information processing device described in (8) above. (10) The acquisition unit is, The first audio data is acquired using the One Shot sample selected according to the timbre category specified by the user. The information processing device described in (8) above. (11) The acquisition unit is, The first audio data is acquired using the One Shot sample selected according to the timbre category selected by the user from among the candidate timbre categories presented to the user according to the timbre category. The information processing device described in (10) above. (12) The acquisition unit is, When generating the first audio data, the edited audio data, which has been edited via a screen presented to the user to enable editing of the first audio data, is acquired as the first audio data. An information processing device as described in any one of (1) to (11) above. (13) The acquisition unit is, The audio data based on the edited MIDI data edited by the user via the aforementioned screen is acquired as the first audio data. The information processing device described in (12) above. (14) The acquisition unit is, Audio data that has been pre-synthesized by a predetermined electronic engineering method is acquired as the first audio data. An information processing device as described in any one of (1) to (5) above. (15) Information processing device, The process involves acquiring the first audio data to be processed, A generation process that generates a second audio data of higher quality than the first audio data by inputting the first audio data into a pre-prepared deep generative model, Information processing methods including (16) Computers, The procedure for obtaining the first audio data to be processed, A generation procedure that generates a second audio data of higher quality than the first audio data by inputting the first audio data into a pre-prepared deep generative model, An information processing program designed to function as such. [Explanation of Symbols]

[0201] 1. Information Processing System 10. User terminals 11 Communications Department 12 Input section 13 Output section 14 Control Unit 100 Information Processing Devices 110 Communications Department 120 Storage section 121 One Shot Sample Storage Unit 122 Model Memory Unit 130 Control Unit 131 Acquisition Department 132 1st generation part 133 Editorial Department 134 Second generation part 135 Provision Department 141 Receiving Unit 142 Transmitter N Network

Claims

1. An acquisition unit that acquires the first audio data to be processed, A generation unit that generates a second audio data of higher quality than the first audio data by inputting the first audio data into a pre-prepared deep generation model, An information processing device equipped with the following features.

2. The generating unit is This generates the aforementioned second audio data that mimics a real-world performance. The information processing apparatus according to claim 1.

3. The generating unit is The second audio data is generated using the deep generative model, which has been trained to improve the quality of audio data by adding noise to remove some information and then performing denoising to compensate for the missing information. The information processing apparatus according to claim 1.

4. The generating unit is Using the deep generation model, the second audio data is generated by adding noise of a randomly selected noise level to the first audio data to remove some information, and then denoising to restore the missing information. The information processing apparatus according to claim 1.

5. The generating unit is Using the deep generation model, the second audio data is generated by adding noise at a noise level selected based on user input to the first audio data, causing some information to be lost, and then performing denoising to compensate for the lost information. The information processing apparatus according to claim 1.

6. The generating unit is When generating the first audio data, the second audio data is generated using noise at the noise level selected via a screen presented to the user to enable editing of the first audio data. The information processing apparatus according to claim 5.

7. The acquisition unit is, The first audio data generated based on MIDI data is obtained. The information processing apparatus according to claim 1.

8. The acquisition unit is, The first audio data is obtained by pasting and combining One Shot samples according to the aforementioned MIDI data. The information processing apparatus according to claim 7.

9. The acquisition unit is, The first audio data is acquired using the One Shot sample, which has been prepared in advance as the One Shot sample used for pasting and combining. The information processing apparatus according to claim 8.

10. The acquisition unit is, The first audio data is acquired using the One Shot sample selected according to the timbre category specified by the user. The information processing apparatus according to claim 8.

11. The acquisition unit is, The first audio data is acquired using the One Shot sample selected by the user from among the candidate timbre categories presented to the user according to the timbre category. The information processing apparatus according to claim 10.

12. The acquisition unit is, When generating the first audio data, the edited audio data, which has been edited via a screen presented to the user to enable editing of the first audio data, is acquired as the first audio data. The information processing apparatus according to claim 1.

13. The acquisition unit is, The audio data based on the edited MIDI data edited by the user via the aforementioned screen is acquired as the first audio data. The information processing apparatus according to claim 12.

14. The acquisition unit is, Audio data that has been pre-synthesized by a predetermined electronic engineering method is acquired as the first audio data. The information processing apparatus according to claim 1.

15. Information processing device, The process involves acquiring the first audio data to be processed, A generation process that generates a second audio data of higher quality than the first audio data by inputting the first audio data into a pre-prepared deep generative model, Information processing methods including

16. Computers, The procedure for obtaining the first audio data to be processed, A generation procedure that generates a second audio data of higher quality than the first audio data by inputting the first audio data into a pre-prepared deep generative model, An information processing program designed to function as such.