Information processing apparatus, method, and non-transitory computer-readable medium

The information processing apparatus addresses the challenge of generating high-quality audio from MIDI data by adding noise and denoising it, allowing for realistic and detailed tone synthesis across various music genres and instruments, enhancing music production capabilities.

WO2026070359A1PCT designated stage Publication Date: 2026-04-02SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Conventional MIDI-to-Audio techniques require paired audio and MIDI data for training, which is difficult to obtain, especially for complex music, and cannot designate subtle tone differences, limiting their applicability and usability in music production.

Method used

An information processing apparatus that generates high-quality audio by adding noise to audio data and denoising it using a deep generative model, allowing for the synthesis of realistic audio waveforms from MIDI data and One Shot samples, enabling detailed tone designation and handling various music genres and instruments.

Benefits of technology

Enables the automatic generation of high-quality, realistic audio that mimics professional musical performance, overcoming the limitations of conventional methods by providing a system that can handle diverse music styles and subtle tone variations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025031839_02042026_PF_FP_ABST
    Figure JP2025031839_02042026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing apparatus according to the present disclosure includes: an acquisition unit that acquires first audio data to be processed; and a generation unit that generates second audio data of higher quality than the first audio data by inputting the first audio data to a deep generative model prepared in advance.
Need to check novelty before this filing date? Find Prior Art

Description

INFORMATION PROCESSING APPARATUS, METHOD, AND NON-TRANSITORY COMPUTER-READABLE MEDIUM

[0001] The present disclosure relates to an information processing apparatus, a method, and a non-transitory computer-readable medium.

[0002] In recent years, a technique for automatically generating audio such as music has attracted attention. For example, a technique called Musical Instruments Digital Interface (MIDI), which generates audio from MIDI data in which musical performance information such as which pitch, how strong, and for how long is played, is recorded, is attracting attention.

[0003] C. Meng et al. "SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations" in Proceedings of ICLR, 2022J. Gardner et al. "MT3: multi-task multitrack music transcription" in Proceedings of ICLR, 2021S. Ian et al. "Scaling polyphonic transcription with mixtures of monophonic transcriptions" in Proceedings of ISMIR, 2022R. M. Bittner et al. "A lightweight instrument-agnostic model for polyphonic note transcription and multipitch estimation" in Proceedings of ICASSP, 2022Y.-T. Wu et al. "Multi-instrument automatic music transcription with self-attention-based instance segmentation" in Proceedings of ICML, 2022

[0004] However, in the conventional art, training data of MIDI data for audio is required to generate audio, and when such training data is poor or unavailable, high-quality audio cannot be generated.

[0005] Therefore, the present disclosure proposes an information processing apparatus, an information processing method, and an information processing program that enable automatic generation of high-quality audio for various music.

[0006] To solve the problems described above, an information processing apparatus according to an embodiment of the present disclosure includes: circuitry configured to: receive first audio data; generate intermediate audio data by adding noise of a predetermined noise level to the first audio data; and generate second audio data by denoising the intermediate audio data, wherein the second audio data is higher quality audio data than the first audio data.

[0007] Fig. 1 is a diagram illustrating an outline of an information processing system according to an embodiment.Fig. 2 is a diagram illustrating an example of MIDI data according to the embodiment.Fig. 3 is a diagram illustrating an example of pasting synthesis according to the embodiment.Fig. 4A is a sequence diagram (1) illustrating an outline of the information processing system according to the embodiment.Fig. 4B is a sequence diagram (2) illustrating an outline of the information processing system according to the embodiment.Fig. 4C is a sequence diagram (3) illustrating an outline of the information processing system according to the embodiment.Fig. 5 is a diagram illustrating a configuration example of a user terminal according to the embodiment.Fig. 6 is a diagram illustrating a configuration example of an information processing apparatus according to the embodiment.Fig. 7 is a diagram illustrating an example of a One Shot sample storage unit according to the embodiment.Fig. 8 is a diagram illustrating an example of a model storage unit according to the embodiment.Fig. 9 is a diagram (1) illustrating an example of a UI screen according to the embodiment.Fig. 10 is a diagram (2) illustrating an example of a UI screen according to the embodiment.Fig. 11 is a flowchart (1) illustrating a flow of generation processing (primary processing) in a control unit.Fig. 12 is a flowchart (2) illustrating a flow of generation processing (primary processing) in the control unit.Fig. 13 is a flowchart illustrating a flow of generation processing (secondary processing) in the control unit.Fig. 14 is a hardware configuration diagram illustrating an example of a computer that achieves a function of the information processing apparatus.

[0008] The embodiment of the present disclosure will be described below in detail on the basis of the drawings. Note that, in each embodiment described below, the same parts are designated by the same reference numerals, and duplicate description will be omitted.

[0009] The present disclosure will be described in the order of items described below. 1. Embodiment 1-1. Outline of the information processing apparatus according to the embodiment 1-2. Outline 1 of the information processing system according to the embodiment 1-3. Outline 2 of the information processing system according to the embodiment 1-4. Configuration of the information processing apparatus according to the embodiment 2. Other embodiments 3. Effects of the information processing apparatus according to present disclosure 4. Hardware configuration

[0010] (1. Embodiment) (1-1. Outline of the information processing apparatus according to the embodiment) In recent years, a technique for automatically generating MIDI data from audio has attracted attention. Such a technique is also called Automatic Music Transcription (AMT), and is an important technique in music production and the like. As a technique related to AMT, a technique that enables automatic generation of MIDI data by causing an annotation for audio to be learned as training data is known.

[0011] As described above, while a technique for generating MIDI data from audio is attracting attention, a technique for generating audio from MIDI data is also attracting attention. A technique for generating audio from such MIDI data is also called MIDI-to-Audio, and is an important technique in music production and the like, similarly to AMT.

[0012] As a conventional art related to MIDI-to-Audio, a technique that enables automatic generation of audio by causing MIDI data for audio to be learned as training data is known. As described above, in the conventional art related to MIDI-to-Audio, it is necessary to cause a pair of audio and MIDI data to be learned as training data. For example, the musical performance audio actually played is required with respect to the MIDI data, and this is training data.

[0013] Therefore, it is necessary to provide musical performance audio that matches the timeline of the MIDI data and matches the time of musical performance of the MIDI data. However, in the case of musical performance by a person, it is easy to imagine that a slight deviation occurs even when the person plays along the sheet music while watching the sheet music. For example, it is easy to imagine that a slight deviation occurs due to early or slow release of keys at the time of keystroke. In view of the fact that it is training data for performing machine learning, such a deviation is not preferable, and there is a problem that machine learning is not appropriately performed without musical performance audio that matches the time of musical performance of the MIDI data. In addition, it is also difficult to mechanically edit (or correct or process) such a deviation.

[0014] In addition to the problem that there is a possibility that machine learning cannot be appropriately performed due to such a deviation, for example, as in a case where a creator or the like creates MIDI data for a certain music but does not disclose the data, MIDI data for the music is not necessarily obtained, and there is a problem that a pair of audio and MIDI data cannot be easily obtained.

[0015] As described above, in addition to the fact that it is not easy to obtain many pairs of audio and MIDI data, it is easy to imagine that it is difficult to obtain many pairs of audio and MIDI data that have no deviation and match the musical performance.

[0016] For this reason, in the conventional art related to MIDI-to-Audio, a pair of audio and MIDI data may be obtained for some limited pieces of music, and learning may be possible, but it is difficult to appropriately perform learning with respect to various music with high complexity. For example, some pianos may have a function (such as a sensor) that automatically obtains MIDI data when musical performance is performed, and thus, learning can be performed to some extent for piano music, but there are not many instruments that can be equipped with such a function, and music exists endlessly depending on combinations or the like, and thus it is difficult to appropriately perform learning with respect to any music. In addition, for example, it is easy to imagine that it is difficult to listen to and transcribe (write back) a sound actually played.

[0017] Additionally, the conventional art related to MIDI-to-Audio has a problem that a tone can be designated only in units of categories. For example, in the conventional art, the tone can be designated only in units of instruments such as a piano and a violin. In music production, a subtle tone difference is important. For example, in general, there are various types of pianos such as a grand piano, an upright piano, and a digital piano (electronic piano), and such a subtle tone difference is important in music production. Creators and the like carefully select a tone to produce music, focusing on subtle tones. In the conventional art, it is possible to designate only in rough units such as units of categories, and it is difficult to realize practical use in music production.

[0018] The tone according to the embodiment will be further described. The tone is a difference in expression with respect to music, and is not limited to a difference in an instrument, and includes a difference in expression by a music performer. For example, even in the case of playing using the same sheet music and the same instrument, it is easy to imagine that the expression of music varies depending on the music performer. Therefore, the tone can also be understood as an expression for music. In MIDI-to-Audio, for example, in a case where a voice of a certain artist is designated, singing is performed with the voice of the artist. Since the tone can be separately designated, a creator or the like can produce music while separately considering note information and tone information.

[0019] In addition, in recent years, a music generation technique using artificial intelligence (AI) has been developed. For example, it is possible to generate a final audio waveform only by inputting text. For example, it is possible to generate an audio waveform played by a piano only by inputting "music played with a piano". Additionally, for example, by inputting a music genre such as "classical" or "Jazz", it is possible to generate an audio waveform played with a piano along the genre. Further additionally, for example, by inputting lyrics, it is possible to generate an audio waveform as if a person is singing along the lyrics.

[0020] However, since such a technique does not understand information such as MIDI data and tone, it is not possible to appropriately respond to, for example, a request for changing some notes. In this way, it has not been possible to introduce an important parameter for a creator or the like that can control the content of music.

[0021] Therefore, an object of the present disclosure is to enable automatic generation of high-quality audio for various music including general music. For example, an object is to provide a MIDI-to-Audio technique applicable to various music. In addition, for example, an object is to provide a MIDI-to-Audio technique capable of designating a detailed tone instead of a rough unit such as units of categories.

[0022] In addition, an object of the present disclosure is to improve the performance of MIDI-to-Audio even in a domain (for example, an instrument, a tone, an artist such as a singer, a genre, a language, and the like) in which there is no or poor pair of audio and MIDI data, for example. In addition, for example, an object is to enable a more detailed design of a synthesized tone while sticking to a subtle tone of each sound. In addition, for example, an object is to provide a music production system in which AI and a person easily cooperate with each other and which is excellent in usability.

[0023] Hereinafter, an outline of an information processing apparatus 100 according to the embodiment will be described. The information processing apparatus 100 is an information processing apparatus intended to enable automatic generation of high-quality audio for various music, and may be any apparatus as long as processing in the embodiment can be achieved. The information processing apparatus 100 is achieved by, for example, a server apparatus, a cloud system, or the like, and executes information processing according to the embodiment. For example, the information processing apparatus 100 is achieved by a server apparatus, a cloud system, or the like that provides or manages a specific service that allows a creator or the like to freely perform music production or the like. For example, the information processing apparatus 100 is achieved by a server apparatus, a cloud system, or the like that provides or manages a specific service that provides support (for example, making various tools for music production and the like usable) for a creator or the like to freely perform music production or the like.

[0024] As an example, the information processing apparatus 100 uses audio sample data called "One Shot" in the MIDI standard. Hereinafter, such audio sample data called "One Shot" will be appropriately referred to as a "One Shot sample". The One Shot sample is audio data in which a sound is registered in advance, and is, for example, audio data of a sound divided in a certain unit (in this case, pitch) such as a sound of "do" or a sound of "re". For example, the One Shot sample is audio data of each sound such as a sound of "do" and a sound of "re". Note that the One Shot sample is not limited to audio data divided by pitch, and may be audio data of sound divided in any unit. In addition, the One Shot sample may be audio data obtained by recording an actual sound in real space in advance.

[0025] Then, for example, the information processing apparatus 100 performs pasting synthesis on the One Shot samples along the MIDI data to generate an audio waveform, and converts the audio waveform into high-quality audio via a deep generative model or the like (for example, an extended model or the like). For example, the information processing apparatus 100 performs various types of data processing on MIDI data transmitted from an external information processing apparatus (such as a user terminal 10 to be described below), and performs conversion into high-quality audio desired by the user.

[0026] In the following embodiment, converting into high-quality audio refers to, for example, converting into realistic audio close to real musical performance. For example, the audio obtained by performing pasting synthesis on the One Shot samples may not be smooth at a joint (transition) or the like of the One Shot samples, and the sound may change unnaturally and become mechanical (artificial) audio. Therefore, for example, it indicates that the audio is converted so as to be smooth even at a joint of the One Shot samples. For example, it indicates converting audio having a noticeable sense of pasting and having no change in dynamics (strength or expression of sound) or tone to audio close to musical performance by a professional music performer. For example, in the case of a piano, dynamics and tone can be variously changed depending on how to move a hand or an arm and how to move the whole body, and thus, it indicates conversion into audio close to such musical performance.

[0027] Note that these examples are merely examples, and it is not particularly limited thereto. That is, the high-quality audio may be any audio as long as the audio is realistic and close to real musical performance as compared with the audio having a mechanical and noticeable sense of pasting.

[0028] In addition, in the following embodiment, the One Shot sample may be, for example, audio data prepared in advance for a specific service, or may be audio data input by a user for each audio generation. In a case where audio data prepared in advance for a specific service is used for the One Shot sample, the user can generate audio by preparing MIDI data. In addition, when audio data input by the user is used for the One Shot sample, the user prepares the One Shot sample in addition to MIDI data. The following embodiment may be implemented in any form on the premise that the user prepares MIDI data. Note that a method in which the user directly uses an audio waveform synthesized by a predetermined electronic technique (synthesizer or the like) or the like without preparing MIDI data will be described below as a variation.

[0029] In addition, in the following embodiment, MIDI data may be, for example, data created by entering a tune (melody) or the like conceived by the user with a keyboard or the like. In addition, MIDI data may be, for example, data acquired from a specific service that discloses MIDI data of existing music and edited by a user.

[0030] (1-2. Outline 1 of the information processing system according to the embodiment) Next, an outline of an information processing system 1 for describing the technique of the present disclosure will be described with reference to Fig. 1. Fig. 1 is a diagram illustrating an outline of the information processing system 1 according to the embodiment.

[0031] In Fig. 1, the information processing apparatus 100 generates an audio waveform OW1 obtained by performing pasting synthesis on audio data OD1 to OD4, which are One Shot samples, along MIDI data MD1. Then, the information processing apparatus 100 performs processing P1 of adding noise to the audio waveform OW1 such that a portion without reality (a portion with low quality) disappears, thereby generating an audio waveform OW2. Then, the information processing apparatus 100 performs processing P2 for removing noise on the audio waveform OW2 to generate an audio waveform OW3. In Fig. 1, the audio waveform OW3 is a target audio. That is, the quality of the audio is higher than that of the audio waveform OW1. Note that the pieces of the processing P1 and P2 may be performed as one piece of processing. Hereinafter, details of each processing illustrated in Fig. 1 will be described.

[0032] The information processing apparatus 100 roughly performs two pieces of processing. One is processing of generating the audio waveform OW1, and the other is processing of generating the audio waveform OW3. Note that, since the audio waveform OW2 is an audio waveform obtained in the process of generating the audio waveform OW3, the audio waveform OW2 is included in the second piece of processing here. The first processing of generating the audio waveform OW1 is processing for obtaining input information necessary for the second processing of generating the audio waveform OW3. First, the processing of generating the audio waveform OW1 is described, but there is a plurality of variations in the first processing of generating the audio waveform OW1. Hereinafter, the processing of generating the audio waveform OW1 using MIDI data and One Shot samples among the plurality of variations will be described as an example, and the other variations will be described at the end. Note that, hereinafter, the first processing will be described as "primary processing", and the second processing will be described as "secondary processing" as appropriate.

[0033] MIDI data will be described with reference to Fig. 2. Fig. 2 is a diagram illustrating an example of MIDI data according to the embodiment. The MIDI data is data such as digital sheet music, and the horizontal axis indicates time and the vertical axis indicates pitch. The pitch corresponds to a key in the case of a piano, for example.

[0034] In Fig. 2, alphanumeric characters such as "F3", "B3", "D4", and "F4" on the vertical axis are information indicating pitch. In addition, information indicated by each rectangle corresponds to a note. For example, a rectangle corresponding to "F3" on the vertical axis indicates that the note is the pitch of "F3". In addition, the length of the rectangle indicates the time during which the note is played. For example, the length of the rectangle of note N1 is about 2.3 seconds to about 3.4 seconds on the horizontal axis, indicating that it is a note to be played for about 2.3 seconds to about 3.4 seconds.

[0035] In addition, notes N1 to N4 are arranged in parallel on the horizontal axis, which indicates that notes of a plurality of pitches are played simultaneously. Specifically, it indicates that pitches "G3#", "B3", "D4", and "D5" corresponding to the notes N1 to N4 are simultaneously played. For example, in the case of a piano, it indicates that keys corresponding to "G3#", "B3", "D4", and "D5" are simultaneously pressed. For example, it indicates that the notes of the pitches of "G3#", "B3", "D4", and "D5" start to be played at around 2.3rd second from the start and the musical performance of the note N4 ends before around 3rd second. Note that, in Fig. 2, only the first four rectangles are denoted by reference numerals, but this is for convenience of description, and reference numerals may be appropriately assigned to other rectangles.

[0036] Next, pasting synthesis of One Shot samples will be described with reference to Fig. 3. Fig. 3 is a diagram illustrating an example of pasting synthesis according to the embodiment. In Fig. 3, description will be made assuming that by designation of a certain tone category, One Shot samples of various pitches of the tone are obtained. That is, description will be made assuming that a One Shot sample group of a certain tone is obtained. For example, when the designated tone category is a grand piano, a sample group of One Shot samples with pitches such as "F3" to "G5#" corresponding to the grand piano is obtained, and when the designated tone category is a digital piano, a sample group of One Shot samples with pitches such as "F3" to "G5#" corresponding to the digital piano is obtained.

[0037] MIDI data MD1 indicates that note N11 is played from time t1 to time t2, note N12 is played from time t2 to time t3, note N13 is played from time t4 to time t6, and note N14 is played from time t5 to time t7. Note that times t1 to t6 are times arranged in the order of timelines, and times t3 to t4 are portions where a sound does not exist.

[0038] In the note N11, audio data OD1 is selected from the One Shot sample group on the basis of the pitch of the note N11, and a part (a part of the audio data OD1) cut out corresponding to times t1 to t2 is pasted. Similarly, in the note N12, audio data OD2 is selected, and a part (a part of the audio data OD2) cut out corresponding to times t2 to t3 is pasted.

[0039] Similarly, in the note N13, audio data OD3 is selected, and a part (a part of the audio data OD3) cut out corresponding to times t4 to t6 is pasted. Similarly, in the note N14, audio data OD4 is selected, and a part (a part of the audio data OD4) cut out corresponding to times t5 to t7 is pasted. Note that, at times t5 to t6, a part of the audio data OD3 and a part of the audio data OD4 overlap each other when pasted.

[0040] In addition, as for cutting, in each case, the head portion of the audio data is cut out, but the cut-out portion is not particularly limited. For example, in the case of the note N11, the head portion of the audio data OD1 is cut out and pasted, but the middle portion may be cut out and pasted, or the end portion may be cut out and pasted, and it is not particularly limited to this example. In addition, for example, in a case where there is an important portion in the audio data and the portion is designated in advance, the portion may be preferentially cut out and pasted. In addition, for example, in a case where there is an important portion in the audio data and the portion is not designated in advance, the portion may be estimated (for example, a portion having the most dynamics may be estimated as an important portion from the audio waveform), and the estimated portion may be preferentially cut out and pasted.

[0041] In addition, in the MIDI data MD1, since times t3 to t4 are portions where a sound does not exist, the portions are portions where a sound does not exist also in the audio waveform obtained by pasting synthesis. In addition, in the MIDI data MD1, since times t5 to t6 are portions where sounds overlap, the portions are portions where sounds overlap also in the audio waveform obtained by pasting synthesis. Specifically, the note N14 is played during the musical performance of the note N13.

[0042] Note that, in Fig. 3, for convenience of description, it has been described that by designation of a certain tone category, One Shot samples of various pitches of the tone are obtained, but it is not limited to a case where one certain tone category is designated, and a plurality of tone categories may be designated. For example, as the audio data OD1 is a piano and the audio data OD2 is a guitar, the audio waveform OW1 may be audio that is first played with a piano and then changed to a guitar in the middle. As described above, it is not limited to the case where the audio waveform is synthesized using the One Shot samples of one certain tone category, and the audio waveform may be synthesized using the One Shot samples of a plurality of tone categories. For example, the One Shot sample corresponding to each note may be selected from among the One Shot samples of various pitches corresponding to each of the plurality of tone categories to synthesize the audio waveform.

[0043] In addition, an audio waveform of a plurality of tone categories may be synthesized by designating a plurality of tone categories, or audio waveforms obtained by designating one certain tone category may be combined so as to finally synthesize an audio waveform of the plurality of tone categories. In this manner, an audio waveform of a plurality of tone categories may be synthesized by combining audio waveforms themselves obtained by designating the tone category. By using the One Shot sample, the editing property of synthesis is enhanced, and thus, it can be applied to various audio generation.

[0044] In addition, the information processing apparatus 100 may present the audio waveform OW1 generated in this manner to the user to enable editing of the MIDI data MD1 and the One Shot samples such as the audio data OD1 to OD4. For example, the information processing apparatus 100 may enable editing of the One Shot sample by Attack, Decay, Sustain, and Release (ADSR), a filter, or the like. As a result, for example, dynamics, tone, and the like can be edited. For example, the user can actually listen to the audio waveform obtained by pasting synthesis and change a portion that the user wants to change. For example, changes can be made such that, when the user actually listens, when the user feels that a certain note is unnecessary, the note information is deleted from the MIDI data MD1, and when the user feels that a certain note is necessary, its note information is added to the MIDI data MD1.

[0045] The processing in which the information processing apparatus 100 generates the audio waveform OW1 obtained by performing pasting synthesis on the audio data OD1 to OD4 along the MIDI data MD1 has been described above. Hereinafter, processing in which the information processing apparatus 100 converts the audio waveform OW1 into a high-quality audio waveform will be described.

[0046] The audio waveform OW1 includes, though implicitly, the MIDI data MD1, tone information, and the like. The information processing apparatus 100 converts the audio waveform OW1 into a realistic audio waveform while maintaining the MIDI data MD1 and the tone information.

[0047] As an example of a method for converting into a realistic audio waveform, there is a method called a noise / denoise approach. In the noise / denoise approach, noise is added so that a portion without reality disappears. On the other hand, when a large amount of noise is added, an important portion of the MIDI data MD1 and the tone information disappears, so that an appropriate level of noise is added. Therefore, it is necessary to perform adjustment so that the noise to be added becomes an appropriate level. This corresponds to the processing P1 illustrated in Fig. 1. The information processing apparatus 100 generates the audio waveform OW2 by adding noise adjusted to an appropriate level to the audio waveform OW1.

[0048] In addition, the information processing apparatus 100 specifies a noise portion from the audio waveform OW2 to which noise is added, and generates the audio waveform OW3 using the deep generative model that restores audio without noise. This corresponds to the processing P2 illustrated in Fig. 1. Although details will be described below, the information processing apparatus 100 generates the audio waveform OW3 by removing noise so as to obtain realistic audio. At this time, the information processing apparatus 100 may receive a condition such as in which direction noise is removed, and generate the audio waveform OW3 on the basis of the condition. For example, in a case where a condition such as "realistic musical performance" is received, the information processing apparatus 100 may generate the audio waveform OW3 by removing noise so as to obtain realistic musical performance audio.

[0049] Here, details of the processing P1 and P2 illustrated in Fig. 1 will be described. The information processing apparatus 100 generates the audio waveform OW3 using a trained deep generative model prepared in advance.

[0050] The information processing apparatus 100 adds a noise level of time of "T, T-1,..., 1, 0" to the audio waveform OW1. The noise level changes stepwise with time of "T, T-1,..., 1, 0". Then, a noise level at a certain time among the above is added to the audio waveform OW1. For example, the information processing apparatus 100 adds the noise level of time "t" selected from time "T" when the noise level is the highest to time "0" when the noise level is the lowest. For example, the information processing apparatus 100 adds the noise level at time "t" near the middle from time "T" to time "0". In addition, for example, the information processing apparatus 100 causes the user to select time "t" (or may select information corresponding to time "t") from time "T" to time "0", and adds the noise level of the selected time "t". Note that the user's selection of the noise level may be selection using any method. For example, selection based on operation of a user interface (UI) such as a slider is exemplified.

[0051] In a case where the user is caused to select the noise level, the information processing apparatus 100 may cause the user to select the noise level at any stage. For example, the information processing apparatus 100 may perform selection in advance before generating the audio waveform OW1, such as at the time of inputting the MIDI data MD1, or may perform selection at the time of generating (immediately after generating) the audio waveform OW1. For example, the information processing apparatus 100 may present the audio waveform OW1 to the user at the time of generating the audio waveform OW1, and may select the noise level at the same time when editing the MIDI data MD1 and the One Shot sample. For example, the information processing apparatus 100 may simultaneously select the noise level on the same UI screen when enabling editing of the MIDI data MD1 or the One Shot sample. In addition, for example, the information processing apparatus 100 may select the noise level on a UI screen different from the UI screen for presentation of the audio waveform OW1 to the user or for editing of the MIDI data MD1 or the One Shot sample.

[0052] It can also be understood that although some information is missing when noise is added in the above-described processing, the audio waveform OW3 with reality can be generated by adding denoise so as to compensate for the missing portion. By causing the user to select the noise level, for example, it is possible to generate the audio waveform in consideration of the intention or desire of the user, such as how much the change is desired or whether the change is allowed. By causing the user to freely select such parameters, it is possible to provide a music production system excellent in usability.

[0053] The information processing apparatus 100 generates the audio waveform OW3 using a deep generative model disclosed in NPL 1. For example, the information processing apparatus 100 generates the audio waveform OW3 using a deep generative model of Text-to-Music or the like that is conditioned with text and outputs audio. For example, the information processing apparatus 100 generates the audio waveform OW3 in consideration of the purpose such as "high-quality violin performance" when the user inputs information according to the purpose such as "high-quality violin performance" (information indicating what reality is desired) as the text. That is, the information processing apparatus 100 generates the audio waveform OW3 by removing noise so as to achieve "high-quality violin performance".

[0054] For example, the information processing apparatus 100 may generate the audio waveform OW3 using not only such a deep generative model, but another noise / denoise approach (for example, a denoising autoencoder (DAE) or the like). For example, the information processing apparatus 100 may generate the audio waveform OW3 by using a method of applying a noise / denoise approach while maintaining the MIDI data MD1 and the tone information (it may explicitly not be a method for maintaining the MIDI data MD1 or the tone information). For example, the information processing apparatus 100 may generate the audio waveform OW3 using a deep generative model conditioning the MIDI data MD1 or tone information, or the like. In addition, for example, the information processing apparatus 100 may generate the audio waveform OW3 by using a noise / denoise approach or the like having the same deep generative model but a different method.

[0055] Note that, in the method disclosed in NPL 1, since noise is added randomly (at random), noise is not added to all portions without reality, but the information processing apparatus 100 may perform processing in which noise is added to portions without reality as much as possible. The information processing apparatus 100 may perform processing of adding noise while performing adjustment so that noise is added to portions without reality as much as possible.

[0056] Note that, in the audio waveform OW2 of Fig. 1, black lines that do not exist in the audio waveform OW1 are added, and the black lines indicate noise added to the audio waveform OW1 for convenience of description. In addition, the audio waveform OW3 of Fig. 1 is expressed by one piece of clear tone information, which indicates that the audio waveform OW3 is smooth and high-quality audio for convenience of description. In practice, similarly to the audio waveforms OW1 and OW2, the pitch may also change in the middle of the audio waveform OW3.

[0057] The secondary processing of converting into a high-quality audio waveform has been described above. Hereinafter, a variation of the processing of generating the audio waveform OW1, which is the primary processing, will be described. In the embodiment described above, the processing of generating the audio waveform OW1 using MIDI data and One Shot samples has been described as an example, and the other variations will be described. Here, two variations will be described as other variations, but it is not particularly limited to these examples, and the audio waveform OW1 may be generated (acquired) using any method. Note that, in any of the variations, the audio waveform OW1 is input information of the secondary processing.

[0058] (Variation 1 of information processing: processing using MIDI data and tone category) In the embodiment described above, the information processing apparatus 100 may use a tone category to generate the audio waveform OW1. The information processing apparatus 100 may generate the audio waveform OW1 using MIDI data and a tone category. For example, when the user designates MIDI data and a tone category, the information processing apparatus 100 may select One Shot samples on the basis of the designated tone category. For example, the information processing apparatus 100 may randomly select One Shot samples, or in a case where an optimum One Shot sample is registered in advance for each tone category, the information processing apparatus 100 may select the registered One Shot sample.

[0059] In addition, when the user designates a tone category, the information processing apparatus 100 may present a plurality of candidates and select a One Shot sample from the candidates on the basis of a further designated tone category. For example, in a case where the tone category designated by the user is a piano, for example, the information processing apparatus 100 may present a list of candidates such as "grand piano", "upright piano", "digital piano", and the like together with a comment such as "In general, there is a plurality of categories for piano. Which category would you like?", and cause the user to select a candidate. At this time, for example, the information processing apparatus 100 may present a list of candidates after narrowing down the candidates to some extent according to the user information (for example, preference information of the user for music, tone, instrument, and the like, a designation history of tone category, and the like), or may present a list of candidates such that a specific candidate is preferentially displayed at a higher level according to the user information.

[0060] (Variation 2 of information processing: processing using synthesizer or the like) In the embodiment described above, the information processing apparatus 100 may directly acquire an audio waveform synthesized in advance by an electronic technique such as a synthesizer or the like to obtain the audio waveform OW1. For example, the information processing apparatus 100 may acquire the audio waveform designated by the user when the user designates the audio waveform synthesized by the synthesizer or the like. At this time, the user may provide the audio waveform synthesized by the synthesizer or the like to the information processing apparatus 100 when designating the audio waveform. Then, the information processing apparatus 100 may set the audio waveform provided from the user as the audio waveform OW1.

[0061] In this manner, the information processing apparatus 100 may set the audio waveform synthesized by the synthesizer or the like as the audio waveform OW1. In addition, the information processing apparatus 100 may set the audio waveform synthesized not only by a synthesizer but also by another electronic technique or the like as the audio waveform OW1. As described above, the audio waveform OW1 may be an audio waveform generated by a synthesizer or another electronic technique.

[0062] The variations of the processing of generating the audio waveform OW1 have been described above. Hereinafter, various variations according to the embodiment will be described.

[0063] (Variation 3 of information processing: scope of application of information processing) The information processing according to the embodiment described above may be achieved by, for example, a web application, application software, a plug-in of a music production tool, or the like. In addition, the information processing according to the embodiment described above may be achieved by, for example, a server apparatus, a cloud system, the user terminal 10 to be described below, or the like. In addition, the information processing according to the embodiment described above may be achieved by, for example, an apparatus, a method, a program, a system, or the like. As described above, the information processing according to the embodiment described above may be achieved in any form. For example, the information processing according to the embodiment described above is not limited to be achieved by the information processing apparatus 100 and the user terminal 10 being separate apparatuses, and may be achieved by the information processing apparatus 100 and the user terminal 10 being integrated.

[0064] (Variation 4 of information processing: generation processing of One Shot sample) In the embodiment described above, the One Shot sample may include a plurality of types of audio data. For example, the One Shot sample may include, for example, a plurality of types of audio data of main audio (such as main melody) and sub audio (such as sub melody). For example, each of the audio data OD1 to OD4 may include a plurality of types of audio data of main audio and sub audio corresponding to each of the audio data OD1 to OD4.

[0065] In addition, the One Shot sample may include, for example, only main audio, only sub audio, or audio data of a type different from the main audio or the sub audio. For example, the One Shot sample may further include audio data of a type that is sub audio of sub audio.

[0066] The information processing apparatus 100 may combine the plurality of types of audio data to generate a One Shot sample. For example, the information processing apparatus 100 may combine the plurality of types of audio data of main audio and sub audio to generate a One Shot sample.

[0067] The information processing apparatus 100 may perform weighting on each of the main audio and the sub audio and generate a One Shot sample according to the weighting. For example, the sound of "do" of a piano and the sound of "do" of a guitar are different although they are both the sound of "do". The information processing apparatus 100 may generate a One Shot sample of a new sound of "do" by weighting each sound. This enables the information processing apparatus 100 to generate a new One Shot sample having both piano and guitar elements. As a result, an audio waveform of various tones can be synthesized.

[0068] The information processing apparatus 100 may generate a One Shot sample by changing the weighting so that the tone changes in the timeline. For example, the information processing apparatus 100 may generate a One Shot sample by changing the weighting for each instrument so that the tone changes in the timeline.

[0069] In a case where there is a plurality of types of each of main audio and sub audio, the information processing apparatus 100 may generate a One Shot sample by appropriately combining the audio data of these types. For example, the information processing apparatus 100 may generate a One Shot sample by appropriately combining one main audio and two sub audios.

[0070] The information processing apparatus 100 may generate a One Shot sample by combining audio data of different types, the same instrument, or the like, or may generate a One Shot sample by combining audio data of different types, different instruments, or the like. For example, the information processing apparatus 100 may combine two types of piano sounds (for example, the sound of a grand piano and the sound of a Jazz piano are combined), may combine the piano sound and the guitar sound, or may combine the piano sound and the singing sound.

[0071] The information processing apparatus 100 may perform pasting synthesis on the One Shot samples thus generated along the MIDI data MD1 to generate an audio waveform. For example, the information processing apparatus 100 may generate the audio waveform on the basis of the information included in the One Shot sample and the information included in the MIDI data MD1. For example, when the information processing apparatus 100 acquires, from the MIDI data MD1, information indicating that the first note is the sound of "do" of the piano and has 1 second of "0:01 to 0:02" (the time from when the key is pressed to when it is released is 1 second), the corresponding One Shot sample may be specified from the One Shot sample group and pasting synthesis may be performed.

[0072] (Variation 5 of information processing: audio waveform editing processing) In the embodiment described above, the information processing apparatus 100 may perform editing (pre-processing for the secondary processing) on the audio waveform OW1. For example, the information processing apparatus 100 may perform editing such as compression and limiting.

[0073] The information processing apparatus 100 may perform editing such as compressing dynamics such that the volume becomes constant to some extent between the audio at the time of musical performance with one note in the audio waveform OW1 and the audio at the time of musical performance with a plurality of notes. In addition, the information processing apparatus 100 may perform editing such as limiting based on a threshold (setting an upper limit of amplitude) so that the volume becomes constant to some extent.

[0074] (Variation 6 of information processing: variation of MIDI data) In the embodiment described above, the MIDI data to be processed may be, for example, data of the entire one piece of music (a certain set of music), data of a certain portion or section of the one piece of music, data of a certain instrument part of the one piece of music, or the like, and is not particularly limited.

[0075] In addition, the MIDI data to be processed may include information related to a release time. The release time is, for example, a time such as of reverberant sound after the hand is released from a key in the case of a piano. For example, how long a sound attenuates after the offset ends depends on an instrument or the like.

[0076] The information processing apparatus 100 may perform pasting synthesis in consideration of release time. For example, the information processing apparatus 100 may determine an audio release time and apply the audio release time to the One Shot sample to perform pasting synthesis. For example, the information processing apparatus 100 may randomly set a release time in a certain range on the basis of MIDI data, and perform pasting synthesis so that the One Shot sample fades out according to the release time.

[0077] (1-3. Outline 2 of the information processing system according to the embodiment) Next, the relationship of the information processing of the information processing apparatus 100 and the user terminal 10 will be described with reference to Figs. 4A to 4C. Figs. 4A to 4C are sequence diagrams (1) to (3) illustrating an outline of the information processing system 1 according to the embodiment. Figs. 4A to 4C are different in that acquired data is MIDI data / One Shot sample (Fig. 4A), MIDI data / tone category (Fig. 4B), and synthesized audio waveform (Fig. 4C), but are similar on the other points, and thus will be collectively described for convenience of description.

[0078] Before describing Figs. 4A to 4C, first, the user terminal 10 will be described. The user terminal 10 is an information processing apparatus used by a user who desires audio data corresponding to predetermined MIDI data. The user is, for example, a user who performs music production while sticking to a subtle tone of each sound, and is a user who desires high-quality audio data corresponding to predetermined MIDI data.

[0079] By providing predetermined MIDI data, the user obtains audio data corresponding to the MIDI data through the information processing according to the embodiment. At this time, the user designates at tone category to obtain audio data corresponding to the MIDI data played in the tone category.

[0080] The user terminal 10 may be any apparatus as long as the processing according to the embodiment can be achieved. In addition, the user terminal 10 may be an apparatus such as a smartphone, a tablet terminal, a notebook PC, a desktop PC, a mobile phone, or a PDA. Figs. 4A to 4C illustrate a case where the user terminal 10 is a smartphone.

[0081] The user terminal 10 is, for example, a smart device such as a smartphone or a tablet, and is a portable terminal apparatus capable of communicating with an arbitrary server apparatus via a wireless communication network such as 4G to 5G (Generation) or Long Term Evolution (LTE). In addition, the user terminal 10 has a screen such as a liquid crystal display, the screen having a screen having a touch panel function, and may receive various operations of display data such as content, such as a tap operation, a slide operation, and a scroll operation, from the user with a finger, a stylus, or the like. In Figs. 4A to 4C, the user terminal 10 is used by a user U1.

[0082] In Figs. 4A to 4C, the information processing apparatus 100 acquires MIDI data to be processed. For example, the information processing apparatus 100 acquires MIDI data provided by the user U1 via the user terminal 10 as MIDI data to be processed.

[0083] In addition, for example, the information processing apparatus 100 may acquire MIDI data input, designated, or selected by the user U1 via a predetermined application, a web, or the like via the user terminal 10 as MIDI data to be processed.

[0084] In addition, for example, the information processing apparatus 100 may acquire MIDI data created by entering a tune or the like conceived by the user U1 or the like with the keyboard or the like as MIDI data to be processed.

[0085] In Fig. 4A, the information processing apparatus 100 acquires a One Shot sample together with MIDI data to be processed (Step S1). For example, the information processing apparatus 100 may acquire a One Shot sample prepared in advance for a specific service, or may acquire a One Shot sample provided by the user U1 via the user terminal 10.

[0086] In Fig. 4B, the information processing apparatus 100 acquires a tone category together with the MIDI data to be processed (Step S2). For example, the information processing apparatus 100 may acquire the tone category designated by the user U1 via the user terminal 10.

[0087] In the case of Step S2, the information processing apparatus 100 selects a One Shot sample on the basis of the tone category (Step S3). For example, the information processing apparatus 100 may select a One Shot sample registered as an optimum One Shot sample for the tone category.

[0088] In the case of Steps S1 and S3, the information processing apparatus 100 generates an audio waveform obtained by performing pasting synthesis on the acquired or selected One Shot sample along the MIDI data to be processed (Step S4).

[0089] Then, the information processing apparatus 100 presents the generated audio waveform to the user U1 (Step S5), and receives editing from the user U1 (Step S6).

[0090] In Fig. 4C, the information processing apparatus 100 acquires an audio waveform synthesized in advance by a synthesizer or the like (Step S7). For example, the information processing apparatus 100 may acquire a synthesized audio waveform by a synthesizer or the like provided by the user U1 via the user terminal 10.

[0091] In the case of Steps S6 and S7, the information processing apparatus 100 converts the audio waveform into a high-quality audio waveform by inputting the audio waveform to a deep generative model prepared in advance and performing noise / denoise (Step S8). Then, the information processing apparatus 100 provides the converted audio data waveform to the user U1 (Step S9).

[0092] By using the information processing according to the embodiment described above, it is possible to enable automatic generation of high-quality audio even for a domain in which a pair of audio and MIDI data does not exist or is poor. That is, it is possible to achieve annotation-free MIDI-to-Audio by using the information processing according to the embodiment described above.

[0093] In the embodiment described above, by using the One Shot sample, it is possible to generate an audio waveform having high editing property and a wide application range. By using the One Shot sample, the editing property of synthesis is enhanced, and thus, it can be applied to various audio generation.

[0094] In addition, music production that matches music in the era of desktop music (DTM) becomes possible by using the One Shot sample. Since the DTM is also one form of music that is often listened to on a daily basis, it is possible to produce familiar music by matching the DTM.

[0095] (1-4. Configuration of the user terminal according to the embodiment) Next, a configuration of the user terminal 10 according to the embodiment will be described with reference to Fig. 5. Fig. 5 is a diagram illustrating a configuration example of the user terminal 10 according to the embodiment. As illustrated in Fig. 5, the user terminal includes a communication unit 11, an input unit 12, an output unit 13, and a control unit 14.

[0096] The communication unit 11 is achieved by, for example, a network interface card (NIC), a network interface controller, or the like. The communication unit 11 is connected to a network N by wire or wirelessly, and transmits and receives information to and from the information processing apparatus 100 via the network N. The network N is achieved by, for example, a wireless communication standard or system such as Bluetooth (registered trademark), the Internet, Wi-Fi (registered trademark), Ultra Wide Band (UWB), or Low Power Wide Area (LPWA).

[0097] The input unit 12 receives various operations from the user. In Figs. 4A to 4C, various operations from the user U1 are received. For example, the input unit 12 may receive various operations from the user via the display surface by a touch panel function. In addition, the input unit 12 may receive various operations from a button provided on the user terminal 10 or a keyboard or a mouse connected to the user terminal 10.

[0098] The output unit 13 is a display screen such as a tablet terminal achieved by, for example, a liquid crystal display, an organic electro-luminescence (EL) display, or the like, and is a display apparatus for displaying various types of information. For example, the output unit 13 displays information transmitted from the information processing apparatus 100.

[0099] The control unit 14 is achieved by, for example, a central processing unit (CPU), a micro processing unit (MPU), a graphics processing unit (GPU), or the like executing a program stored inside the user terminal 10 using random access memory (RAM) or the like as a work area. In addition, the control unit 14 is a controller and may be achieved by, for example, an integrated circuit such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a micro controller unit (MCU).

[0100] As illustrated in Fig. 5, the control unit 14 includes a reception unit 141 and a transmission unit 142, and achieves or executes an action of information processing described below.

[0101] The reception unit 141 receives various types of information from another information processing apparatus such as the information processing apparatus 100. For example, the reception unit 141 receives audio data generated corresponding to MIDI data input, designated, or selected by the user or the like. For example, the reception unit 141 receives audio data generated corresponding to MIDI data created by entering a tune or the like conceived by the user or the like with the keyboard or the like.

[0102] The transmission unit 142 transmits various types of information to another information processing apparatus such as the information processing apparatus 100. For example, the transmission unit 142 transmits MIDI data input, designated, or selected by the user or the like. For example, the transmission unit 142 transmits MIDI data created by entering a tune or the like conceived by the user or the like with the keyboard or the like.

[0103] In addition, for example, the transmission unit 142 transmits One Shot sample information input, designated, or selected by the user or the like. For example, the transmission unit 142 transmits the One Shot sample information input, designated, or selected by the user or the like together with the MIDI data input, designated, or selected by the user or the like.

[0104] In addition, for example, the transmission unit 142 transmits tone category information input, designated, or selected by the user or the like. For example, the transmission unit 142 transmits the tone category information input, designated, or selected by the user or the like together with the MIDI data input, designated, or selected by the user or the like.

[0105] In addition, for example, the transmission unit 142 transmits synthesized audio data input, designated, or selected by the user or the like. For example, the transmission unit 142 transmits audio data synthesized in advance by a synthesizer or another electronic technique.

[0106] (1-5. Configuration of the information processing apparatus according to the embodiment) Next, a configuration of the information processing apparatus 100 according to the embodiment will be described with reference to Fig. 6. Fig. 6 is a diagram illustrating a configuration example of the information processing apparatus 100 according to the embodiment. As illustrated in Fig. 6, the information processing apparatus 100 includes a communication unit 110, a storage unit 120, and a control unit 130. Note that the information processing apparatus 100 may include an input unit (for example, a keyboard, a mouse, or the like) that receives various operations from a manager of the information processing apparatus 100, and a display unit (for example, a liquid crystal display or the like) for displaying various types of information.

[0107] The communication unit 110 is achieved by, for example, an NIC, a network interface controller, or the like. The communication unit 110 is connected to the network N by wire or wirelessly, and transmits and receives information to and from the user terminal 10 or the like via the network N. The network N is achieved by, for example, a wireless communication standard or system such as Bluetooth, the Internet, Wi-Fi, UWB, or LPWA.

[0108] The storage unit 120 is achieved by, for example, a semiconductor memory element such as RAM or flash memory, or a storage apparatus such as a hard disk, solid state drive (SSD), or an optical disk. As illustrated in Fig. 6, the storage unit 120 includes a One Shot sample storage unit 121 and a model storage unit 122.

[0109] The One Shot sample storage unit 121 stores information regarding One Shot samples (corresponding to the audio data OD1 to OD4 in Fig. 1). Here, Fig. 7 illustrates an example of the One Shot sample storage unit 121 according to the embodiment. As illustrated in Fig. 7, the One Shot sample storage unit 121 includes items such as "sample ID", "tone", "pitch", and "One Shot sample".

[0110] "Sample ID" indicates identification information for identifying the One Shot sample. "Tone" indicates a tone. For example, in "tone", information regarding the tone of an instrument or the like to be played is stored. For example, information such as a piano or a violin is stored in "tone". "Pitch" indicates pitch. For example, in "pitch", information regarding the pitch to be played is stored. For example, information such as G3# and B3 is stored in "pitch". "One Shot sample" indicates audio data of a One Shot sample. In the example illustrated in Fig. 7, an example in which conceptual information such as "One Shot sample #1" and "One Shot sample #2" is stored in "One Shot sample" has been described, but in practice, an audio waveform (for example, data based on amplitude and time) or the like is stored.

[0111] The model storage unit 122 stores information regarding the deep generative model applied in the secondary processing. Here, Fig. 8 illustrates an example of the model storage unit 122 according to the embodiment. As illustrated in Fig. 8, the model storage unit 122 includes items such as "model ID" and "model".

[0112] "Model ID" indicates identification information for identifying a deep generative model. "Model" indicates a deep generative model. In the example illustrated in Fig. 8, an example in which conceptual information such as "model #1" and "model #2" is stored in "model" has been described, but in practice, parameters of the deep generative model, and the like are stored. In addition, "model" may store information regarding a noise level to be applied, and the like.

[0113] The control unit 130 is achieved by, for example, a CPU, an MPU, a GPU, the like executing a program (for example, the information processing program according to the present disclosure) stored in the information processing apparatus 100 using RAM or the like as a work area. In addition, the control unit 130 is a controller and may be achieved by, for example, an integrated circuit such as an ASIC, an FPGA, or an MCU.

[0114] As illustrated in Fig. 6, the control unit 130 includes an acquisition unit 131, a first generation unit 132, an editing unit 133, a second generation unit 134, and a provision unit 135, and achieves or executes an action of information processing described below. Note that the internal configuration of the control unit 130 is not limited to the configuration illustrated in Fig. 6, and may be another configuration as long as information processing to be described below is performed.

[0115] The acquisition unit 131 acquires various types of information. For example, the acquisition unit 131 acquires information transmitted from the user terminal 10, an external information processing apparatus such as a server apparatus or a cloud system that provides or manages a specific service, or the like.

[0116] In the primary processing, the acquisition unit 131 acquires, for example, MIDI data (corresponding to MIDI data MD1 in Fig. 1). For example, the acquisition unit 131 acquires MIDI data input, designated, or selected by the user or the like. For example, the acquisition unit 131 acquires MIDI data created by entering a tune or the like conceived by the user or the like with the keyboard or the like.

[0117] In the primary processing, the acquisition unit 131 acquires, for example, One Shot sample information (corresponding to the audio data OD1 to OD4 in Fig. 1). For example, the acquisition unit 131 acquires One Shot sample information input, designated, or selected by the user or the like. For example, the acquisition unit 131 acquires the One Shot sample information input, designated, or selected by the user or the like together with the MIDI data input, designated, or selected by the user or the like.

[0118] In the primary processing, the acquisition unit 131 acquires, for example, tone category information. For example, the acquisition unit 131 acquires tone category information input, designated, or selected by the user or the like. For example, the acquisition unit 131 acquires the tone category information input, designated, or selected by the user or the like together with the MIDI data input, designated, or selected by the user or the like.

[0119] In the primary processing, the acquisition unit 131 acquires, for example, synthesized audio data. For example, the acquisition unit 131 acquires synthesized audio data input, designated, or selected by the user or the like. For example, the acquisition unit 131 acquires audio data synthesized in advance by a synthesizer or another electronic technique.

[0120] In the secondary processing, the acquisition unit 131 acquires, for example, audio data to be processed (hereinafter, "first audio data" as appropriate). The first audio data is, for example, audio data generated on the basis of MIDI data. For example, the first audio data is audio data generated by performing pasting synthesis on One Shot samples along MIDI data.

[0121] The One Shot sample is, for example, audio data prepared in advance as a One Shot sample used for pasting synthesis. For example, the One Shot sample is audio data prepared in advance as a One Shot sample used for pasting synthesis regardless of MIDI data. For example, it is audio data prepared in advance for a specific service or the like as freely acquirable audio data.

[0122] In addition, the One Shot sample is, for example, audio data selected from audio data prepared in advance according to a tone category designated by the user. For example, the One Shot sample is audio data selected from audio data prepared in advance according to a tone category selected by the user from among tone category candidates presented to the user according to a tone category designated by the user.

[0123] For example, in a case where the user designates a tone category in a rough unit such as piano, when tone category candidates such as "grand piano", "upright piano", and "digital piano" are presented to the user together with a comment such as "In general, there is a plurality of categories for piano. Which category would you like?", and the user selects the tone category "digital piano" from the candidates, the corresponding One Shot sample is selected from the One Shot sample group corresponding to the tone category "digital piano".

[0124] In addition, the first audio data is, for example, edited audio data edited via a screen presented to the user to enable editing of the first audio data. For example, the first audio data is audio data newly generated by the user editing at least one data of the MIDI data and the One Shot sample and performing pasting synthesis using the edited data again.

[0125] The first generation unit 132 generates the first audio data to be audio data to be processed in the secondary processing on the basis of the MIDI data acquired by the acquisition unit 131, for example. For example, the first generation unit 132 generates the first audio data on the basis of the MIDI data and the One Shot sample. In addition, for example, the first generation unit 132 generates the first audio data on the basis of the MIDI data and the tone category. For example, the first generation unit 132 generates the first audio data on the basis of the MIDI data and the One Shot sample selected according to the tone category designated by the user.

[0126] In addition, the first generation unit 132 generates the first audio data on the basis of, for example, edited data (edited MIDI data, One Shot sample, or the like) edited by the editing unit 133 to be described below.

[0127] The editing unit 133 performs, for example, processing for enabling editing of the first audio data generated by the first generation unit 132. For example, the editing unit 133 displays a screen for enabling reproduction of the generated first audio data. In addition, for example, the editing unit 133 displays a screen for enabling editing of the MIDI data and the One Shot sample in order to edit the first audio data. At this time, for example, the screen for enabling reproduction of the first audio data may include an operation button for enabling editing of the MIDI data or the One Shot sample. The screen may be transitioned to a screen on which the MIDI data and the One Shot sample can be edited by the user operating (for example, a click, a tap, or the like) the operation button.

[0128] Fig. 9 is a diagram (1) illustrating an example of a UI screen according to the embodiment. Here, an example of the UI screen for performing editing according to the embodiment will be described. A screen G1 is a screen displayed on the user terminal 10. The screen G1 includes, for example, an operation button B1 for reproducing the first audio data generated by the first generation unit 132. The first audio data is reproduced when the operation button B1 is operated. That is, the user can confirm what kind of music the generated first audio data actually is. Note that, in a case where there is image (including still image and moving image) data corresponding to the first audio data, the image data may also be reproduced. For example, it is a case where the user wants to create music that matches image data.

[0129] In addition, the screen G1 includes, for example, an operation button B2 for editing the MIDI data used for synthesizing the first audio data. By operating the operation button B2, a screen on which the MIDI data can be edited is displayed. For example, while listening to what kind of music the generated first audio data actually is, the user performs editing such as deleting a note considered to be unnecessary and adding a note considered to be necessary.

[0130] In addition, the screen G1 includes, for example, an operation button B3 for editing the One Shot sample used for synthesizing the first audio data. By operating the operation button B3, a screen on which the One Shot sample can be edited is displayed. For example, while listening to what kind of music the generated first audio data actually is, the user performs editing such as replacing some of the One Shot samples with One Shot samples of a different tone, or changing dynamics of some of the One Shot samples.

[0131] Note that, although not illustrated in Fig. 9, the screen G1 may include an operation button for enabling change in the tone category. The user may operate the operation button to transition the screen to a screen on which the tone category can be changed.

[0132] In addition, the screen G1 is an example of the UI screen according to the embodiment, and the UI screen for performing editing according to the embodiment is not particularly limited to this example.

[0133] The second generation unit 134 generates audio data (hereinafter, "second audio data" as appropriate) having improved quality as compared with the first audio data on the basis of, for example, the first audio data generated by the first generation unit 132 (or the edited first audio data edited by the editing unit 133). For example, the second generation unit 134 generates the second audio data (for example, the second audio data having improved quality such that the audio is similar to real musical performance in real space from mechanical musical performance) imitating real musical performance in the real space as compared with the first audio data. That is, the second generation unit 134 converts the first audio data into the second audio data.

[0134] For example, the second generation unit 134 generates the second audio data by inputting the first audio data to a deep generative model prepared in advance. For example, the second generation unit 134 generates the second audio data by using a trained deep generative model that has been trained in advance so as to improve the quality of the audio data by adding noise so that some information is missing and performing denoise to compensate for missing some information.

[0135] For example, the second generation unit 134 generates the second audio data by adding noise at a noise level randomly selected to the first audio data so that some information is missing and performing denoise so as to compensate for missing some information.

[0136] For example, the second generation unit 134 generates the second audio data by adding noise at a noise level selected on the basis of the operation by the user to the first audio data so that some information is missing and performing denoise so as to compensate for missing some information. For example, the second generation unit 134 generates the second audio data by adding noise at a noise level selected via the screen presented to the user to enable editing of the first audio data so that some information is missing and performing denoise so as to compensate for missing some information.

[0137] Fig. 10 is a diagram (2) illustrating an example of a UI screen according to the embodiment. Here, an example of the UI screen for adjusting a noise level according to the embodiment will be described. A screen G2 is a screen displayed on the user terminal 10. The screen G2 includes, for example, an operation button B11 for adjusting the noise level. The operation button B11 is a slider, and is operated by the user sliding the operation button B11 left and right. By operating the operation button B11, the noise level to be applied is determined. That is, the user can freely determine parameters such as what level of noise is added, how much change is desired, or whether change is allowed. In Fig. 10, the user can freely determine the noise level from 0 to 100 by operating the operation button B11. In Fig. 10, the user adjusts the noise level to be set to "40". In addition, when an operation button B12 included in the screen G2 is operated, processing of generating the second audio data at the noise level adjusted by the user with the operation button B11 is executed.

[0138] The provision unit 135 provides, for example, the second audio data generated by the second generation unit 134. For example, the provision unit 135 provides information regarding the second audio data to the user who has input, designated, or selected the MIDI data. In addition, for example, the provision unit 135 provides information regarding the second audio data to the user who has input, designated, or selected the One Shot sample. In addition, for example, the provision unit 135 provides information regarding the second audio data to the user who has input, designated, or selected the tone category. In addition, for example, the provision unit 135 provides information regarding the second audio data to the user who has input, designated, or selected the synthesized audio data.

[0139] For example, the provision unit 135 may provide the second audio data in any form as long as the user can acquire the second audio data. For example, the provision unit 135 may directly transmit the second audio data to the user terminal 10, or may transmit the second audio data to a server apparatus, a cloud system, or the like for a specific service that can be acquired by the user.

[0140] Next, processing of each unit constituting the information processing apparatus 100 described above will be described in detail along the flows with reference to Figs. 11 to 13. Fig. 11 is a flowchart (1) illustrating a flow of generation processing (primary processing) in the control unit 130.

[0141] When acquiring the MIDI data from the user, the information processing apparatus 100 determines whether designation of the tone category has been received from the user (Step S11).

[0142] In a case where the information processing apparatus 100 determines that the designation of the tone category has been received from the user (Step S11; YES), the One Shot sample corresponding to the tone category is selected (Step S12).

[0143] On the other hand, in a case where the information processing apparatus 100 determines that the designation of the tone category has not been received from the user (Step S11; NO), the One Shot sample prepared in advance regardless of the tone category is acquired (Step S13).

[0144] The information processing apparatus 100 generates an audio waveform obtained by performing pasting synthesis on the One Shot samples along the MIDI data (Step S14).

[0145] Fig. 12 is a variation of the generation processing in the control unit 130. Fig. 12 is a flowchart (2) illustrating a flow of generation processing (primary processing) in the control unit 130.

[0146] When acquiring the MIDI data from the user, the information processing apparatus 100 determines whether designation of the One Shot sample has been received from the user (Step S21).

[0147] In a case where the information processing apparatus 100 determines that the designation of the One Shot sample has been received from the user (Step S21; YES), the One Shot sample received from the user is acquired (Step S22).

[0148] On the other hand, in a case where the information processing apparatus 100 determines that the designation of the One Shot sample has not been received from the user (Step S21; NO), it is determined whether designation of the tone category has been received from the user (Step S23).

[0149] In a case where the information processing apparatus 100 determines that the designation of the tone category has been received from the user (Step S23; YES), the One Shot sample corresponding to the tone category is selected (Step S24).

[0150] On the other hand, in a case where the information processing apparatus 100 determines that the designation of the tone category has not been received from the user (Step S23; NO), the One Shot sample prepared in advance regardless of the tone category is acquired (Step S25).

[0151] The information processing apparatus 100 generates an audio waveform obtained by performing pasting synthesis on the One Shot samples along the MIDI data (Step S26).

[0152] Subsequently, the information processing apparatus 100 performs generation processing using the deep generative model or the like. Such processing will be described with reference to Fig. 13. Fig. 13 is a flowchart illustrating a flow of generation processing (secondary processing) in the control unit 130.

[0153] When acquiring the first audio data to be processed, the information processing apparatus 100 determines whether designation of the noise level to be applied to the acquired first audio data has been received from the user (Step S31).

[0154] In a case where the information processing apparatus 100 determines that the designation of the noise level has been received from the user (Step S31; YES), it is determined that the noise level received from the user is applied (Step S32).

[0155] On the other hand, in a case where the information processing apparatus 100 determines that the designation of the noise level has not been received from the user (Step S31; NO), it is determined that the noise level is randomly selected and applied (Step S33).

[0156] The information processing apparatus 100 generates the second audio data by adding noise of a predetermined noise level to the acquired first audio data and further performing denoise (Step S34).

[0157] Then, the information processing apparatus 100 provides the generated second audio data (Step S35).

[0158] (2. Other embodiments) The processing according to each embodiment described above may be performed in various different forms other than each embodiment described above.

[0159] Among the pieces of processing described in the embodiment of the present disclosure described above, all or some of the pieces of processing described as being performed automatically can be performed manually, or all or some of the pieces of processing described as being performed manually can be performed automatically by a known method. Additionally, the processing procedures, the specific names, and the information including various data and parameters indicated in the aforementioned document and drawings can be arbitrarily changed unless otherwise specified. For example, the various types of information illustrated in each drawing are not limited to the illustrated information.

[0160] In addition, each component of each apparatus illustrated in the drawings is functionally conceptual, and is not necessarily physically configured as illustrated in the drawings. That is, a specific form of distribution and integration of apparatuses is not limited to those illustrated in the drawings, and all or a part thereof can be functionally or physically distributed and integrated in an arbitrary unit according to various loads, usage situations, and the like.

[0161] In addition, the embodiment of the present disclosure described above can be appropriately combined within an area not contradicting processing contents. In addition, the order of each step illustrated in the sequence diagram or the flowchart of the present embodiment can be changed as appropriate. For example, each step may be processed in time series, processed iteratively, or processed partially in parallel.

[0162] In addition, the effects described in the present specification are merely examples and are not limitative, and there may be other effects.

[0163] (3. Effects of the information processing apparatus according to present disclosure) As described above, the information processing apparatus (the information processing apparatus 100 in the embodiment) according to the present disclosure includes the acquisition unit (the acquisition unit 131 in the embodiment) and the generation unit (the second generation unit 134 in the embodiment). The acquisition unit acquires the first audio data to be processed. The generation unit generates the second audio data of higher quality than the first audio data by inputting the first audio data to the deep generative model prepared in advance.

[0164] As described above, the information processing apparatus according to the present disclosure can enable automatic generation of high-quality audio for various types of highly complicated music including general music. For example, the information processing apparatus can enable automatic generation of high-quality audio even for music in which training data of MIDI data for audio is poor or unavailable.

[0165] In addition, the generation unit generates the second audio data imitating the real musical performance in the real space.

[0166] As described above, the information processing apparatus can enable automatic generation of high-quality audio imitating real musical performance in the real space.

[0167] In addition, the generation unit generates the second audio data by using a deep generative model that has been trained so as to improve the quality of the audio data by adding noise so that some information is missing and performing denoise to compensate for missing some information.

[0168] As described above, the information processing apparatus can add noise such that a portion having a noticeable sense of pasting without reality or a joint portion where a sound changes unnaturally disappears, and thus, it is possible to generate high-quality audio close to real musical performance with high reality.

[0169] In addition, the generation unit generates the second audio data by using the deep generative model by adding noise at a noise level randomly selected to the first audio data so that some information is missing and performing denoise so as to compensate for missing some information.

[0170] As described above, since the information processing apparatus can simplify the operation of the user, it is possible to simplify automation of high-quality audio generation close to real musical performance with high reality.

[0171] In addition, the generation unit generates the second audio data by using the deep generative model by adding noise at a noise level selected on the basis of the operation by the user to the first audio data so that some information is missing and performing denoise so as to compensate for missing some information.

[0172] As described above, the information processing apparatus enables the user to freely select the noise level, thereby enabling high-quality audio generation close to real musical performance with high reality in consideration of the intention and desire of the user.

[0173] In addition, the generation unit generates the second audio data via the noise of the noise level selected via the screen presented to the user to enable editing of the first audio data at the time of generating the first audio data.

[0174] As described above, the information processing apparatus enables the user to actually listen to the audio and select the noise level on the UI screen, so that it is possible to provide a music production system in which AI and a person easily cooperate and which is excellent in usability.

[0175] In addition, the acquisition unit acquires the first audio data generated based on the MIDI data.

[0176] As described above, the information processing apparatus can enable automatic generation of high-quality audio along the MIDI data.

[0177] In addition, the acquisition unit acquires the first audio data generated by performing pasting synthesis on the One Shot samples along the MIDI data.

[0178] As described above, the information processing apparatus can enable automatic generation of high-quality audio along the MIDI data / One Shot sample.

[0179] In addition, the acquisition unit acquires the first audio data using the One Shot samples prepared in advance as One Shot samples used for pasting synthesis.

[0180] As described above, the information processing apparatus can enable automatic generation of high-quality audio along the MIDI data / One Shot sample prepared in advance for a specific service.

[0181] In addition, the acquisition unit acquires the first audio data using the One Shot sample selected according to the tone category designated by the user.

[0182] As described above, the information processing apparatus can enable automatic generation of high-quality audio along the MIDI data / tone category.

[0183] In addition, the acquisition unit acquires the first audio data using the One Shot sample selected according to the tone category selected by the user from among the tone category candidates presented to the user according to the tone category.

[0184] As described above, the information processing apparatus can enable automatic generation of high-quality audio along the MIDI data / One Shot sample selected according to tone category.

[0185] In addition, the acquisition unit acquires, as the first audio data, edited audio data edited via a screen presented to the user to enable editing of the first audio data at the time of generating the first audio data.

[0186] As described above, the information processing apparatus enables the user to actually listen to the audio and perform editing on the UI screen, so that it is possible to provide a music production system in which AI and a person easily cooperate and which is excellent in usability.

[0187] In addition, the acquisition unit acquires, as the first audio data, the audio data based on the edited MIDI data edited by the user via the screen.

[0188] In this manner, the information processing apparatus enables the user to actually listen to the audio and edit the MIDI data on the UI screen, thereby enabling automatic generation of high-quality audio along the edited MIDI data.

[0189] In addition, the acquisition unit acquires audio data synthesized in advance by a predetermined electronic technique as the first audio data.

[0190] As described above, the information processing apparatus can simplify automation of high-quality audio generation by using audio data synthesized in advance by an electronic technique such as a synthesizer.

[0191] (4. Hardware configuration) The information processing apparatus 100 or the like according to the embodiment of the present disclosure described above is achieved by a computer 1000 having the configuration as illustrated, for example, in Fig. 14. The information processing apparatus 100 will be described as an example. Fig. 14 is a hardware configuration diagram illustrating an example of the computer 1000 that achieves the function of the information processing apparatus 100. The computer 1000 includes processing circuitry 1100, RAM 1200, ROM 1300, a secondary storage apparatus 1400, a communication interface 1500, an input / output interface 1600, a display unit 1700, a camera unit 1800, a microphone 1900, and a speaker 2000. Each unit of the computer 1000 is connected by a bus 1050.

[0192] The processing circuitry 1100 operates on the basis of a program stored in the ROM 1300 or the secondary storage apparatus 1400, and controls each unit. For example, the processing circuitry 1100 loads the program stored in the ROM 1300 or the secondary storage apparatus 1400 to the RAM 1200, and executes processing corresponding to various programs.

[0193] The ROM 1300 stores a boot program such as a basic input output system (BIOS) executed by the processing circuitry 1100 when the computer 1000 is activated, a program depending on hardware of the computer 1000, and the like.

[0194] The secondary storage apparatus 1400 is a computer-readable recording medium that non-transiently records a program executed by the processing circuitry 1100, data used by the program, and the like. Specifically, the secondary storage apparatus 1400 is a recording medium that records a program of each processing of the information processing apparatus 100 according to the embodiment of the present disclosure, which is an example of program data 1450.

[0195] The communication interface 1500 is an interface for the computer 1000 to connect to an external network 1550. The communication interface 1500 corresponds to the communication unit 110 included in the information processing apparatus 100. For example, the processing circuitry 1100 receives data from another device or transmits data generated by the processing circuitry 1100 to another device via the communication interface 1500.

[0196] The input / output interface 1600 is an interface for connecting an input / output device 1650 and the computer 1000. For example, the processing circuitry 1100 receives data from the input device such as the microphone 1900 or a touch panel via the input / output interface 1600. In addition, the processing circuitry 1100 transmits data to the output device such as the display unit 1700 or the speaker 2000 via the input / output interface 1600. In addition, the input / output interface 1600 may function as a media interface that reads a program or the like recorded in a predetermined recording medium (media). The media is, for example, an optical recording medium such as a digital versatile disc (DVD) or a phase change rewritable disk (PD), a magneto-optical recording medium such as a magneto-optical disk (MO), a tape medium, a magnetic recording medium, a semiconductor memory, or the like.

[0197] The display unit 1700 is an interface for displaying information processed by the computer 1000. The display unit 1700 is, for example, a liquid crystal display or an organic electro-luminescence (EL) display. In addition, the display unit 1700 may be a touch panel type display apparatus or a video projection apparatus.

[0198] The camera unit 1800 is an interface for the computer 1000 to capture an image. The microphone 1900 is an interface for the computer 1000 to capture a voice. The speaker 2000 is an interface for outputting a voice processed by the computer 1000. Each unit of the computer 1000 is connected by the bus 1050. Each interface is not necessarily provided inside the computer 1000, and may be provided outside the computer 1000 through a network or the like. In addition, each unit constituting the computer 1000 may be controlled by a circuit different from the processing circuitry 1100. For example, the display unit 1700 may be controlled not by the processing circuitry 1100 but by a circuit dedicated to display processing included in the display unit 1700.

[0199] For example, in a case where the computer 1000 functions as the information processing apparatus 100 according to the embodiment of the present disclosure, the processing circuitry 1100 of the computer 1000 executes a program loaded on the RAM 1200 to function as the control unit 130. In addition, the secondary storage apparatus 1400 stores the information processing program according to the present disclosure and various data stored in the storage unit 120. Note that the processing circuitry 1100 reads the program data 1450 from the secondary storage apparatus 1400 and executes the program data, but as another example, these programs may be acquired from another apparatus via the external network 1550. That is, the secondary storage apparatus 1400 is not limited to be placed inside the computer 1000, but may be placed outside the computer 1000. Note that the processing circuitry 1100 is an example of an integrated circuit, and any of the CPU, the MPU, the GPU, the APU, the ASIC, and the FPGA can be regarded as an integrated circuit.

[0200] Note that the present technique can also have the following configurations. (1)  An information processing apparatus, comprising:  circuitry configured to:  receive first audio data;  generate intermediate audio data by adding noise of a predetermined noise level to the first audio data; and  generate second audio data by denoising the intermediate audio data, wherein the second audio data is higher quality audio data than the first audio data. (2)  The information processing apparatus according to (1), wherein the circuitry is configured to receive the predetermined noise level from a user input. (3)  The information processing apparatus according to (1), wherein the circuitry is configured to randomly select the predetermined noise level. (4)  The information processing apparatus according to (1), wherein, to receive the first audio data, the circuitry is configured to:  receive Musical Instruments Digital Interface (MIDI) data;  generate synthesized audio data from the MIDI data and at least one One Shot Sample;  send the synthesized audio data to a user; and  receive the first audio data from the user. (5)  The information processing apparatus according to (4), wherein the circuitry is configured to generate the first audio data by performing pasting synthesis on the at least one One Shot sample based on the MIDI data. (6)  The information processing apparatus according to (4), wherein the circuitry is configured to receive the at least one One Shot sample. (7)  The information processing apparatus according to (4), wherein the circuitry is configured to receive a tone category and select the at least one One Shot sample based on the tone category. (8)  The information processing apparatus according to (4), wherein the circuitry is configured to acquire a predetermined One Shot sample. (9)  The information processing apparatus according to (4), wherein the user edits the synthesized audio data to generate the first audio data sent to the circuitry. (10)  The information processing apparatus according to (9), wherein the user edits at least one of the MIDI data and the at least one One Shot sample. (11)  The information processing apparatus according to (1), wherein the circuitry is further configured to:  receive a condition from a user; and  generate the second audio data further based on the condition. (12)  The information processing apparatus according to (1), wherein the circuitry is configured to denoise the intermediate audio data using a deep generative model. (13)  A method for generating audio data, comprising:  receiving first audio data;  generating intermediate audio data by adding noise of a predetermined noise level to the first audio data; and  generating second audio data by denoising the intermediate audio data, wherein the second audio data is higher quality audio data than the first audio data. (14)  The method according to (13), further comprising receiving the predetermined noise level from a user input. (15)  The method according to (13), further comprising randomly selecting the predetermined noise level. (16)  The method according to (13), wherein, receiving the first audio data includes:  receiving Musical Instruments Digital Interface (MIDI) data;  generating synthesized audio data from the MIDI data and at least one One Shot Sample;  sending the synthesized audio data to a user; and  receive the first audio data from the user. (17)  A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause an information processing apparatus to perform a method comprising:  receiving first audio data;  generating intermediate audio data by adding noise of a predetermined noise level to the first audio data; and generating second audio data by denoising the intermediate audio data, wherein the second audio data is higher quality audio data than the first audio data. (18)  The non-transitory computer-readable medium according to (17), wherein the method further includes receiving the predetermined noise level from a user input. (19)  The non-transitory computer-readable medium according to (17), wherein the method further includes randomly selecting the predetermined noise level. (20)  The non-transitory computer-readable medium according to (17), wherein, receiving the first audio data includes:  receiving Musical Instruments Digital Interface (MIDI) data;  generating synthesized audio data from the MIDI data and at least one One Shot Sample;  sending the synthesized audio to a user; and  receiving the first audio data from the user. (21)  An information processing apparatus including:  an acquisition unit that acquires first audio data to be processed; and  a generation unit that generates second audio data of higher quality than the first audio data by inputting the first audio data to a deep generative model prepared in advance. (22)  The information processing apparatus according to (21), wherein  the generation unit  generates the second audio data imitating real musical performance in a real space. (23)  The information processing apparatus according to (1) or (22), wherein  the generation unit  generates the second audio data by using the deep generative model that has been trained so as to improve quality of audio data by adding noise so that some information is missing and performing denoise to compensate for missing some information. (24)  The information processing apparatus according to any one of (21) to (23), wherein  the generation unit  generates the second audio data by using the deep generative model by adding noise at a noise level randomly selected to the first audio data so that some information is missing and performing denoise so as to compensate for missing some information. (25)  The information processing apparatus according to any one of (21) to (24), wherein  the generation unit  generates the second audio data by using the deep generative model by adding noise at a noise level selected based on an operation by a user to the first audio data so that some information is missing and performing denoise so as to compensate for missing some information. (26)  The information processing apparatus according to (25), wherein  the generation unit  generates the second audio data via noise of the noise level selected via a screen presented to the user to enable editing of the first audio data at a time of generating the first audio data. (27)  The information processing apparatus according to any one of (21) to (26), wherein  the acquisition unit  acquires the first audio data generated based on MIDI data. (28)  The information processing apparatus according to (27), wherein  the acquisition unit  acquires the first audio data generated by performing pasting synthesis on One Shot samples along the MIDI data. (29)  The information processing apparatus according to (28), wherein  the acquisition unit  acquires the first audio data using the One Shot samples prepared in advance as One Shot samples used for the pasting synthesis. (30)  The information processing apparatus according to (28), wherein  the acquisition unit  acquires the first audio data using the One Shot sample selected according to a tone category designated by a user. (31)  The information processing apparatus according to (30), wherein  the acquisition unit  acquires the first audio data using the One Shot sample selected according to the tone category selected by the user from among tone category candidates presented to the user according to the tone category. (32)  The information processing apparatus according to any one of (21) to (31), wherein  the acquisition unit  acquires, as the first audio data, edited audio data edited via a screen presented to a user to enable editing of the first audio data at a time of generating the first audio data. (33)  The information processing apparatus according to (32), wherein  the acquisition unit  acquires, as the first audio data, audio data based on edited MIDI data edited by the user via the screen. (34)  The information processing apparatus according to any one of (21) to (25), wherein  the acquisition unit  acquires audio data synthesized in advance by a predetermined electronic technique as the first audio data. (35)  An information processing method including, by an information processing apparatus:  an acquisition process of acquiring first audio data to be processed; and  a generation process of generating second audio data of higher quality than the first audio data by inputting the first audio data to a deep generative model prepared in advance. (36)  An information processing program for causing a computer to function as:  an acquisition procedure of acquiring first audio data to be processed; and  a generation procedure of generating second audio data of higher quality than the first audio data by inputting the first audio data to a deep generative model prepared in advance.

[0201] 1 Information processing system 10 User terminal 11 Communication unit 12 Input unit 13 Output unit 14 Control unit 100 Information processing apparatus 110 Communication unit 120 Storage unit 121 One Shot sample storage unit 122 Model storage unit 130 Control unit 131 Acquisition unit 132 First generation unit 133 Editing unit 134 Second generation unit 135 Provision unit 141 Reception unit 142 Transmission unit N Network

Claims

1. An information processing apparatus, comprising: circuitry configured to:  receive first audio data;  generate intermediate audio data by adding noise of a predetermined noise level to the first audio data; and  generate second audio data by denoising the intermediate audio data, wherein the second audio data is higher quality audio data than the first audio data.

2. The information processing apparatus according to claim 1, wherein the circuitry is configured to receive the predetermined noise level from a user input.

3. The information processing apparatus according to claim 1, wherein the circuitry is configured to randomly select the predetermined noise level.

4. The information processing apparatus according to claim 1, wherein, to receive the first audio data, the circuitry is configured to:  receive Musical Instruments Digital Interface (MIDI) data;  generate synthesized audio data from the MIDI data and at least one One Shot Sample;  send the synthesized audio data to a user; and  receive the first audio data from the user.

5. The information processing apparatus according to claim 4, wherein the circuitry is configured to generate the first audio data by performing pasting synthesis on the at least one One Shot sample based on the MIDI data.

6. The information processing apparatus according to claim 4, wherein the circuitry is configured to receive the at least one One Shot sample.

7. The information processing apparatus according to claim 4, wherein the circuitry is configured to receive a tone category and select the at least one One Shot sample based on the tone category.

8. The information processing apparatus according to claim 4, wherein the circuitry is configured to acquire a predetermined One Shot sample.

9. The information processing apparatus according to claim 4, wherein the user edits the synthesized audio data to generate the first audio data sent to the circuitry.

10. The information processing apparatus according to claim 9, wherein the user edits at least one of the MIDI data and the at least one One Shot sample.

11. The information processing apparatus according to claim 1, wherein the circuitry is further configured to:  receive a condition from a user; and  generate the second audio data further based on the condition.

12. The information processing apparatus according to claim 1, wherein the circuitry is configured to denoise the intermediate audio data using a deep generative model.

13. A method for generating audio data, comprising: receiving first audio data;  generating intermediate audio data by adding noise of a predetermined noise level to the first audio data; and  generating second audio data by denoising the intermediate audio data, wherein the second audio data is higher quality audio data than the first audio data.

14. The method according to claim 13, further comprising receiving the predetermined noise level from a user input.

15. The method according to claim 13, further comprising randomly selecting the predetermined noise level.

16. The method according to claim 13, wherein, receiving the first audio data includes:  receiving Musical Instruments Digital Interface (MIDI) data;  generating synthesized audio data from the MIDI data and at least one One Shot Sample;  sending the synthesized audio data to a user; and  receive the first audio data from the user.

17. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause an information processing apparatus to perform a method comprising:  receiving first audio data;  generating intermediate audio data by adding noise of a predetermined noise level to the first audio data; and  generating second audio data by denoising the intermediate audio data, wherein the second audio data is higher quality audio data than the first audio data.

18. The non-transitory computer-readable medium according to claim 17, wherein the method further includes receiving the predetermined noise level from a user input.

19. The non-transitory computer-readable medium according to claim 17, wherein the method further includes randomly selecting the predetermined noise level.

20. The non-transitory computer-readable medium according to claim 17, wherein, receiving the first audio data includes:  receiving Musical Instruments Digital Interface (MIDI) data;  generating synthesized audio data from the MIDI data and at least one One Shot Sample;  sending the synthesized audio to a user; and  receiving the first audio data from the user.

Citation Information

Patent Citations

  • Audio data generation using artificial intelligence

    KR102612572B1

  • Generating audio using generative neural networks

    WO2025109032A2