Graphical user interface for generating opponent network music synthesizer
By designing an information processing system that utilizes learning models and IC-GAN generation models, the problems of low controllability and low sound quality in the prior art are solved, and high-quality and controllable musical instrument sound generation are achieved.
Patent Information
- Application Number
- CN202380068133.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-09-07
- Publication Date
- 2025-05-02
AI Technical Summary
The prior art has the problem of low generation controllability when generating music using computers, especially when directly synthesizing the sound of music. On the other hand, although using MIDI to generate instrument sounds has the advantage of independent control, the generated sound quality is low.
An information processing system is designed to extract timbre feature quantities by receiving input sound and pitch information, and to generate musical instrument sounds with pitch using a learning model. The system uses IC-GAN generation model to generate high-quality instrumental sounds in interactive time.
It realizes the generation of high-quality instrumental sounds that reflect the input sound characteristics, with good controllability and timbre consistency, and improves the flexibility and sound quality of music generation.
Smart Images

Figure CN119923686A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims the benefit of Japanese Priority Patent Application JP 2022-164477 filed on Oct. 13, 2022, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The technology disclosed in this specification (hereinafter, “the present disclosure”) relates to an information processing device and an information processing method, a computer program, a sound generating system, and an information terminal that perform information processing related to music production. Background Art
[0004] The development of artificial intelligence (AI) technology is remarkable, and recognition technology of images, voices, etc. using learning models has become common. Recently, image generation technology has also been developed, which uses a generative adversarial network (GAN) to generate complex images. In addition, methods for making music using AI technology are also being sought. For example, a music sound enhancement device that uses a deep neural network (DNN) that reflects the characteristics of the sound of a musical instrument to enhance the sound source (see patent document 1), an information processing method that automatically generates various kinds of music using a learning model generated using a GAN or a variational autoencoder (VAE) (see patent document 2), etc. have been proposed.
[0005] [Citation List]
[0006] [Patent Document]
[0007] Patent Document 1: JP 2019-78864A
[0008] Patent Document 2: JP 2019-78864A
[0009] Non-patent literature
[0010] Non-patent literature 1: Arantxa Casanova, Marl'ene Careil, Jakob Verbeek, Michal Drozdzal and Adririana Romero-Soririano, "Instance-conditioned GAN", in Advances in Neural Information Processing Systems (NeurIPS), 2021.
[0011] Non-patent literature 2: Jesse Engel, CiCinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan, "Neural audio synthesis of musical notes with wavenet autoencoders", in International Conference on Machihine Learning. PMLR, 2017, pp. 1068-1077.
[0012] Non-patent document 3: Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oririol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and KorayKavukcuoglu, "Wavenet: A generative model for raw audio", arXiv prepririntarXiv:1609.03499, 2016.
[0013] Non-patent literature 4: Jesse Engel, Kumar Kririshna Agrawal, Shuo Chen, IshaanGulrajani, Chriss Donahue, and Adam Roberts, "Gansynth: Adversarial neuralaudio synthesis", in International Conference on Learning Representations, 2018.
[0014] Non-patent literature 5: Sean Vasquez and Mike Lewis, “Melnet: A generative model for audio in the frequency domain”, arXiv preprint arXiv:1906.01083, 2019. Summary of the invention
[0015] Technical issues
[0016] The two approaches to generate music using computers are mainly a method of directly synthesizing the sound of music including melody and accompaniment and a method of synthesizing monophonic instrument sounds and playing a musical instrument digital interface (MIDI). In the former, although music can be generated end-to-end, there is a problem of low controllability of the generation. On the other hand, the latter has the following advantages: the generation of MIDI and the design of timbre can be controlled independently, and the quality of the generated sound is high.
[0017] Therefore, for example, the present invention provides an information processing apparatus and an information processing method, a computer program, a sound generating system, and an information terminal that perform information processing related to generation of musical instrument sounds that can be used for MIDI playback.
[0018] Solution to the problem
[0019] The present disclosure is made in view of the above problems, and its first aspect is an information processing system including: a circuit system configured to receive an input sound and pitch information; extract a timbre feature quantity from the input sound; and generate information of a musical instrument sound with a pitch based on the timbre feature quantity and the pitch information.
[0020] The circuit system uses the learned model to generate information about the instrument's sound.
[0021] The circuit system generates information of a musical instrument sound having a pitch using a learning model using the information after preprocessing the input sound and the pitch information as an example condition.
[0022] Furthermore, the circuit system is configured to extract the timbre feature quantity of the input sound so that pitch information is not retained.
[0023] Furthermore, the circuit is configured to extract the timbre feature using a timbre feature extractor that has performed opponent learning with respect to pitch.
[0024] Another aspect of the present invention is directed to an information processing method, comprising: receiving an input sound and pitch information; extracting a timbre feature quantity from the input sound; generating information of a musical instrument sound having a pitch based on the timbre feature quantity and the pitch information;
[0025] Another aspect of the present disclosure is directed to one or more non-transitory computer-readable media, which, when executed by a circuit system, causes the circuit system to: receive an input sound and pitch information; extract a timbre feature quantity from the input sound; and generate information of a musical instrument sound having a pitch based on the timbre feature quantity and the pitch information.
[0026] The computer program can be obtained by defining a computer program described in a computer-readable format so as to implement a predetermined process on a computer. The computer program can be provided to a computer capable of executing various programs or codes through a storage medium provided in a computer-readable format, a communication medium (e.g., a storage medium such as an optical disk, a magnetic disk, a semiconductor memory, etc.), or a communication medium such as a network, etc. Then, by installing the computer program according to the third aspect of the present disclosure in a computer via any medium, a collaborative action is applied on the computer, and operations and effects similar to those of the information processing device according to the first aspect of the present disclosure can be obtained.
[0027] In addition, another aspect of the present invention is directed to a sound generation system, which includes: a terminal configured to request generation of a musical instrument sound; and an information processing device that generates a musical instrument sound, wherein the information terminal is configured to receive an input sound and pitch information; extract a timbre feature quantity from the input sound; and generate information of a musical instrument sound with a pitch based on the timbre feature quantity and the pitch information.
[0028] However, the "system" described herein refers to a logical assembly of multiple devices (or functional modules that implement specific functions), and each of these devices or functional modules may or may not be in a single housing. That is, an apparatus including multiple components or functional modules and an assembly of multiple apparatuses correspond to a "system".
[0029] In addition, another aspect of the present disclosure relates to an information terminal, comprising: a communication interface configured to communicate with an information processing system; and a user interface configured to receive instructions related to a musical instrument sound including generating an input sound and pitch information, wherein the communication interface is configured to send a request to generate a musical instrument sound including an input sound and pitch information to the information processing system, and receive information about the musical instrument sound from the information processing system.
[0030] Advantageous Effects of the Invention
[0031] According to the present disclosure, it is possible to provide an information processing apparatus and information processing method, a computer program, a sound generating system, and an information terminal that perform information processing of generating a musical instrument sound having a pitch reflecting a feature of an arbitrary input sound.
[0032] It should be noted that the effects described in this specification are merely examples, and the effects to be brought about by the present disclosure are not limited thereto. In addition, in addition to the above-described effects, the present disclosure may further exhibit additional effects in some cases.
[0033] Other objects, features, and advantages of the present disclosure will become apparent through more detailed description based on the embodiments described below and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a diagram showing an outline of the present disclosure.
[0035] Figure 2 is a simplified diagram showing a mechanism for generating instrument sounds based on inspiration obtained by mixing two input sounds.
[0036] Figure 3 is a diagram illustrating a functional configuration of a musical instrument sound generating system 300 that generates musical instrument sounds having pitches reflecting characteristics of input sounds.
[0037] Figure 4 is a flowchart showing processing steps for generating a musical instrument sound having a pitch that reflects the characteristics of an input sound.
[0038] Figure 5 is a diagram illustrating an example of a workflow when learning a timbre feature extractor.
[0039] Figure 6 is a diagram illustrating a workflow at the time of learning of a timbre feature extractor in the case of performing opponent learning regarding pitch.
[0040] Figure 7 is a diagram showing an overview of a model in a case where a generator used in the generation unit 303 of the musical instrument sound generation system 300 is learned.
[0041] Figure 8 2 is a diagram illustrating a process flow for reconstructing audio waveform data from a mel-spectrogram.
[0042] Fig. 9 : is a diagram showing the frequency scale conversion process flow of the iterative method of repeatedly updating and correcting to a non-negative value by the gradient method.
[0043] Fig.10 : is a diagram showing the flow of frequency scale conversion processing by an iterative method in the case of setting an initial value of a solution using a least square method with no non-negative value constraint.
[0044] Fig.11 is a diagram showing a configuration of a musical instrument sound generating system 300 including a client server model 1100 .
[0045] Fig.12 1102 and 1101 are diagrams illustrating an exemplary processing sequence performed by the client 1102 and the server 1101 .
[0046] Fig.13 is a diagram showing a configuration example of a GUI screen for generating and reproducing musical instrument sounds with pitches according to an embodiment of the present disclosure.
[0047] Fig.14 It is shown Fig.13 An illustration of an embodiment of operations on a GUI screen is shown.
[0048] Fig.15 It is shown Fig.13 An illustration of an embodiment of operations on a GUI screen is shown.
[0049] Fig.16 It is shown Fig.13 An illustration of an embodiment of operations on a GUI screen is shown.
[0050] Fig.17 It is shown Fig.13 An illustration of an embodiment of operations on a GUI screen is shown.
[0051] Fig.18 It is shown Fig.13 An illustration of an embodiment of operations on a GUI screen is shown.
[0052] Fig.19 It is shown Fig.13 An illustration of an embodiment of operations on a GUI screen is shown.
[0053] Fig. 20 It is shown Fig.13 An illustration of an embodiment of operations on a GUI screen is shown.
[0054] Fig.21 It is shown Fig.13 An illustration of an embodiment of operations on a GUI screen is shown.
[0055] Fig. 22 It is shown Fig.13 An illustration of an embodiment of operations on a GUI screen is shown.
[0056] Fig.23 It is shown Fig.13 An illustration of an embodiment of operations on a GUI screen is shown.
[0057] Fig.24 is a diagram illustrating another configuration example of a GUI screen for generating and reproducing musical instrument sounds with pitches according to an embodiment of the present disclosure.
[0058] Fig.25 It is shown Fig.24 Illustration of a modified example of the GUI screen shown.
[0059] Fig.26 2000 is a diagram showing a specific hardware configuration example of the information processing device 2000 .
[0060] Fig. 27 is a diagram showing an overview of the IC-GAN model. DETAILED DESCRIPTION
[0061] In the following description, the present disclosure will be explained in the following order with reference to the accompanying drawings.
[0062] A. Overview
[0063] B. Sound Generation System
[0064] B-1. System Configuration
[0065] B-2. System Operation
[0066] B-3. System Features
[0067] B-4. System Operation
[0068] C. Extraction of timbre features
[0069] D. Generation of musical instrument sounds with pitch-reflective properties of the input sound
[0070] E. Reconstruction of audio waveform
[0071] F. Operational form of client-server model
[0072] G. GUI Configuration and Operation Example
[0073] G-1. First embodiment
[0074] G-2. Second embodiment
[0075] H. Configuration of information processing equipment
[0076] I. Comparison with Related Studies
[0077] A. Overview
[0078] The present invention is a technology for generating monophonic instrument sounds. The instrument sounds generated based on the embodiments of the present invention are used to realize music production on a computer by playing MIDI. According to the music generation method using the embodiments of the present invention, the generation of MIDI and the design of timbre can be independently controlled, which has the advantage of improving the quality of the generated sound.
[0079] The present disclosure is a technology for generating a monophonic instrument sound, but the instrument sound can be generated based on inspiration obtained from, for example, any input sound generated in a human living space. In addition, the instrument sound generated by the present disclosure is not limited to the instrument sound generated using a real instrument. That is, the instrument sound generated by the present disclosure is a sound that is unlikely to be identified as an instrument but is not identified as a sound generated by an instrument, in other words, a sound that is not identified as a sound generated by a sound other than an instrument.
[0080] Figure 1 The outline of the present disclosure is schematically shown. In the present disclosure, any input sound for which inspiration is given is, for example, various sounds generated in the living environment of human beings, and includes not only natural sounds and environmental sounds, but also artificial sounds artificially generated in advance. For example, the noise can be the sound of a conversation, the singing of a karaoke, the cry of an animal such as a dog or a bird, the environmental sound such as a deer peeking or the sound of rain, the sound of a wind chime or the sound of wind, or the noise of cutting or crushing an object with a chain saw, a heavy machine, etc. For the convenience of computer processing, these optional input sounds are input as audio files (such as wav format files).
[0081] Then, in the present disclosure, based on the inspiration obtained from the input sound, a musical instrument sound is generated. Specifically, according to the present disclosure, for example, a monaural musical instrument sound with a specified pitch in a relatively short time of about one second or several seconds is output as MIDI data. The musical instrument sound generated by the present disclosure can be a single sound generated by an existing musical instrument (e.g., a keyboard instrument, a percussion instrument, a string instrument, a wind instrument, or an electric or electronic musical instrument), but is not limited thereto, and a completely new and unique musical instrument sound can be generated. The unique musical instrument sound generated by the present disclosure is a sound that is unlikely to be identified as a musical instrument but is not identified as a sound generated by a musical instrument (in other words, a sound generated by other musical instruments).
[0082] In the present disclosure, a generative model of deep learning is used to generate instrument sounds as a condition of inputting sounds and pitches. Therefore, according to the present disclosure, there is an effect that users can freely customize instruments and music can be associated with sounds. The deep learning generative model referred to here is a learning model unique to the present disclosure, specifically a generator for generating instrument sounds using the concept of instance-conditioned GAN (IC-GAN) generated by a GAN framework (hereinafter also referred to as the "disclosure model").
[0083] In the past, as methods for synthesizing musical instrument sounds, there were methods such as "synthesizers," which modulate artificially generated periodic oscillator waveforms to control timbre, and "sampling," which records and processes actual instrument sounds to express the authenticity of acoustic instruments that are difficult to synthesize with synthesizers. Sampling can directly utilize any sound to generate music, but it cannot generate completely new timbres or combine the characteristics of multiple sounds.
[0084] On the other hand, according to the present disclosure, it is possible to search the latent space, generate a variety of new and unique instrument sounds, and perform intelligent sound synthesis processing that combines the characteristics of multiple sounds by using a generative model of deep learning. In addition, according to the present disclosure, it is possible to create new and unique instrument sounds by mixing two or three or more arbitrary input sounds in the latent representation of a generative model of deep learning.
[0085] Figure 2A mechanism for generating a musical instrument sound according to the present disclosure as a revelation obtained by mixing two input sounds is schematically shown. Figure 2 In the embodiment described in , the audio waveform of a speaker as the first input sound and the audio waveform of a dog barking as the second input sound are captured as wav format files. First, each input sound is subjected to a timbre feature extractor to obtain a feature vector h respectively. t and h d . Next, each feature vector is synthesized at a mixing ratio specified by a user or the like. Then, a unique instrument sound based on the excitation obtained from the first input sound and the second input sound is generated. Specifically, the learning model (generator) generated by the disclosed model generates a unique instrument sound under the condition of the feature vector of the synthesized timbre and pitch specified by the user.
[0086] The user can control the sound of the instrument by adjusting the mixing ratio of multiple input sounds. Figure 2 In the embodiment shown in the upper part of FIG. , the feature vectors h of the first input sound and the second input sound are t and h d Mix in a ratio of 0.5:0.5 to generate the combined feature vector h s1 Then, we use the generative model of deep learning to extract the feature vector h s1 A unique instrument sound S1 similar to a trumpet having a dog's timbre is generated. Figure 2 In the embodiment shown in the lower part of FIG. , the feature vectors h of the first input sound and the second input sound are t and h d Mix in a ratio of 0.8:0.2 to generate the combined feature vector h s2 Then, a generative model based on deep learning is used to generate a unique instrumental sound S2 similar to a dog being dialed with the timbre of a trumpet.
[0087] The technique for generating musical instrument sounds using a generative model based on deep learning according to the present disclosure has the following features (1) to (4).
[0088] (1) The user's arbitrary vocal affordances can be input into a generative model such as a sampler and effectively generalized to a variety of input voices.
[0089] (2) Multiple sounds can be mixed via a latent space by using a generative model using deep learning.
[0090] (3) A wide range of pitches can be generated with accurate and consistent timbre.
[0091] (4) Instrument sounds can be generated during the interactive time.
[0092] The present disclosure applies a generative model generated by IC-GAN from the perspective of enabling input to the model and improving the generalization characteristics of the input. IC-GAN adjusts the generator and discriminator through the feature quantity of the data point (i.e., the embodiment), and represents the distribution of the entire data as a superposition of the local distribution near the embodiment. IC-GAN is a new GAN learning technology that can achieve input to the model and avoid mode collapse.
[0093] B. Sound Generation System
[0094] B-1. System Configuration
[0095] Figure 3 The functional configuration of the musical instrument sound generating system 300 according to the present invention for generating musical instrument sounds whose pitches reflect the characteristics of input sounds is schematically described.
[0096] See also Figure 3 The musical instrument sound generation system 300 includes a waveform spectrogram conversion unit 301, a timbre feature extraction unit 302, a generation unit 303, and a spectrogram waveform inverse conversion unit 304. The timbre feature extraction unit 302 and the generation unit 303 are implemented using DNN. The musical instrument sound generation system 300 receives an input sound including a short-time audio waveform (wav file), a pitch, and a random number, and outputs a musical instrument sound having a pitch reflecting the characteristics of the input sound. The output musical instrument sound has a length of about one second or several seconds.
[0097] The musical instrument sound generating system 300 receives an input sound including an audio waveform as a wav format file. The spectrogram conversion unit 301 generates a linear spectrogram of the input sound by short-time Fourier transform, and further performs logarithmic scale conversion on the linear spectrogram to generate a mel spectrogram. In the case where two or more input sounds are input to the musical instrument sound generating system 300, the spectrogram conversion unit 301 converts the audio waveform of each input sound into a mel spectrogram.
[0098] Here, the spectrogram corresponds to a so-called voiceprint, in which the frequency spectra of each audio data segment (frame) obtained by extracting the frequency components and amplitude components of the audio signal from the audio waveform by Fourier transform are arranged along the time axis. In the drawings attached to this specification, the spectrogram is illustrated as a two-dimensional graph in which the intensity (amplitude) of the signal component in each of the time component and the frequency component is visualized by shading. In addition, the Mel spectrogram is a logarithmic Mel spectrogram calculated by applying a Mel filter bank to a linear spectrogram, which extracts only specific frequency bands at equal intervals in the Mel scale, focusing on the fact that the human ear does not directly hear the sound of the actual frequency, and hears the sound close to the upper limit of the audible range lower than the actual sound. In addition, the melting scale is a scale based on human hearing (i.e., how to hear the sound).
[0099] The timbre feature extraction unit 302 extracts the timbre feature quantity h from the mel-spectrogram that visualizes and expresses the audio waveform of the input sound. In the case where two or more input sounds are input to the musical instrument sound generation system 300, the timbre feature extraction unit 302 extracts the timbre feature quantities h1, h2, ... from the mel-spectrogram ( Figure 3 not shown).
[0100] The timbre feature extraction unit 302 extracts timbre feature quantities from the mel-spectrogram (image information) using a learning model configured by, for example, a convolutional neural network (CNN) and learned in advance. The timbre feature quantity is specifically an n-dimensional (here n is a positive integer) feature vector.
[0101] The generating unit 303 generates a mel-spectrogram of a musical instrument sound having a pitch reflecting the characteristics of the input sound using the timbre feature quantity, pitch and random number of the input sound as input. When two or more input sounds are input to the musical instrument sound generating system 300 and the timbre feature extraction unit 302 extracts a plurality of timbre feature quantities (feature vectors) h1, h2, ... from the mel-spectrogram of each input sound, a mixture of the timbre feature quantities at a specified mixing ratio is input to the generating unit 303.
[0102] The generation unit 303 generates a mel-spectrogram of a musical instrument sound with a pitch using a generative model of deep learning. Specifically, the generation unit 303 uses a learning model (generator) generated by the model of the present invention to generate a mel-spectrogram of a musical instrument sound with a pitch that reflects the characteristics of the input sound, using the timbre feature quantity and pitch of the input sound as an example condition.
[0103] The spectrogram waveform inverse conversion unit 304 performs Fourier inverse transformation on the mel spectrogram generated by the generation unit 303 to reconstruct the audio waveform data including, for example, a wav format file. The reconstructed audio waveform has a length of about one second or several seconds. There is a problem that the conversion process from the mel spectrogram to the audio waveform is slow, but this will be described in detail later.
[0104] B-2. System Operation
[0105] Figure 4 The processing procedure for generating a musical instrument sound in which the pitch reflects the characteristics of the input sound in the musical instrument sound generating system 300 is described in the form of a flowchart.
[0106] First, the target input sound specified by the user is input to the musical instrument sound generation system 300 (step S401). In this step, for example, the file name of the wav format file as the sound source of the input sound is specified. In addition, in the case where the user specifies more than two input sounds, the wav format file of each input sound is obtained in this step.
[0107] Next, the spectrogram conversion unit 301 converts the audio waveform of the input sound into a mel spectrogram (step S402). That is, the spectrogram conversion unit 301 generates a linear spectrogram of the input sound by Fourier transform, and further performs logarithmic scale conversion on the linear spectrogram to generate a mel spectrogram. In the case where more than two input sounds have been input in step S401, the spectrogram conversion unit 301 generates mel spectrograms for all input sounds in step S402.
[0108] Next, the timbre feature extraction unit 302 extracts a timbre feature value h from the mel-spectrogram of the input sound (step S403 ).
[0109] In the case where two or more input sounds have been input in step S401 (yes in step S404), in step S403, the timbre feature extraction unit 302 extracts timbre feature quantities h1, h2, ... from the mel-spectrograms of all the input sounds, and further mixes the corresponding timbre feature quantities h1, h2, ... to generate the timbre feature quantity h (step S405).
[0110] In the case where the mixing ratio is specified for each input sound, in step S405, the timbre feature quantities h1, h2, ... of each input sound are weighted averaged and mixed according to the specified mixing ratio. In addition, in the case where the mixing ratio is not specified, the timbre feature quantities h1, h2, ... of each input sound may be averaged to perform the mixing process. Here, when N input sounds are input in step S401, the timbre feature quantities h1, h2, ..., h2 are generated from the mel-spectrogram of each input sound. N , and specify the mixing ratio r of the i-th input sound i (where r1+r2+...+r N =1), in step S405, the mixed timbre feature quantity h can be generated according to the following expression (1).
[0111] [Mathematical formula 1]
[0112]
[0113] Next, the pitch information of the musical instrument sound to be generated designated by the user is input to the musical instrument sound generating system 300 (step S406 ). However, in step S401 , the pitch information may be input to the musical instrument sound generating system 300 simultaneously with the input sound.
[0114] Then, the generation unit 303 generates a Mel-spectrogram of a musical instrument sound having a pitch reflecting the characteristics of the input sound based on the timbre feature quantity h obtained in step S403 or S405 and the pitch information obtained in step S406 (step S407). Specifically, using the learning model (generator) generated by the model of the present invention, the generation unit 303 uses the timbre feature quantity and pitch of the input sound as example conditions, and generates a Mel-spectrogram of a musical instrument sound having a pitch reflecting the characteristics of the input sound based on the random number generated by the musical instrument sound generation system 300.
[0115] Then, the spectrogram waveform inverse conversion unit 304 performs Fourier inverse conversion on the mel spectrogram generated by the generation unit 303, reconstructs the audio waveform data including, for example, a wav format file, and outputs the audio waveform data as MIDI data (step S408), and ends the present process. The output instrument sound has a length of about one second or several seconds.
[0116] B-3. System Features
[0117] The above has schematically described the configuration and operation of the sound reproducing system 300. The sound reproducing system 300 has the following features.
[0118] (1) Using the learning model, the musical instrument sound generation system 300 can generate a musical instrument sound having a pitch that reflects the timbre of the input sound within the interaction time.
[0119] (2) By using example adjustment, the quality of the generated instrument sounds and the ability to generate instrument sounds can be improved.
[0120] (3) By performing adversarial learning on the pitch of the timbre feature extractor, pitch accuracy and timbre consistency can be improved.
[0121] B-4. System Operation
[0122] Figure 3 The musical instrument sound generating system 300 shown in is installed on an information processing device including, for example, a computer. The process of generating a learning model (generator) using the model of the present disclosure and the process of generating a musical instrument sound having a pitch reflecting the characteristics of an input sound using the learning model (generator) have a large computational load. Therefore, a client-server model is also assumed as an operation mode of the musical instrument sound generating system 300.
[0123] In this case, on the client side, for example, a wav format file of an input sound is selected (in the case of selecting multiple input sounds, the mixing ratio of each input sound is also specified) and the pitch is specified through the user's graphical user interface (GUI) operation, and the server is requested to generate a musical instrument sound with a pitch reflecting the characteristics of the input sound. On the other hand, on the server side, with the input sound (its timbre feature quantity) and pitch specified from the client side as example conditions, a musical instrument sound with a pitch reflecting the characteristics of the input sound is generated and returned to the client side as a request source. The details of this operation form will be described later (Part F).
[0124] C. Extraction of timbre features
[0125] As described above in Part B, the timbre feature extraction unit 302 extracts the timbre feature quantity h from the mel-spectrogram that visualizes and expresses the audio waveform of the input sound. Specifically, the timbre feature extraction unit 302 is a feature extractor that extracts the timbre feature quantity from the mel-spectrogram (image information) using a learning model configured by CNN and learned in advance.
[0126] It is important to use a high-quality feature extractor to learn the instance-conditioned GAN described in Section D below. The simplest way to obtain a feature extractor is to learn a discriminator that lives with labeled training data, and use the output of the layer immediately before the final fully connected layer as the feature volume.
[0127] However, when the feature extractor learned according to the above method is applied to the timbre feature extraction unit 302, there is a problem that the pitch accuracy and timbre consistency of the sound generated by the musical instrument sound generation system 300 deteriorate. For example, there is a problem that when the pitch of the input sound is "Re", the pitch of the generated sound is "Re" which is close to "Mi". The following two points are considered to be the causes of this problem.
[0128] (a) The feature quantity of the feature extractor learned by the general method includes pitch information.
[0129] (b) Since the pitch specified by the user and the pitch information included in the feature quantity interfere with each other, the learning of the generator (used by the generation unit 303) becomes unstable. For example, in the case where the pitch specified by the user is C4 and the feature quantity extracted by the feature extractor contains G4 as pitch information, the generator at the subsequent stage cannot determine which instrument sound of C4 or G4 can be generated.
[0130] Therefore, in the present invention, learning of the timbre feature extractor is performed so that the timbre feature quantity in which no pitch information remains can be extracted from the mel-spectrogram. Specifically, in the present invention, opponent learning about pitch is performed on the timbre feature extractor so that the timbre feature quantity is not retained in the timbre feature quantity.
[0131] Figure 5 This example illustrates the workflow when learning a timbre feature extractor. Figure 5 In the embodiment described in the above, the timbre feature extractor 501 used in the timbre feature extraction unit 302 is learned together with the musical instrument identifier 502. As described above, the timbre feature extractor 501 extracts the timbre feature quantity h from the mel-spectrogram of the audio waveform. In addition, the musical instrument identifier 502 identifies the musical instrument with the original audio waveform from the timbre feature quantity h. Next, the predicted distribution C output by the musical instrument identifier 502 is compared. pred Distribution of correct answers C gt , and the timbre feature extractor 501 and the musical instrument discriminator 502 are learned by error back propagation. For example, the learning stage in which the musical instrument discriminator 502 is fixed and the timbre feature extractor 501 is learned and the learning stage in which the timbre feature extractor 501 is fixed and the musical instrument discriminator 502 is learned are repeated alternately.
[0132] However, in Figure 5 In the learning method shown, it is difficult to prevent the pitch information from remaining in the feature h extracted by the timbre feature extractor 501. There is a problem that it is difficult to accurately generate the sound of a musical instrument of a specific pitch. This is because, as described in the following section D, the generator G and the discriminator D input both the timbre feature h and the pitch information p, and therefore, if the timbre feature includes the pitch information, it is confused which pitch information is correct, and proper learning becomes difficult.
[0133] Figure 6 The following describes the workflow when learning the timbre feature extractor in the case where the opponent learning about the pitch is performed so that the pitch information of the feature amount is not retained. Figure 6 In the illustrated embodiment, the timbre feature extractor 601 used in the timbre feature extraction unit 302 is learned together with the musical instrument identifier 602 and the pitch identifier 603. Specifically, in the present embodiment, the musical instrument identifier 602 and the pitch identifier 603 are learned simultaneously, and the learning is performed so that the timbre feature quantity extracted by the timbre feature extractor 601 cannot be used to identify the pitch, thereby avoiding the pitch information from being retained in the timbre feature quantity extracted by the timbre feature extractor 601.
[0134] As described above, the timbre feature extractor 601 extracts the timbre feature h from the mel-spectrogram of the audio waveform. In addition, the instrument identifier 602 identifies the instrument having the original audio waveform from the timbre feature h. The learning of the instrument identifier 602 is similar to Figure 5 The workflow described in the description will be omitted here.
[0135] In addition, adversary learning on pitch is performed on the timbre feature extractor 601 so that pitch information is not retained in the timbre feature quantity. The pitch discriminator 603 discriminates the pitch of the original audio waveform from the timbre feature quantity h.
[0136] First, the pitch discriminator 603 performs learning so that the pitch of the original audio waveform can be accurately distinguished from the timbre feature quantity extracted by the timbre feature extractor 601. That is, the predicted distribution C output by the pitch discriminator 603 2,pred Distribution of correct answers C 2,gt By comparing, the pitch discriminator 603 is learned by back-propagating the error. In this way, after the learning of the pitch discriminator 603 is performed, the pitch discriminator 603 is then fixed, and the learning of the timbre feature extractor 601 is performed so that the timbre feature quantity h with no tonality information remaining can be generated. That is, the prediction distribution C output from the pitch discriminator 603 is 2,pred becomes uniformly distributed C 2,uni , in other words, the learning of the timbre feature extractor 601 is performed so that the timbre feature quantity h that cannot be recognized by the key in the pitch discriminator 603 can be generated.
[0137] according to Figure 6 The learning method of the pitch-invariant feature extractor based on adversary learning shown has the effect of avoiding the instability of GAN learning due to the entanglement of timbre and pitch information in the feature quantity space, and can improve the accuracy of pitch and the consistency of timbre.
[0138] The bad learning about the timbre will be described more specifically. φ (x)(=h) is taken as input, and the shallow MLP for instrument identification and pitch identification of the pitch discriminator 603 is respectively composed of C i and C p By alternately optimizing the adversarial learning of the loss functions shown in the following expressions (2) and (3), a timbre feature extractor f that can extract timbre feature quantities that do not contain pitch information can be obtained. φ .
[0139] [Mathematical formula 2]
[0140]
[0141] [Mathematical formula 3]
[0142]
[0143] In the above expressions (2) and (3), i(x) and p(x) represent the instrument label and the pitch label of the sample x, respectively. In addition, CE is the cross entropy, and KL is the Kullback-Leibler divergence.
[0144] The first term of the above expression (2) updates the feature extractor f φ and instrument identification C i , so that the instrument can be correctly identified. On the other hand, the second term of the above expression (2) updates the timbre feature extractor f φ , which allows the use of the feature quantity f φ (x) to distinguish the pitch, that is, the predicted distribution of the pitch is close to the uniform distribution. On the other hand, the above expression (3) updates C p In order to give a given feature value f φ By performing this adversarial learning, a timbre feature extractor capable of extracting timbre feature quantities so that pitch information is not retained can be obtained.
[0145] In fact, when the timbre feature extractor f is learned φ When only the pitch discrimination of the discriminator 603 is fixed and then learned using the above expression (3), it has been confirmed that pitch discrimination can be performed with an accuracy of 17% or more without using opponent learning, while pitch discrimination is reduced to 2.5% in opponent learning on pitch.
[0146] D. Generation of musical instrument sounds with pitch-reflective properties of the input sound
[0147] As described in the above section B, the generation unit 303 uses the timbre feature quantity, pitch and random number of the input sound as input to generate a mel-spectrogram of a musical instrument sound having a pitch reflecting the characteristics of the input sound. Specifically, the generation unit 303 uses the learning model (generator) generated by the model of the present invention to generate a mel-spectrogram of a musical instrument sound having a pitch reflecting the characteristics of the input sound, wherein the timbre feature quantity and pitch of the input sound are example conditions.
[0148] Here, as known in the art, GAN is a deep learning model that makes two neural networks, a discriminator (D) that distinguishes between real data and artificial data, and a generator (G) that generates data from noise, compete for learning. In GAN, there is a problem of mode collapse, in which the quality and diversity of the generated samples are impaired because the generated data is biased to a part of the training data. On the other hand, IC-GAN is a new technology for learning GAN, which solves the problem of mode collapse by adjusting the feature amount corresponding to the data point (instance) by the discriminator D and the generator G and teaching the vicinity of the data point to the discriminator as real data.
[0149] Fig. 27The outline of the IC-GAN model is shown (for example, see Non-Patent Document 1). The discriminator D and the generator G are respectively installed using DNN. In the figure, as an example, the input image x i By feature extractor f φ Mapped to the feature space. Then, as the feature extractor f φ The output image x is obtained i The characteristic quantity h i is input to each of the generator G and the discriminator D. The generator G generates i The extracted feature h i and sampled noise (random number) z to generate image x g In addition, the discriminator D is based on the feature h i and the adjacent image x as the actual sample n Identify the image x generated by the generator G g Then, the generator G enables the discriminator D to learn competitively and be able to distinguish the generated image x g and the adjacent images x generated by the generator G n , the discriminator D is able to make the generated image x g and adjacent images x n As a result, an accurate image X can be generated. g The generator G, the exact image X g It is not allowed to determine the authenticity in the discriminator D.
[0150] In this embodiment, the feature extractor f φ Corresponding to the timbre feature extraction unit 302, the generator G corresponds to the generation unit 303. Then, the input image x i Corresponding to the Mel-spectrogram of the visualized input sound, the feature h i Corresponding to the timbre feature quantity, the generated image x g Corresponds to the mel-spectrogram generated by the generator G.
[0151] The disclosed model is a generative model learned by a GAN framework that uses ideas from IC-GAN to generate instrument sounds (described above). Figure 7 The outline of the disclosed model is shown in the case where the generator used in the generation unit 303 of the musical instrument sound generation system 300 is learned. In this case, the generator G generates a musical instrument sound having a pitch reflecting the characteristics of the input sound from the input sound, the pitch information p, and the noise vector z. In addition, the discriminator D discriminates the true / false of the sound generated by the generator G.
[0152] The input sound is converted into a logarithmic scale Mel spectrogram x by short-time Fourier transform in the spectrogram conversion unit 301. iThen, the timbre feature extraction unit 302 uses the feature extractor f described in the above section C φ The Mel-spectrogram x i Mapped to the timbre feature h i The generator G combines the single vector of the pitch information p and the noise vector z sampled from the standard normal distribution with the timbre feature h i Input together to generate the Mel-spectrogram x g The generated mel-spectrogram is reconstructed into an audio waveform by the spectrogram waveform inverse conversion unit 304 described in Section E described later. The audio waveform is a musical instrument sound having a pitch reflecting the characteristics of the input sound.
[0153] General class conditioning divides the distribution of all data into multiple distributions without overlap according to the number of classes. On the other hand, instance conditioning in IC-GAN (see Non-Patent Document 1) attempts to obtain a complex data distribution by dividing the distribution of the entire data into a large number of local distributions with overlap. By using instance x i The characteristic quantity h i =f φ (x i ) and the monostable vector p of pitch information to adjust both the generator G and the discriminator D, for example x i The local distribution P(x|h i , p) is modeled, and the distribution P(x) of the entire data is expressed as the following expression (4) as its superposition.
[0154] [Formula 4]
[0155]
[0156] The learning process of the generative model follows Non-Patent Document 1. i , in the feature extractor f φ The dataset of the Saint-Genius obo with L2 distance k in the feature space defined by (·) is set to A i At this time, if Figure 7 As shown, based on the uniform distribution from A i Sample adjacent data points x j .x j and the generated sample x as the real sample g Together with the actual sample x j The corresponding fundamental tone p(x j ) is input as a condition to the generator G and the discriminator D. In the present embodiment, the generator G and the discriminator D are optimized by the min-max game shown in the following expression (5).
[0157] [Formula 5]
[0158]
[0159] E. Reconstruction of audio waveform
[0160] The method for reconstructing audio waveform data from a mel-spectrogram mainly includes two methods based on learning and pasting optimization. In the text-speech domain, methods for obtaining a vocoder by learning have been actively studied. However, the present invention is intended to generate instrument sounds of various timbres and pitches, and it is not necessarily easy to obtain a general vocoder that can cope with the sounds generated. On the other hand, there are such research results: by generating a high-resolution mel-spectrogram in the frequency direction, various sounds including music can be synthesized at a certain level of sound quality even when using an optimized audio inversion (see non-patent document 5). Therefore, in the present disclosure, audio waveform data is reconstructed from a mel-spectrogram by adopting an optimized method.
[0161] Figure 8 The general process flow for reconstructing audio waveform data from a mel-spectrogram via an optimization-based approach is schematically shown.
[0162] The mel spectrogram is a logarithmic mel spectrogram calculated by applying a mel filter bank, which extracts only specific frequency bands at equal intervals in the mel scale of the linear spectrogram based on human hearing. Therefore, the frequency scale conversion unit 801 converts the mel spectrogram generated by the generation unit 303 into a linear spectrogram on the frequency scale. Next, the phase recovery unit 802 uses, for example, the known Griffin-Lim algorithm to restore the phase of the linear spectrogram. Then, the inverse short-time Fourier transform unit (iSTFT) 803 performs an inverse Fourier transform to reconstruct an audio waveform. The audio waveform is an audio waveform of a musical instrument sound having a pitch generated by the musical instrument sound generation system 300.
[0163] exist Figure 8 In the processing flow shown in , specifically, the frequency scale conversion unit 801 performs the frequency scale conversion of the mel-spectrogram to the linear spectrogram, which has the problem that the computational cost is high and the processing becomes a bottleneck. The frequency scale conversion from the mel-spectrogram to the linear spectrogram can be formulated as a least squares problem with non-negative value constraints, as in the following expression (6), but the general solution has a large amount of computation. In the following expression (6), F mel is the Mel filter bank matrix, x mel is a Mel-scale spectrogram, and x lin is a linear scale spectrum plot.
[0164] [Mathematical formula 6]
[0165]
[0166] In addition, a method of obtaining a good solution by an iterative method in which the gradient method is repeatedly updated and corrected to a non-negative value can be considered. However, since the initial value is set by a random number, a sufficient number of iterations is required to converge to a good solution.
[0167] Fig. 9 An overview of the frequency scale conversion processing flow of an iterative method that is updated and corrected to a non-negative value by a repeated gradient method is shown. In the prior art, a software library for performing frequency scale conversion based on the processing flow shown has been provided. In this processing flow, first, an initialization unit 901 initializes a spectrum graph with a random number (for example, the intensity (amplitude) of each point on the time axis and the frequency axis is given as a random number). Then, an update unit 902 updates the intensity (amplitude) of each point on the time axis and the frequency axis by a gradient method, and then a correction unit 903 substitutes 0 into a variable with a negative value. The processing of the update unit 902 and the correction unit 903 is repeated until the calculation result converges. As already mentioned, the iterative method of repeating the update and correction of non-negative values by the gradient method is effective, but the convergence speed is slow.
[0168] Therefore, in the present disclosure, a similar iterative method is basically used, but a solution of the least square method without non-negative value constraints is used instead of a random number for spectrum initialization. Specifically, after obtaining a solution of the unrestricted least square method for high-speed calculation (see the following expression (7)), the solution obtained by correcting the solution to a non-negative value is set as the initial value of the iterative calculation, so that convergence to a good solution with a small number of iterations is possible. Therefore, it has been confirmed by experiment that the solution converges to the same accuracy with about 1 / 10 of the number of iterations.
[0169] [Formula 7]
[0170]
[0171] Fig.10 Schematically shown in a similar Fig. 9 The frequency scale conversion processing flow in which the solution of the least square method without non-negative value constraints is used instead of the random number in the iterative method. First, the initialization unit 1001 initializes the spectrum graph using the solution of the least square method without restrictions, and at this time, the initial value correction unit 1002 substitutes 0 for the variable with a negative value. Then, the update unit 1003 updates the intensity (amplitude) of each point on the time axis and the frequency axis by the gradient method, and then the correction unit 1004 substitutes 0 for the variable with a negative value. The processing of the update unit 1003 and the correction unit 1004 is repeated until the calculation result converges.
[0172] according to Fig.10The frequency scale conversion method shown in , the solution obtained by replacing negative values with 0 in the solution of the least square method is set as the initial value of the iterative calculation, so that it can converge to a satisfactory level through a small number of iterations. As a result, the musical instrument sound generation system 300 can realize the generation of musical instrument sounds within the interactive time.
[0173] F. Operational form of client-server model
[0174] The musical instrument sound generating system 300 is installed on an information processing device including, for example, a computer, etc. The process of generating a learning model (generator) using the model of the present disclosure and the process of generating a musical instrument sound having a pitch reflecting the characteristics of an input sound using the learning model (generator) have a large computational load. Therefore, a client-server model is assumed as one operation mode of the musical instrument sound generating system 300.
[0175] Fig.11 The configuration of the musical instrument sound generation system 300 including the client server model 1100 is schematically illustrated. The client server model 1100 includes a server 1101 that provides a service for generating musical instrument sounds with pitches and one or more clients 1102 that request the generation of musical instrument sounds with pitches. The server 1101 and each client 1102 are interconnected via a network such as a wide area network (WAN), a local area network (LAN), or the Internet.
[0176] Client 1102 includes, for example, an information terminal (edge device) used by a user, such as a smart phone, a tablet computer, or a personal computer (PC). The user mentioned here is, for example, a general user who composes music or performs other musical activities using the unique instrument sound provided from server 1101. On the client 1102 side, for example, a GUI operation such as the selection of a wav format file of an input sound and the designation of a pitch is performed via a GUI screen. At this point, in the case of selecting multiple input sounds, the designation of the mixing ratio of each input sound is also included in the GUI operation. Then, client 1102 requests server 1101 to perform a process of generating a musical instrument sound with a pitch reflecting the characteristics of the input sound.
[0177] The server 1101 includes, for example, an information processing device such as a computer, and is equipped with the main components 301 to 304 of the musical instrument sound generation system 300. In response to a request from the client 1102, the server 1101 generates a musical instrument sound with a pitch reflecting the characteristics of the input sound using a learning model (generator), and returns the musical instrument sound with the pitch to the client 1102 as a request source, wherein the learning model is generated using the disclosed model with the specified input sound and pitch as instance conditions. In addition, in the case of requesting multiple input sounds from the client 102, the server 1101 generates a musical instrument sound with a pitch by using a feature vector obtained by combining the feature vectors of the respective input sounds at a mixing ratio specified by the client 1102, and returns the musical instrument sound to the client 1102.
[0178] Fig.12 An exemplary processing sequence performed by the client 1102 and the server 1101 is shown.
[0179] On the client 1102 side, the user specifies an input sound serving as a sound source for generating a musical instrument sound and a pitch of the musical instrument sound to be generated through GUI operation (SEQ 1201).
[0180] The input sound is specified in the form of specifying the file name of the corresponding wav format file from the preset. For example, a wav format file prepared in advance on the server 1101 side, a wav format file that can be specified and obtained by the server 1101 on the client 1102 side, and a wav format file that can be uploaded to the server 1101 by the client 1102 can also be specified as the preset of the input sound. In addition, the user can specify two or more input sounds, and in the case of specifying multiple input sounds, the user can further specify the mixing ratio of each input sound.
[0181] Then, the client 1102 sends a request for generating a musical instrument sound with a pitch to the server 1101 (SEQ1202). The request includes information of the input sound and pitch specified by the user. In the case where the user specifies multiple input sounds, the request also includes the mixing ratio of the corresponding input sounds.
[0182] When receiving the request from the client 1102 (SEQ 1203), the server 1101 first obtains the input sound specified by the request (SEQ 1204). The server 1101 can obtain the wav format file of the specified input sound from its own local disk or from an external accumulation device via the network. In addition, the server 1101 can obtain the wav format file uploaded from the client 1102.
[0183] Next, the server 1101 converts the audio waveform of the input sound into a mel-spectrogram using the spectrogram conversion unit 301 (SEQ 1205). In the case where the request from the client 1102 specifies multiple input sounds, the server 1101 converts the audio waveforms of all the input sounds into mel-spectrograms.
[0184] Next, the server 1101 extracts the timbre feature quantity h from the mel-spectrogram of the input sound using the timbre feature extraction unit 302 (SEQ 1206). In the case where the request from the client 1102 specifies a plurality of input sounds, the mel-spectrogram generated from each input sound is mixed at a specified mixing ratio to calculate the timbre feature quantity h according to the above expression (1).
[0185] Next, the server 1101 generates a mel-spectrogram of a musical instrument sound containing the pitch specified by the request from the client 1102 based on the timbre feature quantity h extracted from the mel-spectrogram of the input sound using the generation unit 303 (SEQ1207). Specifically, using the learning model (generator) generated by the disclosed model, the generation unit 303 generates a mel-spectrogram of a musical instrument sound having a pitch reflecting the characteristics of the input sound based on the random number generated by the musical instrument sound generation system 300 with the timbre feature quantity and the pitch of the input sound as an example condition.
[0186] Next, the server 1101 reconstructs an audio waveform from the generated mel-spectrogram using the spectrogram waveform inverse conversion unit 304 (SEQ 1208). The audio waveform is a musical instrument sound having a pitch reflecting the characteristics of the input sound, and the output musical instrument sound has a length of about one second or several seconds. It is output as MIDI data.
[0187] Then, the server 1101 returns the data of the generated musical instrument sound with pitch as the request source to the client 1102 (SEQ 1209). It should be noted that the server 1101 can return the information of the mel spectrogram before reconstruction and the reconstructed audio waveform. In addition, the server 1101 can stream the data of the musical instrument sound, or can send the file itself to the client 1102.
[0188] When receiving MIDI data of musical instrument sound with pitch from server 1101 (SEQ 1209), client 1102 can reproduce (listen to user) and store musical instrument sound according to user's GUI operation (SEQ 1210). In addition, the user can further request to generate the next musical instrument sound through GUI operation.
[0189] G. GUI Configuration and Operation Example
[0190] In this section G, a GUI screen and GUI operations for requesting the generation of musical instrument sounds with pitches on the client 1102 side and instructing processes such as reproducing and storing the generated musical instrument sounds with pitches will be described. In the case where the sound reproducing system 300 is installed as a client server model, the GUI operations described in section G are performed on the terminal of the client as described in section F. It is to be noted that in the case where the sound reproducing system 300 is installed on a single information processing device, the GUI operations described in section G are performed using the console of the information processing device.
[0191] G-1. First embodiment
[0192] Fig.13 1 shows a configuration example of a GUI screen for generating and reproducing musical instrument sounds with pitches according to the present disclosure. Note that Fig.13 The configuration of the GUI screen used in the case of combining two input sounds to generate a musical instrument sound with a pitch is shown. In the present disclosure, a musical instrument sound with a pitch can also be generated based on three or more input sounds. The configuration and operation of the GUI screen used in the case of specifying three input sounds to generate a musical instrument sound with a pitch will be described later.
[0193] Fig.13 The GUI screen 1300 shown includes a preset selection unit 1301 as an input / output field, a first input sound information display unit 1302, a second input sound information display unit 1303, a mixing ratio specifying unit 1304, a pitch information specifying unit 1305 and a generated instrument sound information presenting unit 1306.
[0194] The preset selection unit 1301 is a GUI component for selecting the sound source of the first input sound and the second input sound via a pull-down menu. The pull-down menu includes a list of presets (file names of wav format files) that can be selected as input sounds prepared in advance by the musical instrument sound generating system 300 (alternatively, the server 1101) (not shown). The wav format files selected by the user on the pull-down menu are sequentially designated as the first input sound and the second input sound.
[0195] The first input sound information display unit 1302 and the second input sound information display unit 1303 display the file names of the wav format files of the first input sound and the second input sound specified by the preset selection unit 1301. In addition, each of the first input sound information display unit 1302 and the second input sound information display unit 1303 has a drop-down menu for changing the input sound. Each of the drop-down menus of the first input sound information display unit 1302 and the second input sound information display unit 1303 is not a preset menu prepared in advance by the musical instrument sound generation system 300 (alternatively, the server 1101), but a list of file names of wav format files that can be independently selected as input sounds in the client (alternatively, the information processing device). The user can also select the first input sound and the second input sound using the corresponding drop-down menus of the first input sound information display unit 1302 and the second input sound information display unit 1303.
[0196] The first input sound information display unit 1302 and the second input sound information display unit 1303 respectively include play buttons 1302-1 and 1303-1 for instructing the reproduction of the wav format file designated as the input sound. Using these play buttons 1302-1 and 1303-1, the user can reproduce and listen to each wav format file designated as the input sound to confirm whether the sound source is the sound source desired by the user before requesting the generation of the instrument sound.
[0197] The mixing ratio designation unit 1304 is an input field for the user to designate the mixing ratio of the two input sounds set in the first input sound information display unit 1302 and the second input sound information display unit 1303. Fig.13 In the embodiment shown in , the mixing ratio designation unit 1304 includes radio buttons for selectively designating the mixing ratios of the second input sound to the first input sound of 0.0, 0.1, 0.2, ..., 0.8, 0.9, and 1.0. A mixing ratio closer to 0.0 may request the generation of more instrument sounds that capture the characteristics of the first input sound, and a mixing ratio closer to 1.0 may request the generation of more instrument sounds that capture the characteristics of the second input sound. A mixing ratio of 0.0 means that only a single sound of the first input sound is specified, and a mixing ratio of 1.0 means that only a single sound of the second input sound is specified, and instrument sound generation may be requested from the characteristics of the single sound.
[0198] exist Fig.13In the illustrated GUI screen configuration example, the pitch information designation unit 1305 is arranged at the bottom of the screen. The pitch information designation unit 1305 includes a design using a layout of piano keys (hereinafter, referred to as a "keyboard") 1305-1. The user can designate the pitch of the instrument sound to be generated by clicking or touching a key in the keyboard 1305-1. Since the text 1305-2 indicating the pitch to be generated (in Fig.13 In the embodiment shown in , the text “Generate Target MIDI Note / Pitch: 60” is displayed) is displayed near the top of the keyboard 1305-1, and the user can visually confirm the text.
[0199] A pair of plus and minus buttons 1305-3 are provided near the lower left end of the keyboard 1305-1. The user can indicate the up or down of the octave by selecting the "+" button and the "-" button. Thus, the pitch of 88 pitches from A-1 to C7 can be specified.
[0200] In addition, a toggle switch 1305-4 for switching between the two states of "low sound quality" on and off is arranged in the lower center portion of the keyboard 1305-1. When the toggle switch 1305-4 is used to toggle to the on state of "low sound quality", a musical instrument sound reflecting the characteristics of the input sound is generated with low sound quality. On the other hand, when the toggle switch 1305-4 is used to toggle to the off state of "low sound quality", a musical instrument sound reflecting the characteristics of the input sound is generated with high sound quality.
[0201] In addition, the pitch information designation unit 1305 includes an "update" button 1305-5 substantially at the center above the keyboard 1305-1. In the case where the user wishes to generate a musical instrument sound with the same characteristics having another pitch, after designating a pitch corresponding to another desired pitch from the keyboard 1305-1 by clicking or touching, the user can instruct to reproduce the same musical instrument sound with another pitch by selecting the update button 1305-5.
[0202] When the user selects one of the keys on the keyboard 1305-1 on the pitch information specifying unit 1305, a request for generating a musical instrument sound with a pitch is output to the server 1101 (optionally, a process for generating a musical instrument sound with a pitch is activated in the information processing device).
[0203] On the server 1101 side (alternatively, in the information processing device), the audio waveforms of the respective input sounds of the first input sound and the second input sound are converted into mel-spectrograms, timbre feature quantities are extracted from the respective mel-spectrograms and mixed at a specified mixing ratio, and then a musical instrument sound having a tone of specified sound quality (high sound quality or low sound quality) is generated using the timbre feature quantities and pitch information as instance conditions. On the other hand, on the client 1102 side, the musical instrument sound with pitch generated on the server 1101 side is streamed and reproduced (alternatively, it is downloaded and reproduced and output) (however, in the case where the musical instrument sound with pitch is generated inside the information processing device, the information processing device reproduces and outputs the generated sound).
[0204] The generated musical instrument sound information presentation unit 1306 includes a presentation domain 1306-1, which presents information about the generated musical instrument sound in pitch. The "information about the musical instrument sound with pitch" to be presented is not particularly limited. For example, information that visually expresses the characteristics of the audio waveform of the musical instrument sound, such as a mel spectrogram (or a frequency spectrogram), can be displayed in the presentation domain 1306-1 (described later). Of course, instead of the spectrogram, the audio waveform of the musical instrument sound can be displayed in the presentation domain 1306-1.
[0205] In addition, the generated instrument sound information presentation unit 1306 includes a play button 1306-2. The user can reproduce and listen to the generated instrument sound at the pitch by using the play button 1306-2, and check whether the pitch and instrument sound reflect the characteristics of the designated input sound as expected.
[0206] In the following, reference will be made to Figures 14 to 23 Described in Fig.13 An operation example on the GUI screen is shown in FIG.
[0207] Fig.14 14 shows a state where the file names used as the first input sound and the second input sound are sequentially specified from the list of file names of wav format files displayed on the pull-down menu 1401 of the preset selection unit 1301. The files specified in the pull-down menu 1401 are displayed on the first input sound information display unit 1302 and the second input sound information display unit 1303, respectively. Fig.14 In the example shown in , audio files “Input_audio#001.wav” and “Input_audio#002.wav” are specified on the pull-down menu 1401 .
[0208] Fig.15A state in which corresponding file names are displayed on the first input sound information display unit 1302 and the second input sound information display unit 1303 in response to designation of audio files "input_audio#001.wav" and "input_audio#002.wav" on the pull-down menu 1401 is shown. The user can visually confirm the combination of input sounds for generating musical instrument sounds with pitches from the file names displayed on the first input sound information display unit 1302 and the second input sound information display unit 1303. In addition, the user can individually reproduce each of the wav format files "input_audio#001.wav" and "input_audio#002.wav" designated as input sounds by using the play button 1302-1 of the first input sound information display unit 1302 and the play button 1303-1 of the second input sound information display unit 1303, and confirm the combination of input sounds for generating musical instrument sounds with pitches by listening.
[0209] Fig.16 The radio buttons of the mixing ratio designation unit 1304 are shown for designating the state of the mixing ratio of the audio waveforms of "input_audio#001.wav" and "input_audio#002.wav" designated as the first input sound and the second input sound, respectively. The mixing ratio designation unit 1304 includes radio buttons for selectively designating the mixing ratio of the second input sound relative to the first input sound of 0.0, 0.1, 0.2, ..., 0.8, 0.9, and 1.0. A mixing ratio that is closer to 0.0 may request the generation of an instrument sound that captures the characteristics of the first input sound, and a mixing ratio that is closer to 1.0 may request the generation of an instrument sound that captures the characteristics of the second input sound (described above). Fig.16 In the example shown, a mixing ratio of 0.3 is specified.
[0210] Fig.16 Also shown is a state in which the pitch of the musical instrument sound to be generated is specified using the pitch information specifying unit 1305. The pitch information specifying unit 1305 has a design using the layout of a piano keyboard, and a character representing the corresponding pitch is displayed on each key of the keyboard 1305-1. In addition, the up and down of the octave can be indicated using the plus / minus button 1305-3 arranged near the lower left end of the keyboard 1305-1. Therefore, the user can specify 88 splits from A-1 to C7 by combining the operation of the plus / minus button 1305-3 of the keyboard 1305-1 and the split information specifying unit 1305 (as described above). It should be noted that the toggle switch 1305-4 for turning on / off "low sound quality" is switched to "off".
[0211] When the user selects any key of the keyboard 1305-1 (in Fig.16 In the embodiment shown in , when the key "A" is set, a request for generating a musical instrument sound with a pitch is output to the server 1101 (alternatively, a process for generating a musical instrument sound with a pitch is activated in the information processing device). Then, the musical instrument sound with a pitch reflecting the characteristics of the input sound is reproduced in a relatively short time of about 1 second or several seconds generated on the server 1101 side (or generated inside the information processing device). Fig.17 The description is given of a state in which a mel-spectrogram of a musical instrument sound is displayed in a presentation field 1306-1 of a musical instrument sound information presentation unit 1306 generated according to reproduction of the musical instrument sound. The presentation field 1306-1 reflects a mel-spectrogram generated corresponding to a mixing ratio specified by the mixing ratio specifying unit 1304. When another radio button is selected in the mixing ratio specifying unit 1304, on / off display of the radio button is switched in the mixing ratio specifying unit 1304 (not shown), the mel-spectrogram in the presentation field 1306-1 is reflected in the musical instrument sound corresponding to the newly selected mixing ratio, and the user can listen to and confirm the musical instrument sound.
[0212] In the case where the user wants to search for a sample of another musical instrument sound, the user may select a preset again through the preset selection unit 1301 , or individually change the first input sound and the second input sound through the first input sound information display unit 1302 and the second input sound information display unit 1303 . Fig.18 1302 and 1303 respectively. Fig.18 In the example shown in , the first input sound is changed to the audio file "input_audio#101.wav" by the selection operation on the pull-down menu 1302-2 of the first input sound information display unit 1302. In addition, the second input sound is changed to "input_audio#203.wav" by the selection operation on the pull-down menu 1303-2 of the second input sound information display unit 1303. The operations related to the designation of the mixing ratio, the designation of the pitch information, and the reproduction of the generated instrument sound with the pitch after each input sound is individually changed are similar to the above description, so the description thereof is omitted here.
[0213] also, Fig.19An operation example is shown in which a musical instrument sound having another pitch is generated while the user maintains the combination of input sounds. The user again specifies the pitch of the musical instrument sound that is desired to be newly generated with the same combination of input sounds by operating the keyboard 1305-1 and the plus / minus button 1305-3 of the pitch information designation unit 1305, and then selects the update button 1305-1 approximately in the center above the keyboard 1305-1 to instruct the reproduction of the same musical instrument sound with another pitch.
[0214] It should be noted that Fig.18 An operation example of specifying or changing the input sound through the pull-down menus of the first input sound information display unit 1302 and the second input sound information display unit 1303 is shown. A list of file names of preset wav format files is displayed on the pull-down menu. On the other hand, although not preset, a sound source at the user's hand can be selected as an input sound. Fig. 20 An operation embodiment is shown in which a sound source at the user's hand is designated as the first input sound. The "sound source at the user's hand" mentioned here is a wav format file that can be obtained from the local disk of the client 1102 (alternatively, an information processing device) or from an external accumulation device via a network. Since the sound of the instrument to generate the pitch is a relatively short time of about 1 second or several seconds, the sound source is also preferably audio data of a relatively short time of about 1 second or several seconds. Fig. 20 In the embodiment described in , an input box 1302-3 for selecting a sound source at the user's hand appears in the first input sound information display unit 1302. The input box 1302-3 displays a list of sound sources (wav format files) at the user's hand in the left half, and displays the attribute information of the currently selected sound source (highlight file) in the right half. The user can use the left half of the input box 1302-3 to search for the sound source at hand, and check whether the input sound is the expected input sound based on the attribute information and the reproduced sound displayed in the right half. Then, when the selection of any wav format file is confirmed on the input box 1302-3, the input box 1302-3 disappears, and the file name of the wav format file whose selection is confirmed is displayed on the first input sound information display unit 1302 (not shown).
[0215] In the case where the user likes the sound of a musical instrument having a pitch generated on the server 1101 side, the user can download a wav format file as a sound source to the client 1102 (for example, the user's own information terminal). Fig.21An operation example is shown when downloading a musical instrument sound having a pitch generated on the server 1101 side. A mel-spectrogram of the musical instrument sound to be downloaded is displayed in the presentation field 1306-1 of the generated musical instrument sound information presentation unit 1306. The user can instruct to download a musical instrument sound having a favorite pitch by selecting a download button 1306-4 arranged near the lower right end of the presentation field 1306-1.
[0216] In addition, a "download file for 12 pitches" button 1305-5 is arranged substantially at the center below the keyboard 1305-1 of the pitch information designation unit 1305. The "download file for 12 pitches" button 1305-5 is a button for instructing to download not only one specific pitch designated by the key in the keyboard 1305-1 but also the instrument sounds for 12 pitches. When the "download file for 12 pitches" button 1305-5 is selected, the server 1101 side generates the instrument sounds with 12 pitches and downloads the instrument sounds to the requesting client 1102. However, it takes processing time to generate the instrument sounds for 12 pitches.
[0217] Fig. 22 The state in which the "Interpolate All" button 1306-3 arranged substantially at the center above the generated musical instrument sound information presentation unit 1306 is selected is shown. The "Interpolate All" button 1306-3 is a button that indicates combining the first input sound and the second input sound to generate musical instrument sounds having pitches in the order from a mixing ratio of 0.1 to 0.9 (or from 0.0 to 1.0, including a single sound). When the "Interpolate All" button 1306-3 is selected, musical instrument sounds having pitches are automatically generated in sequence according to each mixing ratio on the server 1101 side, and the automatically generated musical instrument sounds are reproduced in sequence on the client 1102 side (alternatively, the automatic generation process of the musical instrument sounds having pitches at each mixing ratio is activated in the information processing device, and the sequentially generated musical instrument sounds are reproduced and output).
[0218] Fig.23An operation example in the case of performing envelope processing is shown. The envelope is a process that gives a typical change over time to the generated instrument sound. When the envelope button 1306-5, which is generally placed at the center below the generated instrument sound information presentation unit 1306, is selected, the presentation field 1306-1 switches from the display of the Mel spectrum graph to the display of the slider for adjusting each parameter of the envelope. The envelope parameters include ADSR (attack, decay, sustain, release). The shorter the attack time, the better the response, and the longer the attack time, the softer the rise. The longer the decay time, the longer the decay occurs, and the shorter the decay time, the shorter the decay occurs. The sustain level is a parameter for controlling volume rather than time, and controls the final volume achieved by continuing to open the note. The longer the release time, the longer the resonance occurs, and the shorter the release time, the clearer the sound. Note that the toggle switch 1305-4 is as described above. When the toggle switch 1305-4 is toggled to the off state of "low sound quality" (in other words, a state of high sound quality), a musical instrument sound reflecting the characteristics of the input sound is generated with high quality and customized by an envelope or the like.
[0219] G-2. Second embodiment
[0220] Fig.24 Another configuration embodiment of a GUI screen for generating and reproducing musical instrument sounds with pitches according to the present disclosure is shown. Fig.13 1 shows a GUI screen used in the case of combining two input sounds to generate a musical instrument sound, but Fig.24 A GUI screen used in the case of combining three input sounds to generate a musical instrument sound having a pitch is shown.
[0221] Fig.24 The illustrated GUI screen 2400 includes a sound source designation field 2410 , a musical instrument sound generation operation field 2420 , and a pre-processing operation field 2430 .
[0222] The sound source designation field 2410 is an operation area for selecting each sound source of the first to third input sounds. A wav format file used as the sound source of each input sound can be selected from presets using a pull-down menu, or a sound source at the user's hand can be selected using an input box (for example, see Fig. 20 ). The sound source selection operation using the pull-down menu and input box is as described above, and its detailed description will be omitted here. Fig.24 In the illustrated embodiment, three wav format files of “first input sound.wav”, “second input sound.wav” and “third input sound.wav” are selected through user operation, and the file name and creation date and time of each of these files are displayed in the sound source designation field 2410.
[0223] The musical instrument sound generation operation domain 2420 is an operation area for making settings when synthesizing the first to third input sounds selected in the sound source designation domain 2410. The musical instrument sound generation operation domain 2420 has a triangular background. The first to third input sounds are assigned to the corresponding vertices 2421 to 2423 of the triangle, and the sampled waveforms of the corresponding input sounds are displayed near the corresponding vertices 2421 to 2423.
[0224] The preprocessing operation domain 2430 is an operation area for performing preprocessing on the sample waveform of each input sound. When any one of the sample waveforms of the first input sound to the third input sound is selected in the musical instrument sound generation operation domain 2420, an operation screen of an equalizer (EQ) and an envelope processing (ADSR) designated for the sound source of the selected input sound are displayed in the preprocessing operation domain 2430, and preprocessing of EQ and ADSR can be performed on the sound source.
[0225] The musical instrument sound generation operation domain 2420 will be described again. When any position 2424 in the background triangle is selected, the mixing ratio of the first input sound to the third input sound is set based on the ratio of the distance between the selected position 2424 and each vertex 2421 to 2423. For example, when a position closer to the vertex 2421 to which the first input sound is assigned in the triangle is selected, the mixing ratio of the first input sound becomes higher, and generation of more musical instrument sounds that capture the characteristics of the first input sound can be requested. In addition, when any vertex position of the triangle is selected, this means that only a single sound of the input sound assigned to the vertex among the first to third vertices is specified, and musical instrument sound generation can be requested from the characteristics of the single sound.
[0226] A semi-transparent circle is first displayed at the position selected in the triangle. Thereafter, when the "Generate" button 2425 above the musical instrument sound generation operation field 2420 is selected, a musical instrument sound generation request is output to the server 1101 (optionally, a process of generating musical instrument sounds is activated in the information processing device). In addition, by selecting the "Generate" button 2425, the mixing ratio of the first input sound to the third input sound specified in the ○ position is determined, and the display of ○ changes from semi-transparent to opaque (not shown).
[0227] exist Fig.24 In the GUI screen 2400 shown, it is assumed that the instrument sound generation request requests the generation of an instrument sound of 88 keys. Of course, the request may be a request for generating an instrument sound having 89 or more keys or 87 or less keys. Alternatively, a GUI component (e.g., a keyboard) for specifying pitch information may be arranged in the instrument sound generation operation domain 2420 or in another domain, and a request for generating an instrument sound having the pitch of only one key specified by the GUI component may be made.
[0228] On the server 1101 side (alternatively, in the information processing device), the audio waveform of each input sound of the first to third input sounds is converted into a mel-spectrogram, the timbre feature quantity is extracted from each mel-spectrogram and mixed at a specified mixing ratio, and then the timbre feature quantity and each pitch information of 88 keys are used as instance conditions to generate a musical instrument sound with a tone of specified sound quality (high sound quality or low sound quality). On the other hand, on the client 1102 side, the musical instrument sound corresponding to the 88 keys generated on the server 1101 side is streamed and reproduced (alternatively, it is downloaded and reproduced and output) (however, in the case where the musical instrument sound with the pitch is generated inside the information processing device, the information processing device reproduces and outputs the generated sound).
[0229] Note that when the parameters of any input sound are changed in the pre-processing operation field 2430, the "Generate" button 2425 is highlighted and the user is visually warned that the parameter changes will not be reflected in the instrument sound unless the "Generate" button 2425 is selected to request generation of the instrument sound.
[0230] The user listens to the instrument sounds of 88 keys generated on the server 1101 side, and if the user likes the instrument sound, the user can select the favorite button 2426 and register the instrument sound in the favorite list. Registering to the favorite list may include recording an access method (e.g., a uniform resource identifier (URI) for distinguishing the location of a wav format file stored on the server 1101 side, etc.) to the sound source of the corresponding instrument sound and downloading the wav format file of the instrument sound from the server 1101.
[0231] Fig.25 Show Fig.24 A variation of the GUI screen shown. Fig.25 The GUI screen 2500 shown in FIG. Fig.24 The GUI screen shown in is different in that the musical instrument sound generation operation field 2420 has two triangles 2510 and 2520 arranged in the horizontal direction in the background.
[0232] The triangle 2510 on the left is used to display the sampling waveforms of the first to third input sounds and set the mixing ratio, similar to Fig.24 The triangle in the musical instrument sound generation operation field 2420 of the illustrated GUI screen 2400. Here, a detailed description of the left triangle 2510 is omitted.
[0233] On the other hand, when a generator generated by the disclosed model is used to generate a musical instrument sound under an example condition, a right triangle 2520 is used to fine-tune a random number z input to the generator. Also in the right triangle 2520, the first input sound to the third input sound are assigned to the vertices of the triangle 2520. When any position 2521 is selected in the triangle, the random number z input to the generator is fine-tuned based on the ratio of the distance between the selected position 2521 and each vertex. For example, when a position closer to the vertex in the triangle to which the first input sound is assigned is selected, the random number z is finely adjusted so as to include more features of the first input sound.
[0234] H. Configuration of information processing equipment
[0235] In this Section H, a specific configuration of an information processing device provided for implementing the present disclosure will be described.
[0236] Fig.26 A specific hardware configuration example of the information processing device 2000 is shown. Fig.26 The information processing device 2000 shown in FIG. 1 includes, for example, a PC or the like. The information processing device 2000 can be operated as Figure 3 The musical instrument sound generating system 300 shown in FIG. 1 may be operated as Fig.11 The server 1101 or the client 1102 is shown in FIG.
[0237] Fig.26 The information processing device 2000 shown in the figure includes a CPU 2001, a read-only memory (ROM) 2002, a random access memory (RAM) 2003, a host bus 2004, a bridge 2005, an expansion bus 2006, an interface unit 2007, an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013.
[0238] The CPU 2001 functions as an operation processing device and a control device, and controls the overall operation of the information processing apparatus 2000 according to various programs. The ROM 2002 stores programs (basic input / output system, etc.) and calculation parameters used by the CPU 2001 in a nonvolatile manner. The RAM 2003 is used to load programs used in the execution of the CPU 2001, and temporarily stores parameters such as work data that are appropriately changed in the execution of the program. The programs loaded into the RAM 2003 and executed by the CPU 2001 are, for example, various application programs, an operating system (OS), and the like.
[0239] The CPU 2001, the ROM 2002, and the RAM 2003 are connected to each other via a host bus 2004 including a CPU bus or the like. Then, the CPU 2001 can realize various functions and services by executing various application programs under an execution environment provided by the OS through cooperative operations of the ROM 2002 and the RAM 2003. In the case where the information processing device 2000 is a PC, the OS is, for example, Windows or Unix of Microsoft Corporation. In addition, the application program includes an application program that performs processing as each of the waveform-spectrogram conversion unit 301, the timbre feature extraction unit 302, the generation unit 303, and the spectrogram-waveform inverse conversion unit 304, an application program that performs learning processing of a machine learning model (DNN, etc.) used in each of the timbre feature extraction unit 302 and the generation unit 303, and an application program that processes user operations through a GUI screen, such as Figures 13 to 25 As described in .
[0240] The host bus 2004 is connected to the expansion bus 2006 via the bridge 2005. For example, the expansion bus 2006 is a peripheral component interconnect (PCI) bus or PCI Express, and the bridge 2005 is based on the PCI standard. However, the information processing device 2000 does not necessarily have a configuration in which circuit components are separated by the host bus 2004, the bridge 2005, and the expansion bus 2006, and an implementation in which almost all circuit components are interconnected by a single bus (not shown) may be adopted.
[0241] The interface unit 2007 connects peripheral devices such as the input unit 2008, the output unit 2009, the storage unit 2010, the drive 2011, and the communication unit 2013 according to the standard of the expansion bus 2006. Fig.26 All peripheral devices shown in are necessary, and the information processing device 2000 may further include peripheral devices (not shown). In addition, the peripheral devices may be built in the main body of the information processing device 2000, or some peripheral devices may be externally connected to the main body of the information processing device 2000.
[0242] The input unit 2008 includes an input control circuit, etc., which generates an input signal based on an input from a user and outputs the input signal to the CPU 2001. The input unit 2008 may include an input device such as a keyboard, a mouse, a touch panel, or a microphone. The output unit 2009 includes, for example, a display device such as a liquid crystal display (LCD) device, an organic electroluminescent (EL) display device, and a light emitting diode (LED). At least a part of the devices of the input unit 2008 and the output unit 2009 is used to perform GUI operations to specify input sounds and pitches and to instruct the generation of musical instrument sounds with pitches, or to instruct the reproduction, downloading, etc. of the generated musical instrument sounds with pitches.
[0243] The storage unit 2010 includes, for example, a large-capacity storage device such as a solid-state drive (SSD) or a hard disk drive (HDD), but may include an external storage device. The storage unit 2010 stores files such as programs (applications, OS, etc.) executed by the CPU 2001 and various data. In addition, the storage unit 2010 stores a wav format file of an audio waveform as a sound source of a musical instrument sound, and stores MIDI data of the generated musical instrument sound.
[0244] The removable recording medium 2012 is a cartridge storage medium such as a micro SD card. The driver 2011 performs read and write operations on the loaded removable recording medium 2012. The driver 2011 outputs data read from the removable recording medium 2012 to the RAM 2003 and the storage unit 2010, and writes data on the RAM 2003 and the storage unit 2010 to the removable recording medium 2012. The removable recording medium 2012 is used to read a wav format file of an audio waveform as a sound source of a musical instrument sound, and is used to store MIDI data of the generated musical instrument sound.
[0245] The communication unit 2013 is a device that performs wireless communication such as Wi-Fi (registered trademark), Bluetooth (registered trademark) or a cellular communication network such as 4G or 5G. In the case where the information processing device 2000 operates as the server 1101, mutual communication between the clients 1102 is performed via the communication unit 2013. In addition, the communication unit 2013 may include a terminal such as a universal serial bus (USB) or a high-definition multimedia interface (HDMI (registered trademark)), and may further include a function of performing data communication with a USB device such as a scanner or printer, a display, etc.
[0246] The series of processes described in this specification can be performed by hardware, software, or a configuration combining hardware and software. In the case where the processes are performed by software, a program recorded with a sequence of processes related to the implementation of the present disclosure is installed and executed in a memory in dedicated hardware incorporated in a computer. It is also possible to install the program in a general-purpose computer capable of performing various types of processes and cause the computer to perform processes related to the embodiments of the present disclosure.
[0247] The program may be pre-stored in a recording medium provided in the computer, such as an HDD, SSD, or ROM as a recording medium. Alternatively, the program may be temporarily or permanently stored in a removable recording medium such as a floppy disk, a compact disk read-only memory (CD-ROM), a magneto-optical (MO) disk, a digital versatile disk (DVD), a Blu-ray disk (BD) (registered trademark), a magnetic disk, a universal serial bus (USB) memory, etc. By using such a removable recording medium, a program related to an implementation of the present disclosure may be provided as so-called package software.
[0248] In addition, the program can be transferred from the download site to the computer via a network such as a wide area network (WAN), represented by a cellular network, a local area network (LAN), or the Internet in a wireless or wired manner. In the computer, the program thus transferred can be received and installed in a mass storage device such as an HDD or SSD in the computer.
[0249] I. Comparison with Related Studies
[0250] This Section I describes a comparison of the present disclosure with other studies on the generation of instrumental sounds.
[0251] NSynth (see non-patent document 2) uses an automatic encoder based on Wavenet (see non-patent document 3) to generate waveforms of musical instrument sounds, but there is a problem of slow generation due to autoregressive sampling, and artifacts may appear in the generated sound. On the other hand, the present invention can generate musical instrument sounds that reflect the input sound within the interactive time.
[0252] GANSynth (see non-patent document 4) can improve the generation speed and sound quality by generating a spectrogram containing phase information using an image generation model, but since GANSynth is an unconditional generation model and does not accept input, it is difficult to search for the desired timbre in a complex latent space. On the other hand, since the present invention is a generation model with instance conditions, it can receive an input sound and search for a timbre reflecting the input sound in a complex latent space.
[0253] Industrial Applicability
[0254] The present disclosure has been described in detail with reference to specific embodiments. However, the present disclosure should not be construed as being limited to the above-mentioned embodiments, and it is apparent that, without departing from the gist of the present disclosure, those skilled in the art may modify and replace the embodiments. In addition, the effects described in this specification are only embodiments, and the effects brought about by the present disclosure are not limited, and there may be other effects not described in this specification.
[0255] The present disclosure can be applied to, for example, personal computers, electronic musical instruments, etc., which perform processing related to music generation such as compositions or music editing, and can generate unique instrument sounds from arbitrary sound inspirations to freely customize instruments or assign meanings to music through sounds.
[0256] In short, the present disclosure has been described in the form of embodiments so far, and the contents described in this specification should not be interpreted in a limiting manner. In order to determine the gist of the present disclosure, the claims should be considered.
[0257] It should be noted that the present disclosure may have the following configurations.
[0258] (1) An information processing system comprising:
[0259] The circuit system is configured as
[0260] Receive input sound and pitch information;
[0261] extracting a timbre feature quantity from the input sound; and
[0262] Information of musical instrument sounds having pitches is generated based on the timbre feature quantity and the pitch information.
[0263] (2) An information processing system according to (1), wherein:
[0264] The circuit system is configured to generate information of musical instrument sounds using the learned model.
[0265] (3) The information processing system according to any one of (1) to (2), wherein:
[0266] The circuit system is configured to generate information of a musical instrument sound using a learning model with information generated by preprocessing an input sound and pitch information as example conditions.
[0267] (4) The information processing system according to any one of (1) to (3), wherein:
[0268] The circuit system is configured to extract the timbre feature amount so that pitch information is not retained.
[0269] (5) The information processing system according to any one of (1) to (4), wherein:
[0270] The circuit system is configured to extract timbre features using a timbre feature extractor that has performed opponent learning on pitch.
[0271] (6) The information processing system according to any one of (1) to (5), wherein the circuit system is configured as follows:
[0272] Convert the input sound into a mel-spectrogram; and
[0273] The timbre feature quantity of the input sound is extracted based on the mel-spectrogram from the input sound.
[0274] (7) The information processing system according to (6), wherein the circuit system is configured as follows:
[0275] Generate a mel-spectrogram of a musical instrument sound having a pitch using the timbre feature quantity and the pitch information, and
[0276] Construct an audio waveform from a mel-spectrogram.
[0277] (8) The information processing system according to (7), wherein the circuit system is configured as follows:
[0278] Convert the Mel-spectrogram to the frequency scale of the linear spectrogram;
[0279] restoring the phase of the linear spectrogram; and
[0280] After restoring the phase of the linear spectrogram, an inverse Fourier transform is performed on the linear spectrogram.
[0281] (9) The information processing device according to (8), wherein the circuit system is configured as follows:
[0282] The solution corrected to a non-negative value is set to the solution of the least square method without a non-negative value as the initial value of the iterative calculation; and
[0283] The frequency scale conversion is performed according to an iterative method that is repeatedly updated according to a gradient method and corrected to a non-negative value.
[0284] (10) The information processing system according to any one of (1) to (9), wherein the circuit system is configured as follows:
[0285] receiving input of a plurality of input sounds;
[0286] extracting a timbre feature quantity from each input sound; and
[0287] The musical instrument sound information is generated based on the tone color feature amount obtained by mixing the tone color feature amounts of a plurality of input sounds and the pitch information.
[0288] (11) The information processing system according to (10), wherein the circuit system is configured as follows:
[0289] receiving information about a mixing ratio of a plurality of input sounds; and
[0290] The musical instrument sound information is generated based on a tone color feature amount obtained by mixing a plurality of tone color feature amounts, a mixing ratio, and pitch information.
[0291] (12) The information processing system according to any one of (1) to (11), wherein:
[0292] The circuit system is configured to receive input sound and pitch information based on user operation.
[0293] (13) The information processing system according to any one of (1) to (12), wherein:
[0294] The circuit system is configured to output information of musical instrument sounds.
[0295] (14) The information processing system according to any one of (1) to (13), wherein:
[0296] The circuit system is configured to display a user interface configured to receive user input corresponding to the input sound and pitch information.
[0297] (15) An information processing system according to (14), wherein:
[0298] The user interface is configured to receive a first input corresponding to a first input sound and a second input corresponding to a second input sound, and
[0299] The user interface is configured to receive a mixing ratio corresponding to the first input sound and the second input sound.
[0300] (16) An information processing system according to (15), wherein:
[0301] The graphical user interface includes at least a first graphic and a second graphic, wherein:
[0302] The first graphic is configured to receive a first input corresponding to a first input sound and a second input corresponding to a second input sound, and
[0303] The second graph is configured to receive a tone color feature amount.
[0304] (17) An information processing method comprising:
[0305] Receive input sound and pitch information;
[0306] extracting a timbre feature quantity from the input sound; and
[0307] Information of musical instrument sound having a pitch is generated based on the timbre feature quantity and the pitch information.
[0308] (18) one or more non-transitory computer-readable media that, when executed by a circuit system, causes the circuit system to:
[0309] Receive input sound and pitch information;
[0310] extracting a timbre feature quantity from the input sound; and
[0311] Information of musical instrument sound having a pitch is generated based on the timbre feature quantity and the pitch information.
[0312] (19) A sound generation system comprising:
[0313] a terminal configured to request generation of a musical instrument sound; and
[0314] An information processing device for generating musical instrument sounds, wherein:
[0315] The information terminal is configured as:
[0316] Receive input sound and pitch information;
[0317] extracting a timbre feature quantity from the input sound; and
[0318] Information of musical instrument sounds having pitches is generated based on the timbre feature quantity and the pitch information.
[0319] (20) An information terminal comprising:
[0320] a communication interface configured to communicate with an information processing system; and
[0321] A user interface configured to receive instructions related to a musical instrument sound including generating an input sound and pitch information, wherein
[0322] The communication interface is configured as:
[0323] sending a request for generating a musical instrument sound including an input sound and pitch information to an information processing system, and
[0324] Receive information about the sound of the musical instrument from the information processing system.
[0325] Reference Numbers List
[0326] 300 Sound Generation System
[0327] 301 Waveform Spectrum Conversion Unit
[0328] 302 Tone feature extraction unit
[0329] 303 Generation Unit
[0330] 304 Spectrum waveform inverse conversion unit
[0331] 501 Timbre Feature Extractor
[0332] 502 Instrument Identifier
[0333] 601 Timbre Feature Extractor
[0334] 602 Instrument Identifier
[0335] 603 Pitch Discriminator
[0336] 801 Frequency Scale Conversion Unit
[0337] 802 Phase Recovery Unit
[0338] 803 Inverse Short-Time Fourier Transform Unit (iSTFT)
[0339] 901 Initialization unit
[0340] 902 Update Unit
[0341] 903 Calibration Unit
[0342] 1001 Initialization unit
[0343] 1002 Initial value correction unit
[0344] 1003 Update Unit
[0345] 1004 Calibration Unit
[0346] 1100 Sound Generation System (Client Server Model)
[0347] 1101 Server
[0348] 1102 Client
[0349] 2000 Information Processing Equipment
[0350] 2001CPU
[0351] 2002ROM
[0352] 2003RAM
[0353] 2004 Host Bus
[0354] 2005 Bridge
[0355] 2006 Expansion Bus
[0356] 2007 Interface Unit
[0357] 2008 Input Unit
[0358] 2009 Output Unit
[0359] 2010 Storage Unit
[0360] 2011 Driver
[0361] 2012 Removable Recording Media
[0362] 2013 Communications Unit.
Claims
1. An information processing system, comprising: A circuit system configured to receive input sound and pitch information; extracting a timbre feature quantity from the input sound; and Information of a musical instrument sound having a pitch is generated based on the timbre feature amount and the pitch information.
2. The information processing system according to claim 1, wherein: The circuit system is configured to generate information of the musical instrument sound using a learning model.
3. The information processing system according to claim 1, wherein: The circuit system is configured to generate the information of the musical instrument sound using a learning model with the information generated by pre-processing the input sound and the pitch information as instance conditions.
4. The information processing system according to claim 1, wherein: The circuit system is configured to extract the timbre feature amount so that pitch information is not retained.
5. The information processing system according to claim 1, wherein: The circuit system is configured to extract the timbre feature using a timbre feature extractor that has performed opponent learning on pitch.
6. The information processing system according to claim 1, wherein: The circuit system is configured to: Converting the input sound into a mel-spectrogram; and A timbre feature quantity of the input sound is extracted based on a mel-spectrogram from the input sound.
7. The information processing system according to claim 6, wherein: The circuit system is configured to: generating a mel-spectrogram of a musical instrument sound having a pitch using the timbre feature quantity and the pitch information, and An audio waveform is constructed based on the mel-spectrogram.
8. The information processing system according to claim 7, wherein: The circuit system is configured to: Converting the Mel-spectrogram into a frequency scale in a linear spectrogram; restoring the phase of the linear frequency spectrum; and After restoring the phase of the linear spectrogram, an inverse Fourier transform is performed on the linear spectrogram.
9. The information processing device according to claim 8, wherein: The circuit system is configured to: The solution corrected to a non-negative value is set to the solution of the least square method without a non-negative value as the initial value of the iterative calculation; and The frequency scale conversion is performed according to an iterative method which is repeatedly updated according to a gradient method and corrected to a non-negative value.
10. The information processing system according to claim 1, wherein: The circuit system is configured to: receiving input of a plurality of input sounds; extracting a timbre feature quantity from each input sound; and The musical instrument sound information is generated based on the tone color feature amount obtained by mixing the tone color feature amounts of the plurality of input sounds and the pitch information.
11. The information processing system according to claim 10, wherein: The circuit system is configured to: receiving information about a mixing ratio of a plurality of input sounds; and The musical instrument sound information is generated based on a tone color feature amount obtained by mixing a plurality of tone color feature amounts, the mixing ratio, and the pitch information.
12. The information processing system according to claim 1, wherein: The circuit system is configured to receive the input sound and the pitch information based on a user operation.
13. The information processing system according to claim 1, wherein: The circuit system is configured to output information of the musical instrument sound.
14. The information processing system according to claim 1, wherein: The circuit system is configured to display a user interface configured to receive a user input corresponding to the input sound and the pitch information.
15. The information processing system according to claim 14, wherein: The user interface is configured to receive a first input corresponding to a first input sound and a second input corresponding to a second input sound, and The user interface is configured to receive a mixing ratio corresponding to the first input sound and the second input sound.
16. The information processing system according to claim 15, wherein: The graphical user interface includes at least a first graphic and a second graphic, wherein: The first graph is configured to receive the first input corresponding to the first input sound and the second input corresponding to the second input sound, and the second graph is configured to receive the tone color feature amount.
17. An information processing method, comprising: Receive input sound and pitch information; extracting a timbre feature quantity from the input sound; and Information of a musical instrument sound having a pitch is generated based on the timbre feature amount and the pitch information.
18. One or more non-transitory computer readable media, which, when executed by a circuit system, cause the circuit system to: Receive input sound and pitch information; extracting a timbre feature quantity from the input sound; and Information of a musical instrument sound having a pitch is generated based on the timbre feature amount and the pitch information.
19. A sound generation system comprising: a terminal configured to request generation of a musical instrument sound; as well as An information processing device for generating musical instrument sounds, wherein: The information terminal is configured as follows: Receive input sound and pitch information; extracting a timbre feature quantity from the input sound; and The information of the musical instrument sound having the pitch is generated based on the timbre feature amount and the pitch information.
20. An information terminal, comprising: a communication interface configured to communicate with an information processing system; as well as A user interface configured to receive instructions related to a musical instrument sound including generating an input sound and pitch information, wherein The communication interface is configured to send a request for generating the musical instrument sound including the input sound and the pitch information to the information processing system, and The information of the musical instrument sound is received from the information processing system.
Citation Information
Patent Citations
Musical sound emphasis device, convolution auto encoder learning device, musical sound emphasis method, and program
JP2019078864A
Panel manufacturing device, panel manufacturing method, and panel
JP2022164477A