Information processing apparatus and control method thereof

The information processing apparatus simplifies tone color adjustment for beginners by using natural language input and a learned model to present tone color candidates, addressing the complexity of conventional synthesizers.

JP7710657B2Active Publication Date: 2025-07-22YAMAHA CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021034735
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-04
Publication Date
2025-07-22
Estimated Expiration
2041-03-04

AI Technical Summary

Technical Problem

Conventional synthesizers are difficult for beginners to use due to the complexity of adjusting tone color using numerous buttons and knobs, making it hard to find and adjust waveform data and effect parameters.

Method used

An information processing apparatus that uses a natural language input module, a tone color estimation module, and a presentation module to facilitate easy tone color adjustment by allowing users to input adjectives, which are processed through a learned model to output tone color data, presenting candidates for selection.

Benefits of technology

Enables beginners to easily adjust tone color output by simplifying the process through natural language input and candidate presentation, reducing the need for complex button and knob operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710657000001
    Figure 0007710657000001
  • Figure 0007710657000002
    Figure 0007710657000002
  • Figure 0007710657000003
    Figure 0007710657000003
Patent Text Reader

Abstract

To provide an information processing device that enables even beginners to easily adjust tone to be output, and its control method.SOLUTION: An information processing system 100 includes: an operation unit 105 and a display unit 108 integrally configured as a touch panel display where natural language including adjectives is input by a user; and an estimation unit 203 that outputs tone data based on user-input natural language by using a learned model that outputs tone data from the adjectives, which is a function executed by a GPU 102.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus that adjusts a tone color output based on tone color data, and a control method thereof.

Background Art

[0002] Conventionally, a synthesizer capable of outputting a tone color adjusted using tone color data including waveform data and effect parameters has been known.

[0003] For example, Patent Document 1 discloses a music performance apparatus that outputs a sound with a pitch and a tone color corresponding to the coordinate position of an input means that contacts a display unit that performs two-axis display of pitch and tone color when the input means contacts the display unit.

[0004] Also, for example, Patent Document 2 discloses a tone color setting system that can automatically perform tone color setting suitable for a psychological state such as a user's mood or emotion based on the user's actual performance.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, even when using the technologies of Patent Documents 1 and 2, it is difficult to find waveform data of an instrument type that a beginner wants to use for performance or adjust the tone color with effect parameters by operating buttons and knobs provided in large numbers on conventional synthesizers.

[0007] In view of the above circumstances, an object of the present invention is to provide an information processing apparatus and a control method thereof that can easily adjust the tone color to be output even for beginners.

Means for Solving the Problems

[0008] To achieve the above object, an information processing apparatus according to an aspect of the present invention includes an input module for receiving a natural language including an adjective as user input, and based on the natural language input by the user using a learned model that outputs tone color data from the adjective A plurality of a tone color estimation module that outputs tone color data , a presentation module that presents the plurality of timbre data to a user as candidates for timbre data selected by the user; and is provided with.

Effects of the Invention

[0009] According to the present invention, even a beginner can easily adjust the tone color to be output.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. Each of the embodiments described below is merely an example of a configuration capable of realizing the present invention. Each of the following embodiments can be appropriately modified or changed according to the configuration of the apparatus to which the present invention is applied and various conditions. Also, not all combinations of elements included in the following embodiments are essential for realizing the present invention, and some of the elements can be appropriately omitted. Therefore, the scope of the present invention is not limited by the configurations described in the following embodiments. Also, a configuration combining a plurality of configurations described in the embodiments can be adopted as long as there is no contradiction between them.

[0012] The information processing apparatus 100 according to the present embodiment is realized by a synthesizer, but is not limited thereto. For example, the information processing apparatus 100 may be an information processing apparatus (computer) such as a personal computer or a server that transmits tone color data to be set to an external synthesizer.

[0013] Here, the tone color data in the present embodiment is data including at least one of waveform data of various musical instruments such as a piano, an organ, and a guitar, and effect parameters such as chorus, reverb, and distortion.

[0014] Roughly speaking, the information processing apparatus 100 in the present embodiment sets candidates for tone color data used for tone color adjustment based on the natural language input by the user when the user adjusts the tone color for playing on the information processing apparatus 100, and displays a list of each candidate in a state where the sample tone color can be reproduced. Then, when the user selects a candidate whose reproduced sample tone color is the tone color that the user wants to use for performance from among the candidates displayed in the list, the information processing apparatus 100 performs tone color adjustment so that the sample tone color becomes the tone color when playing on the information processing apparatus 100.

[0015] FIG. 1 is a block diagram showing the hardware configuration of the information processing apparatus 100 according to the embodiment of the present invention.

[0016] As shown in FIG. 1, the information processing apparatus 100 of the present embodiment includes a CPU 101, a GPU 102, a ROM 103, a RAM 104, an operation unit 105, a microphone 106, a speaker 107, a display unit 108, and an HDD 109, which are connected to each other via a bus 110. Although not shown in FIG. 1, the information processing apparatus 100 includes a keyboard that can be played by a user.

[0017] The CPU 101 is one or more processors that control each part of the information processing apparatus 100 by using the RAM 104 as a work memory according to a program stored in the ROM 103, for example.

[0018] Since the GPU 102 can perform efficient calculations by parallel processing of data, the process of performing learning using a learning model, as will be described later, is performed by the GPU 102.

[0019] The RAM 104 is a volatile memory and is used as a main memory of the CPU 101, a temporary storage area such as a work area, and the like.

[0020] The microphone 106 converts the collected sound into an electrical signal (voice data) and supplies it to the CPU 101. For example, the microphone 106 collects the voice consisting of natural language spoken by the user toward the microphone 106, and supplies the converted voice data to the CPU 101.

[0021] The speaker 107 emits a tone with adjusted timbre during performance using the information processing apparatus 100, when executing step S402 in FIG. 4 described later, when executing step S509 in FIG. 5 described later, and the like.

[0022] The HDD 109 is a non-volatile memory, and tone data, other data, various programs for the CPU 101 to operate, and the like are stored in respective predetermined areas. Note that the HDD 109 may be any non-volatile memory that can store the above data and programs, and may be, for example, another memory such as a flash memory.

[0023] The operation unit 105 and the display unit 108 are integrally configured as a touch panel display that receives user operations on the information processing apparatus 100 and displays various types of information. However, the operation unit 105 and the display unit 108 may be independent user interfaces. For example, the operation unit 105 may be composed of a keyboard or a mouse, and the display unit 108 may be composed of a display.

[0024] The bus 110 is a signal transmission path that connects the above-described hardware elements of the information processing apparatus 100 to each other.

[0025] FIG. 2 is a block diagram showing the functional configuration of the information processing apparatus 100.

[0026] In FIG. 2, the information processing apparatus 100 includes a learning unit 201, an input unit 202, an estimation unit 203, and an output unit 204.

[0027] The input unit 202 (input module) is a function executed by the CPU 101 that outputs an adjective input by the user to the estimation unit 203.

[0028] Specifically, the input unit 202 displays the I / F 601 (FIG. 6(a)) on the display unit 108, and acquires the natural language input in the I / F 601 by the user using the operation unit 105. Thereafter, the input unit 202 performs morphological analysis of the acquired natural language, extracts the adjective input by the user, and outputs the extracted adjective to the estimation unit 203.

[0029] Note that the input unit 202 is not limited to this embodiment as long as it can acquire the adjective input by the user. For example, the adjective input by the user may be acquired based on the natural language spoken by the user collected by the microphone 106, or the I / F 602 (FIG. 6(b)) including tags of a plurality of adjectives may be displayed on the display unit 108, and the adjective of the tag selected by the user using the operation unit 105 may be acquired as the adjective input by the user.

[0030] Details of the processing by the input unit 202 will be described later with reference to FIG. 4.

[0031] The learning unit 201 is a function executed by the GPU 102, which is composed of a learning model constituted by a CVAE (conditional variational auto encoder), which is a type of neural network. The GPU 102 trains the learning model that constitutes the learning unit 201 by supervised learning using training data consisting of effect parameters and adjectives tagged thereto, and outputs the parameters of the decoder of the generated learned model to the estimation unit 203.

[0032] The learning model that constitutes the learning unit 201 has an encoder and a decoder. Here, the encoder is a neural network that extracts a latent variable z tagged with an adjective (label y) in the latent space from the training data when an effect parameter (input data x) tagged with an adjective (label y) is input as training data. Also, the decoder is a neural network that reconstructs an effect parameter (output data x̃) tagged with an adjective (label y) when a latent variable z tagged with an adjective (label y) is input. The GPU 102 compares the input data x and the output data x̃ and adjusts the parameters of the encoder and decoder that constitute the learning unit 201. Also, for each label y, the parameters of the encoder are adjusted so that clusters by the latent variable z in the latent space shown in FIG. 3 are formed. The GPU 102 repeats such processing, optimizes the parameters of the learning model that constitutes the learning unit 201, trains the learning model, and generates a learned model. The details of the training process of the learning model by the GPU 102 will be described later with reference to FIG. 4.

[0033] The estimation unit 203 (voice color estimation module) is the same neural network as the decoder of the learned model generated in the learning unit 201 (hereinafter simply referred to as the decoder), which is a function executed by the GPU 102.

[0034] When the parameters are output from the learning unit 201 to the estimation unit 203, the GPU 102 updates the parameters of the decoder that constitutes the estimation unit 203 with those parameters.

[0035] Also, when the adjectives input by the user are output from the input unit 202 to the estimation unit 203, the GPU 102 acquires the latent variable z in the latent space shown in FIG. 3, which is tagged with the adjective, and inputs this to the decoder that constitutes the estimation unit 203, thereby reconstructing (estimating) the effect parameters (tone color data) tagged with the adjective. Then, the GPU 102 outputs the reconstructed effect parameters to the output unit 204. The details of the estimation process of the tone color data by the GPU 102 will be described later with reference to FIG. 5.

[0036] Note that the neural networks used in the learning unit 201 and the estimation unit 203 are not particularly limited, and examples thereof include DNN, RNN / LSTM, Recurrent Neural Network, and CNN (Convolutional Neural Network). Also, instead of the neural network, other models such as HMM (hidden Markov model) and SVM (support vector machine) may be used.

[0037] Also, although the learning unit 201 is configured only with a CVAE for performing supervised learning, it may be configured to include a VAE (variational auto encoder) or a GAN (Generative Adversarial Networks). In this case, in the learning unit 201, semi-supervised learning is executed, in which unsupervised learning by a VAE or a GAN, that is, learning using clustering with the effect parameters not tagged with adjectives as training data, is combined with supervised learning by a CVAE.

[0038] Also, the learning unit 201 and the estimation unit 203 may be a single device (system).

[0039] Furthermore, although the learning unit 201 and the estimation unit 203 are executed by the GPU 102, which is a single processor in this embodiment, the GPU 102 may be composed of a plurality of processors to perform distributed processing. Also, it may be a function executed in cooperation with the CPU 101, not just the GPU 102.

[0040] The output unit 204 (presentation module) is a function executed by the CPU 101 that lists (presents) a plurality of effect parameters output from the estimation unit 203 as candidates for effect parameters used for tone color adjustment when the user plays using the information processing apparatus 100.

[0041] Specifically, the output unit 204 displays on the display unit 108 an I / F 603 (FIG. 6(c)) including a plurality of tabs associated with each candidate effect parameter. As shown in FIG. 6(c), each tab of the I / F 603 is provided with a playback button associated with a sample sound when the tone color is adjusted by each effect parameter. After that, when one of the playback buttons on the I / F 603 is pressed by the user, the output unit 204 sets the tab provided with the playback button to the user-selected state and plays the sample tone color associated with the playback button. The user presses each playback button displayed on the I / F 603 and presses the determination button 604 when the desired sample tone color is played. When the determination button 604 is pressed, the output unit 204 determines to use the effect parameter associated with the currently user-selected tab for tone color adjustment of the information processing apparatus 100.

[0042] Details of the processing by the output unit 204 will be described later with reference to FIG. 5.

[0043] FIG. 3 is a diagram showing a state in which each effect parameter included in the collected training data is mapped onto a latent space.

[0044] When a learned model is generated in the learning unit 201 by the GPU 102, the effect parameter (input data x) is mapped as a latent variable z in the latent space. Many of these latent variables z are included in one of the clusters formed for each label y. In the present embodiment, as shown in FIG. 3, in the latent space, there are formed a cluster 301 of the adjective "beautiful", which is one of the labels y tagged to the input data x, and a cluster 302 of the adjective "glittering", which is also one of the labels y, etc.

[0045] In the present embodiment, the case where the input data x to the learning unit 201 is only the effect parameter has been described, but it is not limited to this if it is sound color data. For example, the input data x to the learning unit 201 may be sound color data consisting of only waveform data, a combination of waveform data and effect parameters, or a sound color data set including a plurality of sound color data.

[0046] FIG. 4 is a flowchart showing the training process of the learning model in the present embodiment.

[0047] This process is executed by the CPU 101 reading out the program stored in the ROM 103 and using the RAM 104 as a working memory.

[0048] First, in step S401, the CPU 101 acquires the effect parameter from the HDD 109. Note that the effect parameter may be acquired from the outside via a communication unit (not shown) in FIG. 1.

[0049] In step S402, the CPU 101 acquires the adjective to be tagged for each of the effect parameters collected in step S401.

[0050] Here, the adjective to be tagged is specifically acquired as follows.

[0051] First, the CPU 101 adjusts the timbre of the waveform data of the piano, which is the default waveform data, using each of the collected effect parameters, causes the adjusted timbre to be emitted from the speaker 107, and causes the display unit 108 to display the I / F 601 (FIG. 6(a)).

[0052] Thereafter, when the CPU 101 detects that the operator has input a character to the I / F 601 using the operation unit 105 for an adjective recalled from the timbre emitted from the speaker 107, the CPU 101 acquires the input character as an adjective to be tagged. The adjective acquired here may be singular or plural.

[0053] Note that since the adjectives to be tagged are acquired by the above method, in view of the common general knowledge in the art at the time of filing, the correlation between the timbre data included in the training data and the adjectives tagged thereto is inferred.

[0054] In step S403, the CPU 101 tags the adjectives acquired in step S402 to the effect parameters acquired in step S401, and generates training data. Note that a data set including such effect parameters and the adjectives tagged thereto may be obtained using crowdsourcing.

[0055] In step S404, the CPU 101 inputs the training data generated in step S403 to the learning unit 201, causes the GPU 102 to learn the learning model constituting the learning unit 201, and generates a learned model. Thereafter, the GPU 102 outputs the parameters of the decoder of the learned model from the learning unit 201 to the estimation unit 203, updates the parameters of the decoder constituting the estimation unit 203, and then ends this process.

[0056] Furthermore, in the present embodiment, the timbre to be emitted from the speaker 107 in step S402 was the one obtained by adjusting the timbre of the waveform data of the piano, but the timbre of the waveform data of a plurality of instrument types may be adjusted. In this case, for the same effect parameter, adjectives tagged for each instrument type are acquired in step S402. Also, in step S404, the learned model is generated for each instrument type.

[0057] Next, the estimation process of the timbre data in the present embodiment, which is executed after the process of FIG. 4, will be described with reference to FIG. 5.

[0058] FIG. 5 is a flowchart showing the estimation process of the timbre data in the present embodiment.

[0059] This process is executed by the CPU 101 reading out the program stored in the ROM 103 and using the RAM 104 as a working memory.

[0060] First, in step S501, the CPU 101 causes the display unit 108 to display the I / F 601, and acquires the natural language in which the user inputs characters to the I / F 601 using the operation unit 105. Thereafter, arbitrary morphological analysis is performed on the acquired natural language, and the adjectives input by the user are extracted.

[0061] For example, when a natural language such as "beautiful piano sound" is input as characters to the I / F 601, three words, "beautiful", "piano", and "sound", are acquired by morphological analysis of the input natural language, and the word "beautiful" is extracted as the adjective input by the user.

[0062] Also, when a natural language such as "glittering and beautiful piano sound" is input as characters to the I / F 601, two words, "glittering" and "beautiful", are extracted as the adjectives input by the user.

[0063] Also, in step S501, if the adjectives input by the user can be acquired, it is not limited to the method of this embodiment. For example, instead of displaying I / F601, I / F602 that displays a plurality of adjectives acquired in the process of step S402 as tags that can be selected by the user may be displayed, and the adjectives displayed in the tags selected by the user may be acquired as the adjectives input by the user. Further, instead of displaying I / F601, voice data including natural language spoken by the user with the microphone 106 may be converted into text data using any voice recognition technology, any morphological analysis may be performed on the text data, and the adjectives input by the user may be extracted.

[0064] Next, in step S502, the CPU 101 acquires from the latent space the latent variables tagged with the adjectives extracted in step S501, and inputs the latent variables tagged with the adjectives to the decoder that constitutes the estimator 203. Thereby, the GPU 102 is caused to output the effect parameters tagged with the adjectives from the decoder that constitutes the estimator 203. If there are a plurality of adjectives extracted in step S501, all of those adjectives are input to the decoder that constitutes the estimator 203.

[0065] For example, when the adjective "beautiful" is extracted in step S501, the effect parameters tagged with the adjective "beautiful", which are reconstructed by the latent variable z such as the latent variable that forms the cluster 301 shown in FIG. 3 in the latent space, are output from the estimator 203.

[0066] Also, for example, when the adjectives "beautiful" and "glittering" are extracted in step S501, the effect parameters tagged with these two adjectives, which are reconstructed by the latent variable z such as the latent variable that forms the cluster 301 shown in FIG. 3 in the latent space and with which these two adjectives are tagged, are output from the estimator 203.

[0067] If the learned model has been generated for each musical instrument type in step S404 and not only adjectives but also musical instrument types have been extracted in step S501, the adjectives extracted in step S501 are input to the decoder for the musical instrument type extracted in the estimation unit 203.

[0068] In step S503, the CPU 101 sets candidates for effect parameters used by the user for tone color adjustment from among the plurality of effect parameters output in step S502. In the present embodiment, those randomly specified from among the plurality of effect parameters output in step S502 are set as candidates for the effect parameters used by the user for tone color adjustment. Incidentally, among the plurality of effect parameters output in step S502, those having a likelihood equal to or greater than a threshold value may be set as candidates for the effect parameters used by the user for tone color adjustment.

[0069] In step S504, the CPU 101 determines whether there is a user input of a musical instrument type. Specifically, if there is a musical instrument type among the words obtained by any morphological analysis in step S501, it is determined that there is a user input of a musical instrument type.

[0070] For example, when the natural language "beautiful piano sound" is input as text to the I / F 601 in step S501, in step S504, the CPU 101 determines that there is a user input of the musical instrument type "piano".

[0071] If there is a user input of a musical instrument type (YES in step S504), the process proceeds to step S505, and the CPU 101 acquires the waveform data of the musical instrument type input by the user from the HDD 109, and proceeds to step S507.

[0072] Still, in this case, the CPU 101 further narrows down (selects) the candidates set in step S503 according to the musical instrument type input by the user. For example, when the musical instrument type input by the user is "piano", usually "distortion" is not used for tone color adjustment. Therefore, if "distortion" is included in the set candidates, it is removed from the candidates.

[0073] On the other hand, when there is no user input of the musical instrument type (NO in step S504), it proceeds to step S506. The CPU 101 acquires the waveform data of the musical instrument type "piano" set by default from the HDD 109 and proceeds to step S507. Note that the waveform data of the musical instrument type set by default is not limited to this embodiment, and may be waveform data of other musical instrument types such as organs and guitars. Also, in step S506, the CPU 101 may display a plurality of tags each describing a plurality of musical instrument types on the display unit 108, and acquire the waveform data of the musical instrument type displayed on the tag selected by the user from the HDD 109.

[0074] In step S507, the CPU 101 list-displays the candidates of the effect parameters set in step S503 on the display unit 108. Specifically, as shown in the I / F 603 of FIG. 6, the candidates of the effect parameters set in step S503 are respectively displayed as user-selectable tabs such as "tone color 1" tab, "tone color 2" tab, ···. Also, a playback button is provided on each tab.

[0075] In step S508, the CPU 101 determines whether there is a playback instruction for one of the candidates of the effect parameters set in step S503. Specifically, it determines whether any of the playback buttons provided on each tab of the I / F 603 has been pressed. When there is a playback instruction for one of the candidates (YES in step S508), it proceeds to step S509.

[0076] In step S509, the CPU 101 reverses the color of the tab (or the part of its playback button) on which the playback button has been pressed on the display unit 108, notifies the user that the tab has been selected by the user, and adjusts the tone color using the effect parameters of the candidate for which a playback instruction has been given and the waveform data acquired in either step S505 or S506, and causes the speaker 107 to emit sound (playback) as a sample tone color.

[0077] In step S510, the CPU 101 determines whether or not the candidate for which a playback instruction has been given has been selected by the user as the effect parameter to be used for tone color adjustment. Specifically, after causing the sample tone color to be emitted by the speaker 107 in step S508, when the decision button 604 is pressed in the I / F 603 without any other playback button being pressed, it is determined that the candidate for which a playback instruction has been given has been selected by the user as the effect parameter to be used for tone color adjustment.

[0078] That is, when there is a playback instruction for one of the other candidates without the decision button 604 being pressed (NO in step S510, YES in step S508), the processing after step S509 is repeated. On the other hand, when the decision button 604 is pressed without any playback instruction for one of the other candidates (YES in step S510), the CPU 101 performs tone color adjustment so that the reproduced sample tone color becomes the tone color when played on the information processing apparatus 100, and then proceeds to step S511.

[0079] In step S511, the CPU 101 causes the GPU 102 to perform additional learning of the learned model generated by the learning unit 201 based on the adjective extracted in step S501 and the effect parameter to be used for tone color adjustment selected by the user in step S510. After that, after updating the parameters of the decoder that constitutes the estimator 203 with the parameters of the decoder part of the learned model after the additional learning, this process ends. As a result, the more the user performs tone color adjustment by the process of FIG. 5 when performing performance on the information processing apparatus 100, the more customized candidate effect parameters will be listed on the I / F 603.

[0080] According to the present embodiment, when a user inputs a natural language representing a tone color that the user wants to use for the performance of the information processing apparatus 100 as text into the I / F601 on the display unit 108, the CPU 101 sets candidates for effect parameters that the user uses for tone color adjustment based on the input natural language, and displays playback buttons for playing back sample tone colors of the respective candidates on the I / F603. When the user presses the playback button displayed on the I / F603 to play back the sample tone color and confirms that it is the tone color that the user wants to use for the performance of the information processing apparatus 100, the user can adjust the tone color when performing a performance using the information processing apparatus 100 only by pressing the determination button 604. That is, even when the user is a beginner and it is difficult to adjust the effect parameters that the user wants to use for the performance of the information processing apparatus 100 by operating the buttons and knobs provided in large numbers in a conventional synthesizer, the tone color when performing a performance using the information processing apparatus 100 can be easily adjusted.

[0081] Also, without operating the buttons and knobs provided in large numbers in a conventional synthesizer, the waveform data of the instrument type when performing a performance with the information processing apparatus 100 can be easily set.

[0082] Note that the method of additional learning performed in step S511 is not particularly limited. For example, the training data generated in step S403 may be updated based on the content selected by the user using the I / F603 in the process of FIG. 5, or reinforcement learning may be performed in which the fact that the user has been selected in step S510 is given as a reward.

[0083] In this embodiment, the information processing apparatus 100 performs all the processes shown in FIGS. 4 and 5, but the configuration is not limited thereto. For example, the information processing apparatus 100 may be connected to a mobile terminal (not shown) such as a tablet or a smartphone, or a server (cloud) (not shown), and may cooperate with these, that is, the processes for each apparatus may be shared so that the processes may be performed anywhere. For example, a learned model may be generated in the cloud, and the I / F 601 in FIG. 6 may be displayed on the mobile terminal.

[0084] The training of the learning model and the additional learning of the learned model in the learning unit 201 can be performed by any machine learning method. For example, methods such as Gaussian process regression (Bayesian optimization), policy gradient method which is a kind of policy iteration method, and genetic algorithm which is a method that mimics the process of biological evolution may be adopted.

[0085] In addition, a storage medium storing each control program represented by software for achieving the present invention may be read by each apparatus so as to achieve the same effect. In that case, the program code itself read from the storage medium realizes the novel functions of the present invention, and the non-transitory computer-readable recording medium storing the program code constitutes the present invention. Further, the program code may be supplied via a transmission medium or the like. In that case, the program code itself constitutes the present invention. Note that as the storage medium in these cases, in addition to a ROM, a floppy disk, a hard disk, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a magnetic tape, a non-volatile memory card, etc. can be used. The "non-transitory computer-readable recording medium" includes those that hold a program for a certain period of time, such as a volatile memory (for example, DRAM (Dynamic Random Access Memory)) inside a computer system that becomes a server or a client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line.

Explanation of Signs

[0086] 100 Information processing apparatus, 101 CPU, 102 GPU, 105 Operation unit, 107 Speaker, 108 Display unit, 109 HDD, 201 Learning unit, 202 Input unit, 203 Estimation unit, 204 Output unit

Claims

1. An input module for receiving natural language input including adjectives from a user, A timbre estimation module that outputs a plurality of timbre data based on the natural language input by the user using a trained model that outputs timbre data from adjectives, An information processing apparatus comprising: a presentation module that presents the plurality of timbre data to the user as candidates for timbre data to be selected by the user.

2. The information processing apparatus according to claim 1, wherein the presentation module pronounces candidates for the timbre data.

3. The information processing apparatus according to claim 2, wherein the candidates for the timbre data are constituted by at least one of waveform data and effect parameters.

4. The information processing apparatus according to claim 3, wherein the candidates for the timbre data are a timbre data set including a plurality of timbre data.

5. The information processing apparatus according to claim 3 or 4, wherein when the candidates for the timbre data are constituted only by effect parameters, the presentation module pronounces the effect parameters, which are the candidates for the timbre data, in combination with default waveform data.

6. The information processing apparatus according to claim 5, wherein when the natural language input by the user includes an instrument type, the presentation module pronounces the effect parameters, which are the candidates for the timbre data, in combination with waveform data of the instrument type instead of the default waveform data.

7. The information processing apparatus according to claim 6, wherein the presentation module narrows down the candidates for the timbre data according to the instrument type.

8. The information processing apparatus according to any one of claims 1 to 7, wherein additional learning of the trained model is performed based on the timbre data selected by the user from among the candidates for the timbre data and the adjectives included in the natural language input by the user.

9. The information processing apparatus according to any one of claims 1 to 8, wherein the timbre estimation module obtains a latent variable with an adjective included in the natural language input by the user tagged from a latent space, and inputs the obtained latent variable into the trained model to output the plurality of timbre data.

10. Obtain natural language input by the user that includes adjectives, Output a plurality of timbre data based on the natural language input by the user using a trained model that outputs timbre data from adjectives, A control method realized by a computer that presents the plurality of timbre data to a user as candidates for timbre data selected by the user.

Citation Information

Patent Citations

  • Electronic musical instrument

    JP1994149243A

  • Tone color selecting device and tone color adjusting device

    JP1997325773A

  • Timbre setting device and program

    JP2006030414A

  • Method and device for constituting musical sound contents, and program and recording medium therefor

    JP2006235201A

  • Apparatus for musical performance

    JP2007156109A