Sound generation method, sound generation system, and program

The audio generation method uses trained generative models to process note and text data, enabling the creation of instrument sounds with diverse acoustic characteristics, addressing the limitations of conventional synthesis techniques.

JP7740068B2Active Publication Date: 2025-09-17YAMAHA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022036293
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-09
Publication Date
2025-09-17
Estimated Expiration
2042-03-09

AI Technical Summary

Technical Problem

Conventional sound synthesis techniques struggle to generate synthetic sounds with diverse acoustic characteristics beyond those dictated by musical scores.

Method used

An audio generation method that processes first and second control data sequences representing musical notes and text characteristics using trained generative models to produce instrument sounds with varied acoustic traits.

Benefits of technology

Generates instrument sounds with diverse acoustic characteristics by incorporating text-related musical expressions, enhancing the variety and realism of synthesized sounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007740068000001
    Figure 0007740068000001
  • Figure 0007740068000002
    Figure 0007740068000002
  • Figure 0007740068000003
    Figure 0007740068000003
Patent Text Reader

Abstract

To generate an acoustic data string of musical instrument sound having various acoustic characteristics.SOLUTION: A sound generation system includes: a control data string acquisition part 30 for acquiring a first control data string X for indicating a feature of a note string and a second control data string Y for indicating the feature of a text corresponding to the note string; and an acoustic data string generation part 33 for generating an acoustic data string Z for indicating musical instrument sound of the note string having an acoustic characteristic corresponding to the feature of the text which the second control data string Y indicates by processing the first control data string X and the second control data string Y by a trained generation model Mb.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to techniques for generating an acoustic data sequence representing a musical instrument sound. [Background technology]

[0002] Techniques for synthesizing desired sounds have been proposed in the past. For example, Non-Patent Document 1 discloses a technique for generating synthetic sounds corresponding to a sequence of musical notes using a trained generative model. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Blaauw, Merlijn, and Jordi Bonada. "A NEURAL PARAMETRIC SINGING SYNTHESIZER." arXiv preprint arXiv: 1704.03809v3 (2017) Summary of the Invention [Problem to be solved by the invention]

[0004] However, conventional synthesis techniques only generate singing sounds that follow musical scores, and it is difficult to generate synthetic sounds with diverse acoustic characteristics. In consideration of the above circumstances, one aspect of the present disclosure aims to generate an acoustic data sequence of musical instrument sounds with diverse acoustic characteristics. [Means for solving the problem]

[0005] In order to solve the above problems, an audio generation method according to one embodiment of the present disclosure acquires a first control data sequence representing the characteristics of a sequence of notes and a second control data sequence representing the characteristics of text corresponding to the sequence of notes, and processes the first control data sequence and the second control data sequence using a trained first generation model to generate an audio data sequence representing the instrument sound of the sequence of notes having acoustic characteristics corresponding to the characteristics of the text represented by the second control data sequence.

[0006] An audio generation system according to one embodiment of the present disclosure includes a control data sequence acquisition unit that acquires a first control data sequence representing characteristics of a sequence of notes and a second control data sequence representing characteristics of text corresponding to the sequence of notes, and an audio data sequence generation unit that processes the first control data sequence and the second control data sequence using a trained first generation model to generate an audio data sequence representing the instrument sound of the sequence of notes having acoustic characteristics corresponding to the characteristics of the text represented by the second control data sequence.

[0007] A program according to one embodiment of the present disclosure causes a computer system to function as a control data sequence acquisition unit that acquires a first control data sequence representing characteristics of a sequence of notes and a second control data sequence representing characteristics of text corresponding to the sequence of notes, and an acoustic data sequence generation unit that processes the first control data sequence and the second control data sequence using a trained first generative model to generate an acoustic data sequence representing the instrument sound of the sequence of notes having acoustic characteristics corresponding to the characteristics of the text represented by the second control data sequence. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a block diagram illustrating a configuration of an information system according to a first embodiment. [Figure 2] FIG. 1 is a block diagram illustrating an example of the functional configuration of a sound generation system. [Figure 3] FIG. 10 is an explanatory diagram of the operation of a control data string acquisition unit. [Figure 4] FIG. 2 is a block diagram illustrating a configuration of a second generation unit. [Figure 5] 10 is a flowchart illustrating a detailed procedure of a synthesis process. [Figure 6] FIG. 1 is a block diagram illustrating an example of the functional configuration of a machine learning system. [Figure 7] 10 is a flowchart illustrating a detailed procedure of a learning process. [Figure 8] FIG. 10 is an explanatory diagram of the operation of a control data sequence acquisition unit in the second embodiment. [Figure 9] FIG. 2 is a schematic diagram of phoneme data. [Figure 10] FIG. 11 is a schematic diagram of a second control data sequence Y in the third embodiment. [Figure 11] FIG. 10 is an explanatory diagram of a generative model in a modified example. [Figure 12] FIG. 10 is a block diagram illustrating a functional configuration of a sound generation system according to a modified example. [Figure 13] FIG. 10 is a block diagram illustrating a functional configuration of a sound generation system according to a modified example. [Figure 14] 10A and 10B are explanatory diagrams illustrating the operation of a control data sequence acquisition unit in a modified example. DETAILED DESCRIPTION OF THE INVENTION

[0009] A: First embodiment 1 is a block diagram illustrating the configuration of an information system 100 according to the first embodiment. The information system 100 includes a sound generation system 10 and a machine learning system 20. The sound generation system 10 and the machine learning system 20 communicate with each other via a communication network 200 such as the Internet.

[0010] [Sound Generation System 10] The sound generation system 10 is a computer system that generates a performance sound of a specific piece of music (hereinafter referred to as a "target sound") The target sound in the first embodiment is a musical instrument sound having the timbre of a musical instrument.

[0011] The sound generation system 10 includes a control device 11, a storage device 12, a communication device 13, an operation device 14, and a sound emission device 15. The sound generation system 10 is realized by an information terminal such as a smartphone, a tablet terminal, or a personal computer. The sound generation system 10 may be realized by a single device, or may be realized by multiple devices configured separately from each other.

[0012] The control device 11 is composed of one or more processors that control each element of the sound generation system 10. For example, the control device 11 is composed of one or more types of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a sound processing unit (SPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC). The control device 11 generates an acoustic signal A that represents the waveform of a target sound.

[0013] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is configured with a known storage medium such as a magnetic storage medium or a semiconductor storage medium. The storage device 12 may also be configured with a combination of multiple types of storage media. Note that the storage device 12 may be a portable storage medium that is detachable from the sound generation system 10, or a storage medium that the control device 11 can access via the communication network 200 (e.g., cloud storage).

[0014] The storage device 12 stores music data D representing a piece of music. The music data D includes musical score data G and a word string T. The musical score data G specifies the time sequence of notes that make up the piece of music. Specifically, the musical score data G specifies the pitch and sound duration for each of the multiple notes in the piece of music. The sound duration is specified, for example, by the start point and duration of the note. The word string T specifies text corresponding to the piece of music. Specifically, the word string T specifies one or more characters for each of the multiple notes in the piece of music. A word string T is made up of multiple characters corresponding to different notes. For example, a music file that complies with the MIDI (Musical Instrument Digital Interface) standard is used as the music data D. Note that the music data D may also specify information such as performance symbols that express musical expression.

[0015] The communication device 13 communicates with the machine learning system 20 via a communication network 200. Note that the communication device 13, which is separate from the sound generation system 10, may be connected to the sound generation system 10 by wire or wirelessly.

[0016] The operation device 14 is an input device that accepts operations by a user. For example, an operator operated by a user or a touch panel that detects contact by a user is used as the operation device 14.

[0017] The sound emitting device 15 reproduces the target sound represented by the acoustic signal A. The sound emitting device 15 is, for example, a speaker or headphones. For convenience, a D / A converter that converts the acoustic signal A from digital to analog and an amplifier that amplifies the acoustic signal A are not shown in the figure. Furthermore, the sound emitting device 15, which is separate from the sound generation system 10, may be connected to the sound generation system 10 by wire or wirelessly.

[0018] 2 is a block diagram illustrating an example of the functional configuration of the sound generation system 10. The control device 11 executes a program stored in the storage device 12 to realize multiple functions for generating the sound signal A (a control data sequence acquisition unit 30, a sound data sequence generation unit 33, and a signal generation unit 34).

[0019] FIG. 3 is an explanatory diagram of the operation of the control data sequence acquisition unit 30. The control data sequence acquisition unit 30 acquires a first control data sequence X and a second control data sequence Y. Specifically, the control data sequence acquisition unit 30 acquires the first control data sequence X and the second control data sequence Y in each of a plurality of unit periods U on the time axis. Each unit period U is a period (the hop size of the frame window) with a time length sufficiently short compared to the duration of each note in the music piece. For example, the window size is 2 to 20 times the hop size (the window is longer), the hop size is 2 to 20 milliseconds, and the window size is 20 to 60 milliseconds. The control data sequence acquisition unit 30 of the first embodiment includes a first generation unit 31 and a second generation unit 32.

[0020] The first generation unit 31 generates first control data X from a note data sequence N for each unit period U. The note data sequence N used for generation is a portion of the musical score data G that corresponds to each unit period U. The note data sequence N corresponding to any one unit period U is a portion of the note data sequence of the music data D that includes note data of notes that include that unit period U (hereinafter referred to as "target notes"). In other words, the note data sequence N specifies a note sequence of the music data D that includes the target note and at least one of the notes preceding and following the target note.

[0021] Each piece of first control data X is data in any format that represents the characteristics of the note sequence specified by the note data sequence N. The first control data X in any one unit period U is information indicating the characteristics of a note indicated by the note data of a target note that includes the unit period U among the multiple notes in a piece of music. For example, the characteristics indicated by the first control data sequence X include the characteristics (e.g., pitch and, optionally, duration) of the note that includes the unit period. The first control data sequence X also includes information about notes other than the target note. For example, the first control data sequence X includes the characteristics (e.g., pitch) indicated by the note data of at least one of the notes before and after the note that includes the unit period. The first control data sequence X may also include the pitch difference between the target note and the note immediately before or after it. If there is no note before or after it that should be included and there is a rest, the characteristics of that rest may be included instead of the note.

[0022] The first generation unit 31 generates the first control data sequence X by performing a predetermined calculation process on the musical note data sequence N. The first generation unit 31 may generate the first control data sequence X by using a generative model configured by a deep neural network (DNN) or the like. The generative model is a statistical estimation model that learns the relationship between the musical note data sequence N and the first control data sequence X by machine learning. The first control data sequence X is data that specifies the musical conditions of the target sound to be generated by the sound generation system 10.

[0023] The second generation unit 32 generates second control data Y required for the current unit period U from the word sequence T in synchronization with or prior to the progress of the unit period U. The second control data Y for each unit period U is data in any format that represents the characteristics of the current phrase that includes that unit period U in the word sequence T. Specifically, the second control data Y includes a phrase vector V of that phrase included in the word sequence T. The phrase vector V is a vector that represents the position of each phrase in the semantic space. The closer the meanings of multiple phrases are, the closer the positions of the phrase vectors V of those phrases are in the semantic space. A phrase represented by a phrase vector V is composed of one or more words. In other words, the phrase vector V is data that represents the characteristics of one word or one phrase (a time series of multiple words) in the word sequence T.

[0024] FIG. 4 is a block diagram illustrating the configuration of the second generation unit 32. The second generation unit 32 includes a language analysis unit 321 and an information generation unit 322. The language analysis unit 321 divides the word string T represented by the word string T into multiple words using natural language processing such as morphological analysis. The language analysis unit 321 sequentially generates phrase data Q. The phrase data Q is data identifying a phrase composed of one or more words in the word string T, or data representing a character string of the phrase. The information generation unit 322 generates a phrase vector sequence V for the phrase represented by the phrase data Q. As illustrated in FIG. 3, in each unit period U within a period corresponding to one phrase in the music piece, the phrase vector V of the phrase is repeatedly used as the second control data Y. Note that in each unit period U within a period in which no notes or word string T are set in the music piece, a zero vector is generated as the second control data Y.

[0025] As illustrated in FIG. 4, the information generation unit 322 uses a generative model Ma to generate a word vector sequence V. The generative model Ma is a trained model that has learned, through machine learning, the latent relationship between word data Q as input and a word vector sequence V in semantic space as output. The generative model Ma outputs a word vector sequence V in response to input of word data Q. The information generation unit 322 processes the word data Q using the trained generative model Ma to generate word vectors V for each word and phrase, and outputs the word vectors V for the corresponding unit period as a word vector sequence V. As can be understood from the above explanation, the second generation unit 32 generates a word vector sequence V representing words included in the word sequence T as a second control data sequence Y using the generative model Ma. With the above configuration, the second control data sequence Y can be easily generated using the generative model Ma. Note that the generative model Ma is an example of a "second generative model."

[0026] The generative model Ma in the first embodiment is a statistical estimation model such as a deep neural network. For example, a word vector sequence V for a phrase consisting of one word is generated using a technique (Word2Vec) described in Tomas Mikolov et al., "Efficient Estimation of Word Representations in Vector Space," arXiv:1301.3781 [cs.CL], 2013. Furthermore, a word vector sequence V for a phrase (i.e., a sentence) consisting of multiple words is generated using a technique (Doc2Vec) described in Quoc Le and Tomas Mikolov, "Distributed Representations of Sentences and Documents," CoRR, abs / 1405.4053, pp.1-9, 2014.

[0027] 2, through the above processing by the control data sequence acquisition unit 30, control data C is generated for each unit period U. The control data C for each unit period U includes first control data X generated by the first generation unit 31 for that unit period U and second control data Y generated by the second generation unit 32 for that unit period U. The control data C is data obtained by concatenating, for example, the first control data X and the second control data Y with each other.

[0028] The performance of an instrument is basically determined by a sequence of notes on a musical score. Furthermore, research conducted by the inventors of the present application has confirmed that even when performers play the same sequence of notes on an instrument, if the text attached to the sequence of notes is different, the musical expression of the instrument sound produced by the performance of the instrument also tends to differ. That is, while it is natural that the musical expression of a singing voice depends on the text (i.e., lyrics), there is also a tendency for the musical expression of an instrument sound, which is generally assumed not to be affected by the text, to actually depend on the text. Based on the above findings, in the first embodiment, an audio signal A of a target sound is generated in accordance with a first control data sequence X representing the characteristics of the sequence of notes and a second control data sequence Y representing the characteristics of a word sequence T corresponding to the sequence of notes.

[0029] The sound data sequence generator 33 in FIG. 2 generates sound data sequence Z using the control data sequence C (first control data sequence X and second control data sequence Y). The sound data sequence Z is data in any format that represents a target sound. Specifically, the sound data sequence Z represents a target sound that corresponds to the sequence of notes represented by the first control data sequence X and has acoustic characteristics according to the features of the word sequence T represented by the second control data sequence Y. In other words, the instrument sound that would be produced if a performer played the sequence of notes on an instrument with the word sequence T in mind is generated as the target sound.

[0030] Specifically, the acoustic data sequence Z is data that represents the envelope of the frequency spectrum of the target sound. Specifically, acoustic data Z corresponding to each unit period U is generated in accordance with the control data C for that unit period U. Each piece of acoustic data Z corresponds to a waveform sample series for one frame window that is longer than the unit period. As explained above, the acquisition of control data C by the control data sequence acquisition unit 30 and the generation of acoustic data Z by the acoustic data sequence generation unit 33 are performed for each unit period U.

[0031] A generative model Mb is used by the acoustic data sequence generator 33 to generate the acoustic data sequence Z. The generative model Mb estimates the acoustic data Z for each unit period according to the control data C for that unit period. The generative model Mb is a trained model that has learned, by machine learning, the latent relationship between the control data sequence C as input and the acoustic data sequence Z as output. In other words, the generative model Mb outputs an acoustic data sequence Z that is statistically appropriate for the control data sequence C from the perspective of that relationship. The acoustic data sequence generator 33 generates acoustic data Z for each unit period U by processing the control data C using the generative model Mb.

[0032] The generative model Mb is realized by a combination of a program that causes the control device 11 to execute a calculation to generate an acoustic data sequence Z from a control data sequence C, and multiple variables (weights and biases) that are applied to the calculation. The program and multiple variables that realize the generative model Mb are stored in the storage device 12. The multiple variables of the generative model Mb are set in advance by machine learning. The generative model Mb is an example of a "first generative model."

[0033] The generative model Mb is configured, for example, by a deep neural network. For example, any type of deep neural network, such as a recurrent neural network (RNN) or a convolutional neural network (CNN), may be used as the generative model Mb. The generative model Mb may also be configured by combining multiple types of deep neural networks. In addition, additional elements, such as a long short-term memory (LSTM) or attention, may be incorporated into the generative model Mb.

[0034] The signal generating unit 34 generates an acoustic signal A of the target sound from the time series of the acoustic data sequence Z. The signal generating unit 34 converts the acoustic data sequence Z into a waveform signal in the time domain by, for example, performing an operation including a discrete inverse Fourier transform, and generates the acoustic signal A by concatenating the waveform signals for successive unit periods U. Note that the signal generating unit 34 may generate the acoustic signal A from the acoustic data sequence Z using, for example, a deep neural network (a so-called neural vocoder) that has learned the relationship between the acoustic data sequence Z and each sample of the acoustic signal A. The acoustic signal A generated by the signal generating unit 34 is supplied to the sound emitting device 15, and the target sound is reproduced from the sound emitting device 15.

[0035] 5 is a flowchart illustrating the detailed procedure of a process (hereinafter referred to as a "synthesis process") Sa in which the control device 11 generates an acoustic signal A. The synthesis process Sa is executed in each of a plurality of unit periods U.

[0036] When the synthesis process Sa is started, the control device 11 (control data sequence acquisition unit 30) acquires the music data D from the storage device 12 (Sa1). The control device 11 (first generation unit 31) generates first control data X for the unit period U from a note data sequence N corresponding to the unit period U in the musical score data G of the music data D (Sa2). Furthermore, the control device 11 (second generation unit 32) generates second control data Y for each unit period U from a word sequence T of the music data D (Sa3). Note that the order of generating the first control data X (Sa2) and the second control data Y (Sa3) may be reversed.

[0037] The control device 11 (acoustic data sequence generator 33) processes control data C including first control data X and second control data Y using a generative model Mb to generate acoustic data Z for a unit period U (Sa4). The control device 11 (signal generator 34) generates an acoustic signal A for the unit period U from the acoustic data Z (Sa5). From the acoustic data sequence Z for each unit period, a signal is generated that spans a time length longer than the unit period, and by overlapping and adding these signals, an acoustic signal A spanning multiple unit periods is generated. The time difference (hop size) between the previous and next frame windows corresponds to the unit period. The control device 11 supplies the acoustic signal A to the sound emitting device 15 to reproduce the target sound (Sa6).

[0038] As explained above, in the first embodiment, in addition to the first control data sequence X representing the characteristics of the note sequence, the second control data sequence Y representing the characteristics of the word sequence T corresponding to the note sequence is used to generate the sound data sequence Z. Therefore, compared to a configuration in which the sound data sequence Z is generated from only the first control data sequence X, it is possible to generate a sound data sequence Z of a target sound having a variety of acoustic characteristics according to the word sequence T corresponding to the note sequence. For example, even if the note data sequence N is the same, it is possible to generate a sound data sequence Z of a target sound with different acoustic characteristics by changing the word sequence T. In particular, in the first embodiment, the second control data sequence Y includes a word vector sequence V representing words in the word sequence T. That is, the word vector sequence V reflecting the meaning of the word sequence T is used as the second control data sequence Y. Therefore, it is possible to generate a sound data sequence Z of a target sound whose acoustic characteristics reflect the meaning of the words in the word sequence T.

[0039] [Machine Learning System 20] 1 is a computer system that establishes, by machine learning, a generative model Mb used by the sound generation system 10. The machine learning system 20 includes a control device 21, a storage device 22, and a communication device 23.

[0040] The control device 21 is composed of one or more processors that control each element of the machine learning system 20. For example, the control device 21 is composed of one or more types of processors such as a CPU, a GPU, an SPU, a DSP, an FPGA, or an ASIC.

[0041] The storage device 22 is one or more memories that store programs executed by the control device 21 and various data used by the control device 21. The storage device 22 is configured with a known recording medium, such as a magnetic recording medium or a semiconductor recording medium. The storage device 22 may also be configured with a combination of multiple types of recording media. Note that the storage device 22 may be a portable recording medium that is detachable from the machine learning system 20, or a recording medium that the control device 21 can access via the communication network 200 (e.g., cloud storage).

[0042] The communication device 23 communicates with the sound generation system 10 via the communication network 200. Note that the communication device 23, which is separate from the machine learning system 20, may be connected to the machine learning system 20 by wire or wirelessly.

[0043] FIG. 6 is an explanatory diagram of the function of the machine learning system 20 in establishing a generative model Mb. The storage device 22 stores multiple pieces of basic data B corresponding to different pieces of music. Each of the multiple pieces of basic data B includes music data D and a reference signal R. The music data D is data representing a sequence of notes in a specific piece of music (hereinafter referred to as a "reference piece of music") that is being played using a waveform represented by the reference signal R. As mentioned above, the music data D includes musical score data G and a word sequence T. The musical score data G specifies the time sequence of notes that make up the reference piece of music. The word sequence T specifies the word sequence T that corresponds to the reference piece of music.

[0044] The reference signal R is a signal representing the waveform of an instrument sound produced by an instrument when a performer plays a reference piece of music while referring to the word string T. For example, a skilled performer plays the reference piece of music while adding musical expressions corresponding to the word string T. The reference signal R is generated by recording the instrument sound produced by the instrument under the above circumstances. After recording the reference signal R, the position of the reference signal R on the time axis is adjusted. Therefore, the instrument sound represented by the reference signal R is an instrument sound having acoustic characteristics corresponding to the word string T.

[0045] The control device 21 executes a program stored in the storage device 22 to realize a plurality of functions (a training data acquisition unit 41, a learning processing unit 42) for generating the generative model Mb.

[0046] The training data acquisition unit 41 generates multiple pieces of training data L from multiple pieces of basic data B. One piece of training data L is generated for each reference piece of music. Therefore, multiple pieces of training data L are generated from multiple pieces of basic data B corresponding to different reference pieces of music. The learning processing unit 42 establishes a generative model Mb through machine learning using the multiple pieces of training data L.

[0047] Each of the multiple training data L is composed of a combination of a training control data sequence Ct and a training acoustic data sequence Zt. The control data sequence Ct is composed of a combination of a first training control data sequence Xt and a second training control data sequence Yt. The first control data sequence Xt is an example of a "first training control data sequence," and the second control data sequence Yt is an example of a "second training control data sequence." Furthermore, the acoustic data sequence Zt is an example of a "training acoustic data sequence."

[0048] For each unit period U, the training data acquisition unit 41 generates first control data Xt for the unit period U from the note data sequence Nt. The note data sequence Nt used to generate the first control data Xt for each unit period U is a portion of the note data sequence of the musical score data G that includes note data for target notes that fall within the unit period U. In other words, the note data sequence Nt includes note data for the target note in the reference music piece and at least one of the note data for the note preceding it and the note data for the note following it. Similar to the first control data sequence X described above, the first control data sequence Xt is data that represents the characteristics of the reference note sequence represented by the note data sequence Nt. The training data acquisition unit 41 generates the first control data sequence Xt for each unit period U from the note data sequence Nt using the same process as the first generation unit 31.

[0049] The second control data Yt for one unit period U indicates a word vector V estimated for a word in a word sequence T that corresponds to the unit period U. The training data acquisition unit 41 generates, by the same process as the second generation unit 32, a second control data sequence Yt for each unit period U that indicates a word vector sequence V estimated from the word sequence T.

[0050] The acoustic data Zt for one unit period U represents the waveform of one frame of the reference signal R corresponding to that unit period U. The training data acquisition unit 41 generates the acoustic data sequence Zt from the reference signal R. As can be understood from the above explanation, the acoustic data sequence Zt represents the waveform of the musical instrument sound produced by the musical instrument when the reference note sequence corresponding to the first control data sequence Xt is played under the phrase represented by the second control data sequence Yt. In other words, the acoustic data sequence Zt is the ground truth of the acoustic data sequence that should be output by the generative model Mb in response to the input of the control data sequence Ct.

[0051] 7 is a flowchart of a process Sb (hereinafter referred to as the "learning process") in which the control device 21 establishes a generative model Mb through machine learning. For example, the learning process Sb is started in response to an instruction from the operator of the machine learning system 20. The control device 21 executes the learning process Sb to implement the learning processing unit 42 in FIG. 6.

[0052] When the learning process Sb starts, the control device 21 selects one of a plurality of training data L (hereinafter referred to as "selected training data L") (Sb1). As illustrated in Fig. 6, the control device 21 processes a control data sequence Ct of the selected training data L using an initial or provisional generative model Mb (hereinafter referred to as "provisional model Mb0") to generate an acoustic data sequence Z (Sb2).

[0053] The control device 21 calculates a loss function that represents the error between the acoustic data sequence Z generated by the provisional model Mb0 and the acoustic data sequence Zt of the selected training data L (Sb3). The control device 21 updates multiple variables of the provisional model Mb0 so that the loss function is reduced (ideally minimized) (Sb4). The backpropagation algorithm, for example, is used to update each variable according to the loss function.

[0054] The control device 21 determines whether a predetermined termination condition is met (Sb5). The termination condition is that the loss function falls below a predetermined threshold, or that the amount of change in the loss function falls below a predetermined threshold. If the termination condition is not met (Sb5: NO), the control device 21 selects unselected training data L as new selected training data L (Sb1). That is, the process of updating multiple variables of the provisional model Mb0 (Sb1 to Sb4) is repeated until the termination condition is met (Sb5: YES). If the termination condition is met (Sb5: YES), the control device 21 ends the learning process Sb. The provisional model Mb0 at the time the termination condition is met is determined to be the trained generative model Mb.

[0055] As can be understood from the above explanation, the generative model Mb learns the latent relationship between the input control data sequence Ct and the output acoustic data sequence Zt. Therefore, the trained generative model Mb outputs an acoustic data sequence Z that is statistically plausible for the unknown control data sequence C in terms of this relationship.

[0056] The control device 21 transmits the generative model Mb established by the above processing to the sound generation system 10 from the communication device 23. Specifically, multiple variables that define the generative model Mb are transmitted to the sound generation system 10. The control device 11 of the sound generation system 10 receives the generative model Mb transmitted from the machine learning system 20 via the communication device 13 and stores the generative model Mb in the storage device 12.

[0057] B: Second embodiment A second embodiment will be described. Note that, for elements in the following exemplary aspects that have the same functions as those in the first embodiment, the same reference numerals as those in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.

[0058] The control device 11 in the sound generation system 10 of the second embodiment, like the first embodiment, includes a control data sequence acquisition unit 30 that acquires a control data sequence C, a sound data sequence generation unit 33 that generates a sound data sequence Z from the control data sequence C, and a signal generation unit 34 that generates a sound signal A from the sound data sequence Z.

[0059] FIG. 8 is an explanatory diagram of the operation of the control data sequence acquisition unit 30 in the second embodiment. Similar to the first embodiment, the first generation unit 31 of the control data sequence acquisition unit 30 generates first control data X from a note data sequence N for each unit period U. In the second embodiment, the function of the second generation unit 32 differs from that of the first embodiment. The second generation unit 32 in the first embodiment generates a phrase vector sequence V representing each phrase in a word sequence T as a second control data sequence Y. On the other hand, the second generation unit 32 in the second embodiment generates phoneme data P representing each phoneme in the word sequence and outputs the phoneme data P as second control data Y for each unit period corresponding to the duration of the phoneme. That is, the second generation unit 32 analyzes the word sequence T to generate phoneme data P indicating the phoneme type and duration for each phoneme, and outputs the phoneme data P as second control data Y for each unit period U. Similar to the first embodiment, the control data sequence C includes a first control data sequence X and a second control data sequence Y.

[0060] FIG. 9 is a schematic diagram of phoneme data P. The phoneme data P specifies one of multiple (K types) phonemes. Specifically, the phoneme data P is composed of K elements E (E1 to EK) (K is a natural number equal to or greater than 2) corresponding to different types of phonemes. The phoneme data P specifying any one type of phoneme is a one-hot vector in which one element E corresponding to the phoneme among the K elements E1 to EK is set to "1" and the remaining (K-1) elements E are set to "0". Note that a one-cold vector in which the "1" and "0" of each element E are interchanged may also be used as the phoneme data P.

[0061] The second generation unit 32 estimates the type and duration of each phoneme of characters included in the word string T at each point in time through phoneme analysis processing, and generates phoneme data P that specifies the phoneme. Any known technology may be employed for the phoneme analysis processing. As illustrated in FIG. 8 , in each unit period U within a period corresponding to one phoneme in a piece of music, the phoneme data P indicating the phoneme is repeatedly used as the second control data Y. The boundaries of the periods of each phoneme of characters included in the word string T at each point in time are estimated using a statistical model such as an HMM (Hidden Markov Model) or an SVM (Support Vector Machine). It is also possible to use a rule-based method for identifying the boundaries of each phoneme using a reference table that stores the relationship between each character constituting the word string T and the boundaries of each phoneme. Furthermore, the boundaries of each phoneme manually specified by the creator of the music data D may be specified by the music data D.

[0062] The second control data string Y (phrase vector string V) of the first embodiment reflects the meaning of the word string T, but does not reflect information about the pronunciation of the word string T (i.e., phonemes). On the other hand, the second control data string Y (phoneme data P) of the second embodiment reflects information about the pronunciation of the word string T (i.e., phonemes), but does not reflect the meaning of the word string T. The phrase vector string V and the phoneme data P are comprehensively expressed as data representing the characteristics of the word string T.

[0063] The second embodiment is the same as the first embodiment except that the phrase vector sequence V is replaced with phoneme data P. For example, the process in which the sound data sequence generator 33 generates the sound data sequence Z from the control data sequence C and the process in which the signal generator 34 generates the sound signal A from the sound data sequence Z are the same as those in the first embodiment. The synthesis process Sa and the learning process Sb are also the same as those in the first embodiment.

[0064] In the second embodiment, in addition to a first control data sequence X representing the characteristics of a note sequence, a second control data sequence Y representing the characteristics of a word sequence T corresponding to the note sequence is used to generate the sound data sequence Z. Therefore, as in the first embodiment, it is possible to generate a sound data sequence Z of a target sound having a variety of acoustic characteristics according to the word sequence T corresponding to the note sequence. In particular, in the second embodiment, the second control data sequence Y includes phoneme data P representing the phonemes in the word sequence T. That is, the phoneme data P reflecting the pronunciation of the word sequence T is used as the second control data sequence Y. Therefore, it is possible to generate a sound data sequence Z of a target sound that reflects non-linguistic characteristics (e.g., characteristics in the time domain or frequency domain) related to the pronunciation of the phonemes in the word sequence T. For example, a target sound that gives the impression that the word sequence T is perceived as an onomatopoeia is generated.

[0065] C: Third embodiment 10 is a schematic diagram of the second control data sequence Y in the third embodiment. The second control data sequence Y includes first data Y1 and second data Y2. The first data Y1 corresponds to the second control data sequence Y in the first embodiment, and the second data Y2 corresponds to the second control data sequence Y in the second embodiment.

[0066] Specifically, the first data Y1 is a phrase vector sequence V representing each phrase included in the word sequence T. For example, in each unit period U within a period corresponding to one phrase in the music piece, the phrase vector V of that phrase is used as the first data Y1. On the other hand, the second data Y2 is phoneme data P representing each phoneme in the word sequence T. For example, in each unit period U within a period corresponding to one phoneme in the music piece, the phoneme data P of that phoneme is used as the second control data Y2.

[0067] The third embodiment also achieves the same effects as the first embodiment. Furthermore, in the third embodiment, the second control data sequence Y includes first data Y1 (phrase vector sequence V) and second data Y2 (phoneme data P). Therefore, it is possible to generate an acoustic data sequence Z of a target sound that reflects both the meaning of each word in the word sequence T and the pronunciation of the phonemes in the word sequence T.

[0068] D: Modification Specific modified embodiments that can be added to each of the embodiments exemplified above are exemplified below. Multiple embodiments arbitrarily selected from the following examples may be combined as appropriate within the scope of not mutually contradicting each other.

[0069] (1) The phoneme data P in the second embodiment is not limited to a vector consisting of K elements E1 to EK. For example, a code sequence (identifier) ​​uniquely assigned to each phoneme may be used as the phoneme data P.

[0070] (2) In the above embodiments, the acoustic data sequence Z represents the frequency characteristics of the target sound, but the information represented by the acoustic data sequence Z is not limited to these examples. For example, the acoustic data sequence Z may represent samples of the target sound. In the above embodiments, the time series of the acoustic data sequence Z constitutes the acoustic signal A. Therefore, the signal generator 34 is omitted.

[0071] (3) In the above embodiments, the control data sequence acquirer 30 generates the first control data sequence X and the second control data sequence Y. However, the operation of the control data sequence acquirer 30 is not limited to these examples. For example, the control data sequence acquirer 30 may receive the first control data sequence X and the second control data sequence Y generated by an external device from the external device via the communication device 13. In addition, in an embodiment in which the first control data sequence X and the second control data sequence Y are stored in the storage device 12, the control data sequence acquirer 30 reads the first control data sequence X and the second control data sequence Y from the storage device 12. As can be understood from the above examples, the "acquisition" by the control data sequence acquirer 30 encompasses any operation of acquiring the first control data sequence X and the second control data sequence Y, such as generating, receiving, and reading the first control data sequence X and the second control data sequence Y. Similarly, the "acquisition" of the first control data sequence Xt and the second control data sequence Yt by the training data acquisition unit 41 also encompasses any operation of acquiring the first control data sequence Xt and the second control data sequence Yt (for example, generating, receiving, and reading).

[0072] (4) In each of the above-described embodiments, a control data sequence C concatenating a first control data sequence X and a second control data sequence Y is supplied to a generative model Mb. However, the input of the first control data sequence X and the second control data sequence Y to the generative model Mb is not limited to the above examples.

[0073] For example, as illustrated in FIG. 11, consider a configuration in which a generative model Mb is composed of a first portion Mb1 and a second portion Mb2. The first portion Mb1 is a portion composed of the input layer and part of the hidden layer of the generative model Mb. The second portion Mb2 is a portion composed of another part of the hidden layer and the output layer of the generative model Mb. In the above configuration, a first control data sequence X may be supplied to the first portion Mb1 (input layer), and a second control data sequence Y may be supplied to the second portion Mb2 together with data output from the first portion Mb1. As can be understood from the above example, concatenation of the first control data sequence X and the second control data sequence Y is not essential to the present disclosure.

[0074] (5) As illustrated in Fig. 12, multiple generative models Mb corresponding to different musical instruments may be selectively used. A generative model Mb corresponding to one type of musical instrument is a trained model trained using a reference signal R of the musical instrument sound produced by that musical instrument. Therefore, the generative model Mb corresponding to each musical instrument outputs an acoustic data string Z representing the musical instrument sound of that musical instrument.

[0075] The user operates the operation device 14 to select one of a plurality of musical instruments. The musical instrument data α in FIG. 12 is data specifying the musical instrument selected by the user. The sound data sequence generator 33 selects a generative model Mb from a plurality of generative models Mb that corresponds to the musical instrument specified by the musical instrument data α, and generates a sound data sequence Z by processing the control data sequence C using that generative model Mb. With the above configuration, it is possible to generate a target sound with a timbre corresponding to one of a plurality of musical instruments.

[0076] 13, a control data sequence C including musical instrument data α in addition to a first control data sequence X and a second control data sequence Y may be input to a single generative model Mb. The generative model Mb in FIG. 13 is established by machine learning using multiple reference signals R corresponding to different musical instruments. Furthermore, the training data L includes not only the first control data sequence Xt and the second control data sequence Yt for training, but also the musical instrument data α specifying the musical instrument corresponding to the reference signal R. Therefore, the target sound represented by the sound data sequence Z is a musical instrument sound having the timbre of the instrument specified by the musical instrument data α.

[0077] (6) In the above-described embodiments, the note data string N is generated from the music data D pre-stored in the storage device 12. However, it is also possible to use a note data string N sequentially supplied from a performance device. The performance device is an input device such as a MIDI keyboard that accepts a performance by a user and sequentially outputs a note data string N corresponding to the user's performance. The sound generation system 10 generates a sound data string Z using the note data string N supplied from the performance device. The synthesis process Sa described above may be executed in real time in parallel with the user's performance on the performance device. Specifically, the second control data string Y and the sound data string Z may be generated in parallel with the user's operation on the performance device.

[0078] (7) In the above-described embodiments, one phrase vector sequence V is generated from one piece of phrase data Q. However, multiple phrase vector sequences V may be generated from one piece of phrase data Q. FIG. 14 is an explanatory diagram of the operation of the control data sequence acquisition unit 30 in this modification. The phrase data sequence Q in FIG. 14 is data representing a phrase sequence obtained by dividing a word sequence T into phrases. In other words, the phrase data Q is data identifying one or more phrase sequences corresponding to one phrase. A phrase is a section obtained by dividing a piece of music according to musical or semantic unity. For example, each phrase is specified in the music data D. However, each phrase of a piece of music may also be defined by analyzing the music data D. As described above, the language analysis unit 321 generates the phrase data sequence Q by dividing a word sequence T into phrases.

[0079] The information generation unit 322 processes the phrase data sequence Q to generate a phrase vector sequence V for each phrase included in the corresponding phrase sequence. As illustrated in FIG. 14 , when one phrase data sequence Q includes multiple phrases, the context of the phrase sequence is analyzed, and multiple phrase vector sequences V corresponding to the respective phrases are generated from the single phrase data sequence Q. As in the first embodiment, a generative model Ma capable of interpreting context is used to generate the phrase vector sequence V by the information generation unit 322. The generative model Ma is a trained model that has learned the relationship between the context of the phrase sequence indicated by the phrase data sequence Q and the phrase vector sequence V indicating the meaning of each phrase in that context. Specifically, the generative model Ma generates a phrase vector sequence V indicating the meaning of each word included in the phrase sequence indicated by the phrase data Q in the phrase sequence. For example, a natural language processing model such as BERT (Bidirectional Encoder Representations from Transformers) is used as the generative model Ma. As can be understood from the above explanation, in this modification, one or more phrase vector sequences V corresponding to the phrase data Q are generated for each phrase.

[0080] For example, FIG. 14 illustrates a phrase data sequence Q1 that specifies a phrase sequence #1 composed of words #1a and #1b. In response to an input of one phrase data sequence Q1, the generative model Ma generates a phrase vector sequence V corresponding to word #1a included in the phrase sequence #1 and a phrase vector sequence V corresponding to word #1b. Furthermore, the phrase data sequence Q2 in FIG. 14 specifies a phrase sequence #2 composed of words #2a, #2b, and #2c. In response to an input of one phrase data sequence Q2, the generative model Ma generates a phrase vector sequence V corresponding to word #2a included in the phrase sequence #2, a phrase vector sequence V corresponding to word #2b, and a phrase vector sequence V corresponding to word #2c. As in the first embodiment, in each unit period U corresponding to a single phrase, the phrase vector V of the phrase is repeatedly used as the second control data Y. In this modification, the context of the phrase sequence is interpreted to generate phrase vectors V that more accurately indicate the meaning of each phrase included in the phrase sequence.

[0081] (8) Although a deep neural network is exemplified in each of the above embodiments, the generative model M(Ma, Mb, Mc) is not limited to a deep neural network. For example, any type of statistical model, such as an HMM (Hidden Markov Model) or an SVM (Support Vector Machine), may be used as the generative model M(Ma, Mb, Mc).

[0082] (9) In each of the above-described embodiments, the machine learning system 20 establishes the generative model Mb, but the function of establishing the generative model Mb (the training data acquisition unit 41 and the learning processing unit 42) may be included in the sound generation system 10. Also, the function of establishing the generative model Ma or the generative model Mc may be included in the sound generation system 10.

[0083] (10) The sound generation system 10 may be realized by a server device that communicates with an information device such as a smartphone or a tablet terminal. For example, the sound generation system 10 receives music data D from the information device and generates an audio signal A by a synthesis process Sa that applies the music data D. The sound generation system 10 transmits the audio signal A generated by the synthesis process Sa to the information device. Note that in a configuration in which the signal generation unit 34 is installed in the information device, the time series of the audio data string Z is transmitted to the information device. In other words, the signal generation unit 34 is omitted from the sound generation system 10.

[0084] (11) As described above, the functions of the sound generation system 10 (the control data sequence acquisition unit 30, the sound data sequence generation unit 33, and the signal generation unit 34) are realized by the cooperation of one or more processors constituting the control device 11 and the programs stored in the storage device 12. Also, the functions of the machine learning system 20 (the training data acquisition unit 41 and the learning processing unit 42) are realized by the cooperation of one or more processors constituting the control device 21 and the programs stored in the storage device 22.

[0085] The programs exemplified above can be provided in a form stored on a computer-readable recording medium and installed on a computer. The recording medium is, for example, a non-transitory recording medium, such as an optical recording medium (optical disk) such as a CD-ROM, but also includes any known type of recording medium, such as a semiconductor recording medium or a magnetic recording medium. Note that a non-transitory recording medium includes any recording medium other than a transitory, propagating signal, and does not exclude volatile recording media. Furthermore, in a configuration in which a distribution device distributes a program via a communication network 200, the recording medium that stores the program in the distribution device corresponds to the non-transitory recording medium described above.

[0086] E: Notes From the above-described exemplary embodiments, the following configurations can be understood, for example.

[0087] An audio generation method according to one aspect (aspect 1) of the present disclosure acquires a first control data sequence representing characteristics of a sequence of notes and a second control data sequence representing characteristics of text corresponding to the sequence of notes, and processes the first control data sequence and the second control data sequence using a trained first generation model to generate an audio data sequence representing the instrument sound of the sequence of notes having acoustic characteristics corresponding to the characteristics of the text represented by the second control data sequence.

[0088] In the above-described aspects, in addition to a first control data sequence representing the characteristics of a sequence of notes, a second control data sequence representing the characteristics of the text corresponding to the sequence of notes is also used to generate an audio data sequence. Therefore, compared to a configuration in which an audio data sequence is generated from only the first control data sequence, it is possible to generate an audio data sequence of musical instrument sounds with a variety of acoustic characteristics according to the text corresponding to the sequence of notes.

[0089] The "first control data sequence" is data (first control data) in any format that represents the characteristics of a note sequence, and is generated, for example, from a note data sequence that represents the note sequence. The first control data sequence may also be generated from a note data sequence that is generated in real time in response to operations on an input device such as an electronic musical instrument. The "first control data sequence" can also be described as data that specifies the conditions of the musical instrument sound to be synthesized. For example, the "first control data sequence" specifies various conditions related to each note that makes up the note sequence, such as the pitch or duration of each note that makes up the note sequence, and the relationship between the pitch of one note and the pitches of other notes located around that note. The "musical instrument sound" is a musical sound that is produced from an instrument when played.

[0090] The "first generative model" is a trained model that has learned the relationship between the first and second control data sequences and the audio data sequence through machine learning. A plurality of training data are used for the machine learning of the first generative model. Each training data set includes a pair of the first and second training control data sequences and a training audio data sequence. The first training control data sequence is data representing the characteristics of a reference note sequence, and the second training control data sequence is data representing the characteristics of text corresponding to the reference note sequence. The training audio data sequence represents the instrument sounds produced by playing the note sequence corresponding to the first training control data sequence and the text corresponding to the second training control data sequence. For example, various statistical estimation models, such as a deep neural network (DNN), a hidden Markov model (HMM), or a support vector machine (SVM), are used as the "first generative model."

[0091] The first control data sequence and the second control data sequence may be input to the first generative model in any form. For example, input data including the first control data sequence and the second control data sequence may be input to the first generative model. Furthermore, in a configuration in which the first generative model includes an input layer, multiple intermediate layers, and an output layer, it is also possible for the first control data sequence to be input to the input layer, and the second control data sequence to be input to the intermediate layer. In other words, combining the first control data sequence and the second control data sequence is not required.

[0092] An "acoustic data sequence" is data (acoustic data) in any format that represents the sound of a musical instrument. For example, data that represents acoustic characteristics (frequency characteristics) such as an intensity spectrum, a Mel spectrum, or Mel-Frequency Cepstrum Coefficients (MFCC) are examples of an "acoustic data sequence." A sample sequence that represents the waveform of a musical instrument sound may also be generated as an "acoustic data sequence."

[0093] "Text corresponding to a sequence of notes" means that the text is associated with the sequence of notes. In other words, the "correspondence" between a sequence of notes and text means, for example, that each note in the sequence of notes is associated with each word in the text in terms of time.

[0094] In a specific example (Aspect 2) of Aspect 1, the first generative model is a model trained using training data including a first training control data sequence representing characteristics of a reference sequence of notes, a second training control data sequence representing characteristics of text corresponding to the reference sequence of notes, and a training audio data sequence representing instrument sounds of the reference sequence of notes. According to the above aspect, it is possible to generate an audio data sequence that is statistically valid in terms of the relationship between the first training control data sequence and the second training control data sequence of the reference sequence of notes and the training audio data sequence representing the instrument sounds of the reference sequence of notes.

[0095] In a specific example (Aspect 3) of Aspect 1 or Aspect 2, the second control data sequence includes a phrase vector sequence representing phrases included in the text. According to the above aspect, the second control data sequence includes a phrase vector sequence representing phrases in the text. Therefore, it is possible to generate an acoustic data sequence of musical instrument sounds in which the meanings of the phrases in the text are reflected in the acoustic characteristics.

[0096] A "phrase vector sequence" is a vector (phrase vector) defined in a language space (semantic space) according to the meaning of a phrase. A "phrase" is a single word or a sequence of multiple words (i.e., a phrase). To generate a phrase vector sequence, for example, a statistical estimation model described in Tomas Mikolov et al., "Efficient Estimation of Word Representations in Vector Space," arXiv:1301.3781 [cs.CL], 2013 (Word2Vec) or Quoc Le, Tomas Mikolov, "Distributed Representations of Sentences and Documents," CoRR, abs / 1405.4053, pp.1-9, 2014 (Doc2Vec) is used.

[0097] In a specific example (Aspect 4) of Aspect 3, when the second control data sequence is acquired, the word vector sequence is generated using a trained second generative model. According to the above aspect, the second control data sequence can be easily generated using the second generative model.

[0098] In a specific example (Aspect 5) of any of Aspects 1 to 4, the second control data string includes phoneme data representing phonemes that make up the text. According to the above aspects, the second control data string includes phoneme data representing phonemes that make up the text. Therefore, it is possible to generate an acoustic data string of musical instrument sounds whose acoustic characteristics reflect non-linguistic characteristics (e.g., characteristics in the time domain or frequency domain) related to the pronunciation of phonemes in the text.

[0099] In a specific example (aspect 6) of any one of aspects 1 to 5, the first control data and the second control data are obtained and the acoustic data is generated in each of a plurality of unit periods on the time axis.

[0100] An audio generation system according to one embodiment (embodiment 7) of the present disclosure includes a control data sequence acquisition unit that acquires a first control data sequence representing characteristics of a sequence of notes and a second control data sequence representing characteristics of text corresponding to the sequence of notes, and an audio data sequence generation unit that processes the first control data sequence and the second control data sequence using a trained first generation model to generate an audio data sequence representing the instrument sound of the sequence of notes having acoustic characteristics corresponding to the characteristics of the text represented by the second control data sequence.

[0101] A program according to one embodiment (embodiment 8) of the present disclosure causes a computer system to function as a control data sequence acquisition unit that acquires a first control data sequence representing characteristics of a sequence of notes and a second control data sequence representing characteristics of text corresponding to the sequence of notes, and an acoustic data sequence generation unit that processes the first control data sequence and the second control data sequence using a trained first generative model to generate an acoustic data sequence representing the instrument sound of the sequence of notes having acoustic characteristics corresponding to the characteristics of the text represented by the second control data sequence. [Explanation of symbols]

[0102] 100...information system, 10...sound generation system, 11...control device, 12...storage device, 13...communication device, 14...operation device, 15...sound emission device, 20...machine learning system, 21...control device, 22...storage device, 23...communication device, 30...control data sequence acquisition unit, 31...first generation unit, 32...second generation unit, 321...language analysis unit, 322, 326...information generation unit, 326...phoneme analysis unit, 33...acoustic data sequence generation unit, 34...signal generation unit, 41...training data acquisition unit, 42...learning processing unit.

Claims

1. obtaining a first control data sequence representing a characteristic of a sequence of notes and a second control data sequence representing a characteristic of a text corresponding to the sequence of notes; generating an acoustic data sequence representing the musical instrument sound of the musical note sequence having acoustic characteristics according to the features of the text represented by the second control data sequence by processing the first control data sequence and the second control data sequence using a trained first generative model; A computer system implemented method for generating sound.

2. The first generative model is a first training control data sequence representing characteristics of a reference sequence of notes, and a second training control data sequence representing characteristics of a text corresponding to the reference sequence of notes; a training audio data sequence representing the musical instrument sounds of the reference musical note sequence; This model is trained using training data containing The method of producing sound according to claim 1.

3. The second control data sequence includes a sequence of phrase vectors representing phrases included in the text. The sound generating method according to claim 1 or 2.

4. In acquiring the second control data sequence, the word vector sequence is generated by a trained second generative model. The method of producing sound according to claim 3.

5. The second control data string includes phoneme data representing phonemes that constitute the text.

5. The sound generating method according to claim 1.

6. In each of a plurality of unit periods on the time axis, acquiring the first control data and the second control data in the acquisition of the first control data sequence and the second control data sequence; and generating individual acoustic data in the generation of the acoustic data sequence.

6. The sound generating method according to claim 1.

7. a control data sequence acquisition unit that acquires a first control data sequence that represents a characteristic of a sequence of notes and a second control data sequence that represents a characteristic of a text corresponding to the sequence of notes; an audio data sequence generation unit that processes the first control data sequence and the second control data sequence using a trained first generative model to generate an audio data sequence representing the musical instrument sound of the musical note sequence having acoustic characteristics corresponding to the features of the text represented by the second control data sequence; 1. A sound generating system comprising:

8. a control data sequence acquisition unit that acquires a first control data sequence that represents characteristics of a sequence of notes and a second control data sequence that represents characteristics of a text corresponding to the sequence of notes; an acoustic data sequence generation unit that processes the first control data sequence and the second control data sequence using a trained first generative model to generate an acoustic data sequence representing the musical instrument sound of the musical note sequence having acoustic characteristics corresponding to the features of the text represented by the second control data sequence; A program that makes a computer system function as a

Citation Information

Patent Citations

  • Electronic musical instrument

    JP1994175658A

  • Score recognizing device

    JP1994332443A

  • Automatic accompaniment device

    JP1996095566A

  • Musical sound outputting device and recording medium therefor

    JP2000227794A