Information processing methods, information processing systems, and programs
A machine learning-based generative model processes acoustic signals using condition and acoustic tokens to simplify and enhance the execution of multiple acoustic tasks, addressing the complexity of existing techniques.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- YAMAHA CORP
- Filing Date
- 2024-10-24
- Publication Date
- 2026-05-12
AI Technical Summary
Existing acoustic signal processing techniques require complex configurations due to the use of mutually independent mechanisms for various processes, leading to inefficiencies and complications.
An information processing method utilizing a machine learning-trained generative model to process acoustic signals based on condition and acoustic tokens, allowing for a simple configuration to perform multiple acoustic processes such as instrument estimation, transcription, chord estimation, and beat estimation.
Enables efficient and simplified execution of diverse acoustic processes on musical sounds with reduced data complexity and error likelihood, utilizing a unified generative model for various tasks.
Smart Images

Figure 2026076630000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for processing acoustic signals.
Background Art
[0002] Various techniques for processing acoustic signals have been proposed conventionally. For example, Patent Document 1 discloses a technique for estimating (i.e., pitch detection) the time series of musical notes of singing voice represented by an acoustic signal by processing the acoustic signal. Also, for example, Patent Document 2 discloses a technique for estimating a musical instrument that produces a performance sound by processing an acoustic signal.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] Processes executed on acoustic signals, such as the techniques of Patent Document 1 or Patent Document 2, are very diverse. Therefore, there is a problem that the configuration required for the processes becomes complicated in a form where each of a plurality of processes on an acoustic signal is realized by a mutually independent mechanism. In view of the above circumstances, one aspect of the present disclosure aims to realize various processes on an acoustic signal with a simple configuration.
Means for Solving the Problems
[0005] To solve the above problems, an information processing method according to one aspect of the present disclosure acquires input data including a condition token that specifies one of a plurality of different processes and an acoustic token that represents the frequency characteristics of an acoustic signal, and processes the input data with a machine learning-trained generative model to generate output data that represents the result of performing the process specified by the condition token from among the plurality of processes on the acoustic signal represented by the acoustic token.
[0006] An information processing system according to one aspect of this disclosure includes an input data acquisition unit that acquires input data including a condition token that specifies one of a plurality of different processes and an acoustic token that represents the frequency characteristics of an acoustic signal, and an output data generation unit that processes the input data using a machine learning-based generative model to generate output data that represents the result of executing the process specified by the condition token from among the plurality of processes on the acoustic signal represented by the acoustic token.
[0007] A program according to one aspect of this disclosure causes a computer system to function as an input data acquisition unit that acquires input data including a condition token that specifies one of a plurality of different processes and an acoustic token that represents the frequency characteristics of an acoustic signal, and an output data generation unit that processes the input data using a machine learning-prepared generative model to generate output data that represents the result of executing the process specified by the condition token among the plurality of processes on the acoustic signal represented by the acoustic token. [Brief explanation of the drawing]
[0008] [Figure 1] This is a block diagram illustrating the configuration of the information processing system in the first embodiment. [Figure 2] This is a block diagram illustrating the functional configuration of an information processing system. [Figure 3] This is a schematic diagram of the input data. [Figure 4] This is an explanatory diagram for music tokens. [Figure 5] This is a specific example of output data B when the condition token specifies the instrument estimation process. [Figure 6] This is a concrete example of output data B when the condition token specifies the transcription process. [Figure 7] This is a specific example of output data B when the condition token specifies the code estimation process. [Figure 8] This is a specific example of output data B when the condition token specifies the estimation process. [Figure 9] This is a specific example of output data B when the condition token specifies the beat estimation process. [Figure 10] This is a flowchart of the output data generation process. [Figure 11] This is an explanatory diagram regarding machine learning for generative models. [Figure 12] This is a diagram illustrating the process of preparing acoustic tokens in the training data. [Figure 13] This is a flowchart of the learning process. [Figure 14] This is a schematic diagram of the output data in the second embodiment. [Figure 15] This is a block diagram illustrating the functional configuration of the information processing system in the third embodiment. [Figure 16] This is a flowchart of the process for generating a condition token in the fourth embodiment. [Figure 17] This is an example of the display on a modified display device. [Modes for carrying out the invention]
[0009] A: First Embodiment FIG. 1 is a block diagram illustrating the configuration of an information processing system 100 in the first embodiment. The information processing system 100 is a computer system that executes various processes (hereinafter referred to as "acoustic processing") on an acoustic signal S. The acoustic signal S is a time signal representing sounds (hereinafter referred to as "musical sounds") that constitute music. For example, the musical sounds represented by the acoustic signal S are a mixed sound of musical sounds emitted by a plurality of musical instruments during the performance of a piece of music. The information processing system 100 is realized by an information device such as a smartphone, a tablet terminal, or a personal computer, for example.
[0010] The information processing system 100 includes a control device 11, a storage device 12, an operation device 13, and a display device 14. The information processing system 100 can be realized as a single device or as a plurality of devices separately configured from each other.
[0011] The control device 11 is composed of one or more processors that control each element of the information processing system 100. For example, the control device 11 is composed of one or more types of processors such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an SPU (Sound Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).
[0012] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is composed of a known recording medium such as a magnetic recording medium or a semiconductor recording medium, for example. The storage device 12 may be composed of a combination of multiple types of recording media. Further, a portable recording medium detachable from the information processing system 100 or a recording medium (e.g., cloud storage) on which the control device 11 can perform writing or reading via a communication network may be used as the storage device 12.
[0013] The storage device 12 stores the acoustic signal S. For example, the acoustic signal S is stored in the storage device 12 as a music file in an arbitrary format such as the WAV format or the MP3 format.
[0014] The operation device 13 is an input device that receives instructions from the user. The operation device 13 is, for example, an operator that the user operates, or a touch panel that detects contact by the user. The display device 14 displays various images under the control of the control device 11. The display device 14 is composed of a display panel such as a liquid crystal panel or an organic EL (Electroluminescence) panel.
[0015] FIG. 2 is a block diagram illustrating a functional configuration of the information processing system 100. The control device 11 realizes a plurality of functions (input data acquisition unit 21, output data generation unit 22, output control unit 23) for processing the acoustic signal S by executing a program stored in the storage device 12.
[0016] The input data acquisition unit 21 generates input data A. FIG. 3 is a schematic diagram of the input data A. As illustrated in FIG. 3, the input data A includes a condition token X and a plurality of acoustic tokens Y. The condition token X is data that specifies an acoustic process (i.e., a task) to be executed on the acoustic signal S. Each of the plurality of acoustic tokens Y is data that represents the frequency characteristics of the acoustic signal S. As illustrated in FIG. 2, the input data acquisition unit 21 includes a condition token generation unit �1 and an acoustic token generation unit 32.
[0017] The condition token generation unit 31 generates the condition token X. The condition token X specifies any one of a plurality of different acoustic processes. Specifically, the condition token generation unit 31 generates a condition token X that specifies the acoustic process selected by the user by operating the operation device 13 among the plurality of acoustic processes.
[0018] The user can input a desired instruction P (prompt) by operating the control device 13. The instruction P is a string of characters that instructs the execution of a desired sound processing on the sound signal S. For example, the instruction P is expressed in natural language, such as "Execute the transcription process". The condition token generation unit 31 generates an embedding vector representing the sound processing instructed by the user via the instruction P, as a condition token X. For example, the condition token generation unit 31 generates the condition token X by processing the instruction P using an existing embedding model.
[0019] As explained above, condition token X specifies the sound processing selected by the user from among several different sound processing processes. The multiple sound processing processes that the user can select include, for example, instrument estimation processing, music transcription processing, chord estimation processing, key estimation processing, and rhythm estimation processing.
[0020] Instrument estimation is an acoustic process that estimates the instrument corresponding to the musical sound represented by an acoustic signal S. In other words, instrument estimation is an acoustic process that estimates the type of instrument (e.g., piano, drums, string instruments, etc.) corresponding to each of the multiple musical tones that make up the musical sound of the acoustic signal S. Users who wish to perform instrument estimation input an instruction P such as "Estimate the instruments included in the music."
[0021] The transcription process is an acoustic process that estimates the time series of notes (i.e., musical score) corresponding to the musical sounds represented by the acoustic signal S. Specifically, the transcription process estimates the time series of notes corresponding to the sounds played by the instrument specified by the user, among the musical sounds represented by the acoustic signal S. In other words, among multiple performance parts performed by different instruments, the time series of notes in the performance part of the instrument specified by the user is estimated. A user who desires the transcription process inputs an instruction statement P that includes the specification of the instrument to be transcribed, such as "Generate a piano score." The condition token X that specifies the transcription process includes the specification of the instrument selected by the user as the target of transcription.
[0022] The chord estimation process is an acoustic process that estimates the time series of chords (harmonies) contained in the musical sound represented by the acoustic signal S. Users who wish to perform the chord estimation process input an instruction P such as "Estimate the chords contained in the music." The key estimation process is an acoustic process that estimates the key corresponding to the musical sound represented by the acoustic signal S. Users who wish to perform the key estimation process input an instruction P such as "Estimate the key of the music."
[0023] The beat estimation process is an acoustic process that estimates the beats corresponding to the musical sounds represented by an acoustic signal S. A beat refers to a point in time that marks a musical division within a piece of music. For example, beats within a piece of music, such as beat points or bar lines, are estimated by the beat estimation process. Users who wish to perform the beat estimation process input an instruction P such as "Estimate the beats of the music."
[0024] The acoustic token generation unit 32 in Figure 2 generates a time series of multiple acoustic tokens Y by analyzing the acoustic signal S. As mentioned above, each acoustic token Y is data representing the frequency characteristics of the acoustic signal S.
[0025] Specifically, as illustrated in Figure 3, the acoustic token generation unit 32 generates a spectrogram F by analyzing the acoustic signal S. The spectrogram F is a time-frequency representation of the acoustic signal S. Specifically, the acoustic token generation unit 32 calculates the spectrogram F by performing a constant-Q transform (CQT) on the acoustic signal S. That is, the spectrogram F in the first embodiment is a constant-Q spectrogram in which the frequency axis is represented on a logarithmic scale.
[0026] The acoustic token generation unit 32 divides the spectrogram F of the acoustic signal S into multiple unit periods U on the time axis and generates a time series of multiple acoustic tokens Y representing different unit periods U. That is, an acoustic token Y is data that represents the frequency characteristics of the acoustic signal S, specifically the portion of the spectrogram F obtained by performing a constant Q transform on the acoustic signal S within one unit period U. The number of acoustic tokens Y constituting the input data A is a variable value depending on the time length of the acoustic signal S. Note that the number of acoustic tokens Y included in the input data A may be just one.
[0027] The time length of the unit period U is the time length corresponding to the beat period of the musical sound represented by the acoustic signal S (hereinafter referred to as the "reference length"). The beat period is the interval between each beat point that is sequential on the time axis (beat interval). The time length of the unit period U is set to an integer multiple or an integer fraction of the reference length. In the first embodiment, the reference length is set to a time length corresponding to 1 / 12 of a beat of the musical sound represented by the acoustic signal S. The acoustic token generation unit 32 estimates the beat period by analyzing the acoustic signal S and generates an acoustic token Y for each unit period U of the reference length corresponding to the beat period. As described above, the acoustic token Y is data that represents the frequency characteristics of the acoustic signal S for each unit period U of the reference length corresponding to the beat period of the musical sound.
[0028] As explained above, input data A includes a condition token X generated by the condition token generation unit 31 and multiple acoustic tokens Y generated by the acoustic token generation unit 32. As previously stated, the acoustic processing specified by condition token X is selected according to instructions from the user. Therefore, while the multiple acoustic tokens Y corresponding to the acoustic signal S are common, multiple types of input data A are generated in which the condition token X differs according to instructions from the user.
[0029] The output data generation unit 22 in Figure 2 generates output data B from input data A. Output data B is data representing the result of performing an acoustic processing specified by the condition token X of input data A on the acoustic signal S represented by multiple acoustic tokens Y of input data A. The output control unit 23 outputs the output data B generated by the output data generation unit 22 to the user. Specifically, the output control unit 23 displays an image representing output data B on the display device 14.
[0030] The output data generation unit 22 of the first embodiment generates output data B by processing input data A with a machine learning-prepared generation model M. The generation model M is a statistical model that has learned the relationship between input data A and output data B through prior machine learning. Specifically, the generation model M is realized by a combination of a program that causes the control device 11 to execute an operation to generate output data B from input data A, and a plurality of variables (e.g., bias and weight values) applied to the operation. The numerical values of each of the plurality of variables are set in advance by machine learning.
[0031] For example, a transformer, which is an encoder-decoder model including a self-attention mechanism (specifically a multi-head attention mechanism), can be used as the generative model M. Transformers are disclosed, for example, in Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, "Attention Is All You Need", 31st Conference on Neural Information Processing Systems (NIPS 2017). Alternatively, a type of transformer called MEGA (Moving Average Equipped Gated Attention) may be adopted as the generative model M. MEGA is disclosed, for example, in Ma, C. Zhou, X. Kong, J. He, L. Gui, G. Neubig, J. May, and L. Zettlemoyer, "Mega: moving average equipped gated attention", arXiv:2209.10655, 2022.
[0032] For processing by the transformer used as the generative model M, tokens such as the conditional token X and acoustic token Y in the input data A are preferred. Similarly, the output data B in the first embodiment is represented by tokens suitable for processing by the generative model M (transformer) (hereinafter referred to as "music token T"). Specifically, the output data B is represented, for example, by a time series of music token T in REMI (Revamped MIDI-derived Events) format.
[0033] Figure 4 is an explanatory diagram of musical tokens T. As illustrated in Figure 4, examples of musical tokens T include instrument token T1, note token T2, note value token T3, root token T4, type token T5, bass token T6, key token T7, measure token T8, beat token T9, and position token T10. In the following explanation, the symbol N represents a variable included in each musical token T.
[0034] The instrument token T1 specifies the type of instrument. The type of instrument in instrument token T1 is specified by a string such as "acp" for acoustic piano, "acg" for acoustic guitar, or "drums" for drum set.
[0035] Note token T2 is a musical token T that represents a musical note. Specifically, note token T2 specifies the type of instrument used to play the note and the pitch of the note (e.g., note number). The type of instrument in note token T2 is specified by a string such as "acp_N" for acoustic piano or "acg_N" for acoustic guitar, similar to how the type of instrument is specified by instrument token T1. The variable N in note token T2 is the note number representing the pitch of the note, and is set to one of several numerical values (N=0~127).
[0036] The note value token T3 (len_N) specifies the note value (duration) of a note. The variable N in the note value token T3 represents the duration of the note as a number of units corresponding to the time length of the beat period of the musical note. The unit of note value represented by the note value token T3 is the same as the base length of the unit period U, for example, the time length corresponding to 1 / 12 of a beat of the musical note. Therefore, for example, the note value token T3 "len_18" means a time length of 1.5 beats. The combination of the note token T2 and the note value token T3 represents a single musical note with specified instrument, pitch, and note value.
[0037] The root token T4, type token T5, and bass token T6 are musical tokens T that represent chords composed of multiple notes. The root token T4 (chord_N) specifies the root note of any chord. The variable N of the root token T4 is set to one of the 12 pitches that make up an octave (C, C#, D, D#, E, F, F#, G, G#, A, A#, B). The type token T5 (type_N) is a musical token T that represents the type of chord. The variable N of the type token T5 specifies one of several chord types such as "major", "M7", or "dim". The bass token T6 (bass_N) is a musical token T that represents the bass note in an on-bass chord. The variable N of the bass token T6 is set to one of the 12 pitches (C, C#, D, D#, E, F, F#, G, G#, A, A#, B), similar to the root token T4. A chord is represented by a combination of the root token T4 and the type token T5. An on-bass chord is represented by a combination of the root token T4, the type token T5, and the bass token T6.
[0038] The key token T7 (key_N) is a musical token T that represents the key of the song. The variable N of the key token T7 specifies one of the 12 pitches (C, C#, D, D#, E, F, F#, G, G#, A, A#, B) and a combination of major or minor.
[0039] The bar token T8 (bar) is a musical token T that represents a bar line (i.e., the first beat of each measure) within a piece of music. The beat token T9 (beat) is a musical token T that represents a beat within a piece of music. The position token T10 (pos_N) is a musical token T that represents a specific point in time within a piece of music. The variable N in the position token T10 represents an arbitrary time length, depending on the number of reference lengths. Specifically, the position token T10 (pos_N) represents a point in time when a time equivalent to N reference lengths has elapsed since the beat immediately preceding the position token T10 (beat token T9).
[0040] As explained above, the output data B (note value token T3, position token T10) expresses time within the music using a reference length unit similar to the temporal unit of the acoustic token Y. In other words, the temporal unit is common between the acoustic token Y and the music token T. Therefore, compared to a configuration where, for example, one of the input data A (acoustic token Y) and output data B uses a time length corresponding to the beat period as its unit, and the other uses a fixed time length unrelated to the beat period as its unit, the machine learning of the generative model M becomes more efficient. Furthermore, since the reference length in the first embodiment is a time length corresponding to the beat period, it is suitable for processing the acoustic signal S representing musical sounds.
[0041] Output data B is represented by the combinations of musical tokens T exemplified above. The combinations of musical tokens T used in output data B differ depending on the type of sound processing specified by the conditional token X of input data A. Specific examples of output data B for each sound processing are described below.
[0042] [Instrument Estimation Processing] Figure 5 shows a specific example of output data B generated when condition token X specifies instrument estimation processing. Output data B corresponding to instrument estimation processing consists of instrument token T1 that specifies the type of instrument used to play the musical sound of the acoustic signal S. The output control unit 23 displays the type of instrument specified by instrument token T1 on the display device 14. As described above, according to the first embodiment, by specifying instrument estimation processing with condition token X, the instrument corresponding to the musical sound represented by the acoustic signal S can be estimated. Note that output data B corresponding to instrument estimation processing does not include time-related musical tokens T (note value token T3, position token T10).
[0043] [Transcription Processing] Figure 6 shows a specific example of output data B generated when condition token X specifies a transcription process. Output data B corresponding to the transcription process represents the time series of notes corresponding to the sounds played by the instrument specified by the user in instruction P (acoustic piano in Figure 6) from the musical sounds represented by the acoustic signal S. Specifically, output data B is represented by note tokens T2 and note value tokens T3 representing each note, and beat tokens T9 or position tokens T10 representing the position of each note within the musical piece.
[0044] For example, the part of output data B in Figure 6 that says "beat_2 pos_6 acp_62 len_6" means that at the point (pos_6) six units of the reference length (i.e., half a beat) after the second beat (beat_2) of the piece, a note (acp_62) with a pitch of "62" to be played on an acoustic piano will be placed with a duration of six units of the reference length (i.e., half a beat) (len_6). The output control unit 23 displays the image (musical score) represented by output data B on the display device 14. In Figure 6, a musical staff is shown as an example of output data B, but a piano roll image represented by output data B may also be displayed on the display device 14.
[0045] As described above, according to the first embodiment, by specifying the transcription process using the conditional token X, it is possible to estimate the time series of notes corresponding to the musical sound represented by the acoustic signal S. In particular, in the first embodiment, by specifying an instrument using the conditional token X, it is possible to generate a musical score for that instrument from the acoustic signal S of a musical sound composed of the sounds of multiple instruments being played.
[0046] [Code estimation process] Figure 7 shows a specific example of output data B generated when condition token X specifies a chord estimation process. Output data B corresponding to the chord estimation process represents the time series of chords contained in the musical sound represented by the acoustic signal S. Specifically, output data B is represented by root note tokens T4 and type tokens T5 (and also bass tokens T6) that represent each chord of the musical sound, and beat point tokens T9 or position tokens T10 that represent the position of the chord in the musical piece.
[0047] For example, the part of output data B in Figure 7 that says "beat_1 chord_0 type_0" means that the chord from the first beat (beat_1) of the song onwards is "C" (Chord_0=C, type_0=major). Also, the part of output data B in Figure 7 that says "beat_3 chord_4 type_0 bass_7" means that the chord from the third beat (beat_3) of the song onwards is "E / G#" (Chord_4=E, type_0=major, bass_7=G#). The output control unit 23 displays the image represented by output data B on the display device 14 (not shown). As described above, according to the first embodiment, by specifying the chord estimation process with the condition token X, the time series of chords corresponding to the musical sound represented by the acoustic signal S can be estimated.
[0048] [Estimation Processing] Figure 8 shows a specific example of output data B generated when condition token X specifies key estimation processing. Output data B corresponding to key estimation processing represents the key of the musical sound represented by the acoustic signal S. Specifically, output data B is represented by a key token T7 representing the key of the musical sound and a beat token T9 or position token T10 representing the starting point of that key.
[0049] For example, the "beat_1 key_0" part in output data B in Figure 8 means that the key of the piece from the first beat onward is "C". The output control unit 23 displays the image represented by output data B on the display device 14 (not shown). As described above, according to the first embodiment, by specifying the key estimation process with condition token X, the key corresponding to the musical sound represented by the acoustic signal S can be estimated.
[0050] [Beat rate estimation process] Figure 9 shows a specific example of output data B generated when condition token X specifies a beat estimation process. Output data B corresponding to the key estimation process represents the beat corresponding to the musical tone represented by the acoustic signal S. Specifically, output data B is represented by a measure token T8, a beat token T9, and a position token T10.
[0051] For example, the "pos_12 bar" part in output data B in Figure 9 means that the bar line is located at a point (pos_12) that is 12 beats (1 beat) long from the previous beat. The output control unit 23 displays the image represented by output data B on the display device 14. As described above, according to the first embodiment, by specifying the beat estimation process with the condition token X, the beat corresponding to the musical sound represented by the acoustic signal S can be estimated.
[0052] As illustrated above, the multiple sound processing specified by condition token X includes time-dependent sound processing (instrument estimation processing) and time-dependent sound processing (transcription processing, chord estimation processing, key estimation processing, and beat estimation processing).
[0053] Figure 10 is a flowchart of the process by which the control device 11 generates output data B (hereinafter referred to as the "output data generation process"). For example, the output data generation process is initiated in response to instructions from the user to the operating device 13.
[0054] When the output data generation process begins, the control device 11 (input data acquisition unit 21) generates input data A (Sa1~Sa3). Specifically, the control device 11 (condition token generation unit 31) generates condition token X in response to instructions from the user to the operating device 13 (Sa1). The control device 11 (acoustic token generation unit 32) generates a time series of multiple acoustic tokens Y by analyzing the acoustic signal S. The control device 11 generates input data A including the condition token X and the multiple acoustic tokens Y (Sa3). Note that the order of generating condition token X (Sa1) and generating acoustic token Y (Sa2) may be reversed.
[0055] The control device 11 (output data generation unit 22) generates output data B by processing input data A using a machine learning-based generative model M (Sa4). Then, the control device 11 (output control unit 23) performs output processing according to the output data B (Sa5). For example, the control device 11 displays the image represented by output data B on the display device 14.
[0056] As described above, in the first embodiment, the generation model M processes input data A, which includes a condition token X that specifies one of several acoustic processing methods and an acoustic token Y that represents the frequency characteristics of the acoustic signal S, thereby generating output data B that represents the result of performing the acoustic processing on the acoustic signal S. In other words, a single generation model M is reused for multiple acoustic processing methods. Therefore, compared to a configuration that requires a separate generation model for each acoustic processing method applied to the acoustic signal S, a variety of acoustic processing methods for the acoustic signal S can be realized with a simple configuration. In particular, in the first embodiment, since the acoustic signal S represents musical sounds, a variety of acoustic processing methods related to musical sounds can be performed.
[0057] Furthermore, in the first embodiment, the acoustic processing selected by the user from among multiple acoustic processing methods is specified by the condition token X. Therefore, it is possible to generate the result (output data B) of executing the user's desired processing on the acoustic signal S.
[0058] Furthermore, in the first embodiment, the output data B is represented by a time series of REMI-formatted music tokens T. Therefore, compared to the form in which the output data B is represented in MIDI (Musical Instrument Digital Interface) format, the amount of data in the output data B is reduced, and as a result, the time required to generate the output data B can be shortened. In the MIDI format, which includes a note-on event that instructs the sounding of a note and a note-off event that instructs the muting of the note, there is a possibility of an error occurring in which either the note-on event or the note-off event is lost. In the first embodiment, since the duration of the note is specified by a single note value token T3, there is also the advantage that the aforementioned error that is problematic in the MIDI format is less likely to occur.
[0059] [Machine learning of generative models] Figure 11 is an explanatory diagram of the machine learning process for the generative model M. The machine learning process for the generative model M utilizes multiple training datasets D. These multiple training datasets D are stored in the memory device 12. Each of the multiple training datasets D consists of a pair of training input data At and training output data Bt.
[0060] The input data At for training data D includes a conditional token X and multiple acoustic tokens Y. The conditional token X is data that specifies the acoustic processing. The acoustic tokens Y are data that represent the frequency characteristics of the training acoustic signal S. The output data Bt for training data D is the correct value (label) that the generative model M should generate for the input data At of training data D. In other words, the output data Bt for training data D represents the result of performing the acoustic processing specified by the conditional token X of training data D on the acoustic signal S represented by the acoustic token Y of training data D.
[0061] Figure 12 is an explanatory diagram of the process for preparing acoustic tokens Y in each training data D. Figure 12 shows an example of a spectrogram Ft generated by a constant-Q transform on the training acoustic signal S. The control device 11 generates reference data R by dividing one spectrogram Ft into units U on the time axis. The control device 11 extracts a predetermined range on the frequency axis (hereinafter referred to as the "target range W") from the reference data R as acoustic token Y. Acoustic token Y is generated for each of the multiple cases in which the position of the target range W on the frequency axis is different.
[0062] In the spectrogram Ft generated by a constant-Q transform on the acoustic signal S, the frequency axis is on a logarithmic scale. Therefore, shifting the target range W on the frequency axis corresponds to changing the pitch. That is, for each of the multiple cases where the target range W is shifted, the acoustic tokens Y extracted from the spectrogram Ft represent frequency characteristics corresponding to different pitches. As described above, since the spectrogram Ft of the first embodiment is a constant-Q spectrogram generated by a constant-Q transform on the acoustic signal S, multiple acoustic tokens Y corresponding to different pitches can be easily generated by the simple process of shifting the target range W. In other words, compared to a form in which multiple acoustic tokens Y corresponding to different pitches are generated by, for example, pitch shifting the acoustic signal S, the processing load required to prepare a large number of input data At can be reduced.
[0063] As illustrated in Figure 11, the control device 11 functions as a learning processing unit 40. The learning processing unit 40 establishes a generative model M using machine learning with multiple training data D. Figure 13 is a flowchart of the process (hereinafter referred to as "learning process") in which the control device 11 (learning processing unit 40) establishes a generative model M using machine learning with multiple training data D.
[0064] When the learning process begins, the control device 11 (learning processing unit 40) selects one of the multiple training data D stored in the memory device 12 (hereinafter referred to as "selected training data D") (Sb1). As illustrated in Figure 11, the control device 11 generates output data B by processing the input data At of the selected training data D using a provisional generative model M (Sb2). The control device 11 calculates a loss function that represents the error between the output data B generated by the above procedure and the output data Bt of the selected training data D (Sb3). The control device 11 updates several variables of the provisional generative model M so that the loss function is reduced (ideally minimized) (Sb4).
[0065] The control device 11 determines whether a predetermined termination condition has been met (Sb5). The termination condition is, for example, that the loss function falls below a predetermined threshold, or that the amount of change in the loss function falls below a predetermined threshold. If the termination condition is not met (Sb5: NO), the control device 11 selects the unselected training data D stored in the memory device 12 as the new selected training data D (Sb1). That is, the process of updating multiple variables of the generative model M (Sb1~Sb4) is repeated until the termination condition is met (Sb5: YES). If the termination condition is met (Sb5: YES), the control device 11 terminates the learning process. The generative model M at the time the termination condition is met is determined to be the machine-learned generative model M.
[0066] As can be understood from the above explanation, the generative model M learns the relationship between input data At and output data Bt in multiple training data sets D. Therefore, the output data generation unit 22, which uses the machine-learned generative model M, outputs statistically valid output data B for unknown input data A, based on the latent relationship between input data At and output data Bt in multiple training data sets D.
[0067] B: Second Embodiment A second embodiment will now be described. For elements whose function is the same as in the first embodiment in each of the embodiments described below, the same reference numerals as in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.
[0068] In the first embodiment, an acoustic token Y is generated for each unit period U of a reference length corresponding to the beat period of the acoustic signal S. In the second embodiment, the acoustic token generation unit 32 divides the acoustic signal S into units U of a predetermined reference length that are unrelated to the musical sound of the acoustic signal S, and generates a time series of multiple acoustic tokens Y that represent different unit periods U. That is, the acoustic token Y in the second embodiment is data representing the frequency characteristics of the acoustic signal S for each unit period U of a predetermined reference length that is unrelated to the musical sound of the acoustic signal S. As described above, the reference length in the first embodiment is a variable length corresponding to the tempo of the musical sound represented by the acoustic signal S, whereas the reference length in the second embodiment is a fixed length that does not depend on the tempo of the musical sound. The reference length in the second embodiment is, for example, 20 milliseconds.
[0069] Figure 14 is a schematic diagram of output data B in the second embodiment. Figure 14 illustrates output data B generated when condition token X specifies beat estimation processing. In output data B of the second embodiment, the position token T10 in the first embodiment is replaced with time token T11. Time token T11(t_N) is data that represents the elapsed time from the beginning of output data B (i.e., a specific point in time within the music) in units of a reference length. Specifically, time token T11(t_N) means the point in time when a time equivalent to N units of the reference length has elapsed from the starting point of output data B. As described above, output data B of the second embodiment is data that represents the result of performing the acoustic processing specified by condition token X on the acoustic signal S, in units of a predetermined reference length.
[0070] The same effects as in the first embodiment are achieved in the second embodiment. Furthermore, in the output data B of the second embodiment, the time token T11, which relates to time, expresses time using the same reference length as the temporal unit of the acoustic token Y. In other words, the temporal unit is common between the acoustic token Y and the music token T. Therefore, in the second embodiment as in the first embodiment, the machine learning of the generative model M is made more efficient compared to a form in which the temporal units differ between the input data A (acoustic token Y) and the output data B.
[0071] C: Third Embodiment Figure 15 is a block diagram illustrating the functional configuration of the information processing system 100 in the third embodiment. The control device 11 executes a program stored in the storage device 12 to realize multiple functions for processing the acoustic signal S (first input data acquisition unit 51, first output data generation unit 52, second input data acquisition unit 53, second output data generation unit 54, and output control unit 55). The first input data acquisition unit 51 and the second input data acquisition unit 53 operate in the same manner as the input data acquisition unit 21 in each of the above embodiments, and the first output data generation unit 52 and the second output data generation unit 54 operate in the same manner as the output data generation unit 22 in each of the above embodiments.
[0072] The user inputs an instruction P, for example, "Perform transcription processing for one instrument in a piece of music," by operating the control device 13. The user's instruction P is achieved through a two-stage acoustic processing, which includes an instrument estimation process that estimates the instrument corresponding to the musical sound represented by the acoustic signal S, and a transcription process that estimates the time series of notes corresponding to the sound played by the instrument estimated by the instrument estimation process.
[0073] The first input data acquisition unit 51 generates input data A1. Input data A1 includes a condition token X1 and a plurality of acoustic tokens Y. Note that input data A1 is an example of "first input data", and condition token X1 is an example of "first condition token".
[0074] The multiple acoustic tokens Y, as in the forms described above, are data representing the frequency characteristics of the acoustic signal S. Specifically, each acoustic token Y represents a portion of the spectrogram F of the acoustic signal S that corresponds to one unit period U.
[0075] Condition token X1 specifies the sound processing that corresponds to instruction P from among several sound processing processes. Specifically, condition token X1 specifies the instrument estimation process. The sound processing specified by condition token X1 (for example, the instrument estimation process) is an example of the "first processing".
[0076] The first output data generation unit 52 processes the input data A1 using the generation model M to generate output data B1, which represents the result of performing instrument estimation processing on the acoustic signal S represented by multiple acoustic tokens Y. Specifically, output data B1 specifies the types of multiple instruments used to play the musical sound of the acoustic signal S, similar to the example in Figure 5. Note that output data B1 is an example of "first output data".
[0077] The second input data acquisition unit 53 in Figure 15 generates input data A2. Input data A2 includes a condition token X2 and a plurality of acoustic tokens Y. Note that input data A2 is an example of "second input data", and condition token X2 is an example of "second condition token".
[0078] The multiple acoustic tokens Y in input data A2 are the same multiple acoustic tokens Y contained in input data A1. In other words, the multiple acoustic tokens Y generated by the analysis of the acoustic signal S are shared by input data A1 and input data A2.
[0079] The condition token X2 of input data A2 specifies the sound processing corresponding to output data B1 from among multiple sound processing processes. Specifically, the condition token X2 specifies a transcription process targeting one of the multiple types of instruments represented by output data B1 (hereinafter referred to as the "specific instrument"). In other words, the transcription process specified by condition token X2 is a sound processing process that estimates the time series of notes corresponding to the performance sounds of the specific instrument estimated by the instrument estimation process, from among the musical sounds represented by the sound signal S. The sound processing specified by condition token X2 (for example, the transcription process) is an example of the "second process".
[0080] The second output data generation unit 54 processes the input data A2 using the generation model M to generate output data B2, which represents the result of performing a musical transcription process on the acoustic signal S represented by multiple acoustic tokens Y. Specifically, output data B2 represents the time series (i.e., musical score) of notes corresponding to the sounds played by a specific instrument among the musical sounds represented by the acoustic signal S. Note that output data B2 is an example of "second output data".
[0081] As described above, a common generation model M is used for both the processing of input data A1 by the first output data generation unit 52 (instrument estimation processing) and the processing of input data A2 by the second output data generation unit 54 (musical transcription processing).
[0082] The output control unit 55 in Figure 15 outputs the output data B2 generated by the second output data generation unit 54 to the user. Specifically, the output control unit 55 displays an image representing the output data B2 (for example, a musical score corresponding to the performance part of a specific instrument) on the display device 14, similar to the example in Figure 5.
[0083] The same effects as in the first embodiment are achieved in the third embodiment. In the third embodiment, output data B1 is generated by processing input data A1 using the generation model M, and output data B2 is generated by processing input data A2, which includes condition token X2 corresponding to output data B1, with the generation model M. In other words, the generation model M can be shared for acoustic processing (instrument estimation processing and transcription processing) that is executed sequentially on the acoustic signal S. In particular, in the third embodiment, since the instrument estimation processing (first processing) and the transcription processing (second processing) are executed sequentially, it is possible to generate a musical score corresponding to the acoustic signal S for an unknown instrument that corresponds to the musical sound represented by the acoustic signal S.
[0084] In the third embodiment, the process of generating sheet music for a specific instrument from an acoustic signal S is divided into an instrument estimation process that estimates the specific instrument from the acoustic signal S, and a transcription process that generates sheet music for the specific instrument. In other words, compared to a configuration in which a generative model M realizes the entire process of generating sheet music for a specific instrument from an acoustic signal S, the content of each acoustic process is simplified, and as a result, the machine learning of the generative model M can be made more efficient. As a result of the more efficient machine learning of the generative model M, the accuracy of the acoustic processing using the generative model M is improved. For example, the accuracy of instrument estimation by the instrument estimation process and the accuracy of time series estimation of musical notes by the transcription process can be improved.
[0085] Furthermore, in the third embodiment, since the instrument estimation process, which estimates a specific instrument from the acoustic signal S, and the music transcription process, which generates a musical score for the specific instrument, are executed sequentially, the output data B1 representing the result of the instrument estimation process can be edited. For example, the second input data acquisition unit 53 modifies the result of the instrument estimation process represented by the output data B1 in accordance with instructions from the user to the operating device 13. The second input data acquisition unit 53 generates a condition token X2 that specifies the music transcription process for the modified specific instrument in accordance with the user's instructions. As described above, according to the third embodiment, it is possible to intervene user operations between sequential acoustic processes (instrument estimation process and music transcription process). Note that the configuration of the second embodiment may also be applied to the third embodiment.
[0086] D: Fourth Embodiment In the first embodiment, a configuration was illustrated in which the sound processing selected by the user from among multiple sound processing methods is specified by a condition token X. In the fourth embodiment, the method of selecting the sound processing specified by the condition token X differs from that of the first embodiment.
[0087] Figure 16 is a flowchart of the process by which the control device 11 (condition token generation unit 31) of the fourth embodiment generates a condition token X. For example, in the generation of condition token X (Sa1) of the output data generation process illustrated in Figure 10, the process in Figure 16 is executed.
[0088] The control device 11 (condition token generation unit 31) calculates a feature quantity Z by analyzing the acoustic signal S (Sa11). The feature quantity Z is a numerical value that represents a feature related to the acoustic characteristics of the acoustic signal S. The control device 11 (condition token generation unit 31) selects one of several acoustic processing methods according to the feature quantity Z of the acoustic signal S (Sa12). The control device 11 (condition token generation unit 31) generates a condition token X that specifies the acoustic processing method selected according to the feature quantity Z (Sa13).
[0089] The combination of the type of feature quantity Z of the acoustic signal S and the acoustic processing to be selected is arbitrary. For example, in a configuration in which the total number of timbres contained in the acoustic signal S is calculated as feature quantity Z, the control device 11 selects instrument estimation processing if feature quantity Z exceeds a threshold, and selects music transcription processing if feature quantity Z falls below the threshold.
[0090] The same effects as in the first embodiment are achieved in the fourth embodiment. Furthermore, in the fourth embodiment, since the acoustic processing selected according to the feature quantity Z of the acoustic signal S is specified by the condition token X, the user does not need to select one of the multiple acoustic processing options. Therefore, the user's burden can be reduced. Note that the configuration of the second or third embodiment may also be applied to the fourth embodiment.
[0091] E: Variation The following are examples of specific modifications that may be added to each of the embodiments exemplified above. Two or more embodiments may be arbitrarily selected from the following examples and merged as appropriate, provided they do not contradict each other.
[0092] (1) In the above-described forms, examples were given in which each acoustic token Y of the input data A represents a constant-Q spectrogram obtained by constant-Q transforming the acoustic signal S. However, the form of the acoustic token Y is not limited to the above examples. For example, a form in which each acoustic token Y represents a time-frequency representation such as a Mels spectrogram or amplitude spectrogram of the acoustic signal S is also conceivable. The acoustic token Y is comprehensively represented as data that represents the frequency characteristics of the acoustic signal S.
[0093] (2) In the above-described embodiments, an example was given in which one input data A is generated for the entire acoustic signal S. However, input data A may also be generated for each of the multiple intervals obtained by dividing the acoustic signal S on the time axis. The input data A corresponding to one interval of the acoustic signal S includes a conditional token X that specifies the acoustic processing and a plurality of acoustic tokens Y that represent the frequency characteristics of the acoustic signal S for that interval.
[0094] (3) For example, a chat screen 70 as illustrated in Figure 17 may be displayed on the display device 14. The input message 71 in Figure 17 includes an instruction sentence P entered by the user in an operation on the operating device 13 and a link 72 to a music file (sound signal S) that the user has specified as the target of processing. The response message 73 to the input message 71 includes a string 74 indicating that processing in accordance with the user's instruction has been completed and a link 75 for the user to access the output data B. The pair of input message 71 and response message 73 is repeated in chronological order. Therefore, the user can obtain the result (output data B) of performing the desired sound processing on the desired sound signal S through interactive exchange of information.
[0095] (4) The content and types of sound processing are not limited to the examples of each form described above. For example, the following sound processing can also be realized by the generation model M. The sound effects exemplified below are suitable for use in situations such as DAW (Digital Audio Workstation) or mixing.
[0096] [a] A process for defining suitable channel names (labels) and channel colors for an acoustic signal S. [b] A process for selecting a suitable sound effect (effector) for an acoustic signal S. [c] The process of setting the parameters of an acoustic effect (e.g., an equalizer) to values suitable for the acoustic signal S. [d] Automation in mixing, such as dry mixing, which automatically sets the temporal changes of parameters related to the acoustic signal S. For example, the process of adjusting parameters so that the sound pressure of multiple sound sources remains constant. [e] The process of arranging the musical sound represented by the acoustic signal S. For example, arranging the musical sound to match a specific genre (e.g., jazz) or musical style (e.g., music box style), arranging by increasing or decreasing the number of notes in the musical sound, or arranging by adjusting to the range of a specific instrument. [f] A process to change the data format of an acoustic signal S. For example, a process to convert an acoustic signal S expressed in WAV or MP3 format into MIDI time-series data.
[0097] (5) In each of the above-described forms, a transformer was used as an example of the generative model M, but the configuration or type of the generative model M is arbitrary. For example, a deep neural network such as a recurrent neural network (RNN) or long short-term memory (LSTM) may be used as the generative model M. The generative model M may also be constructed by a combination of multiple types of statistical models.
[0098] (6) In each of the above-described embodiments, an embodiment in which the information processing system 100 is equipped with a learning processing unit 40 has been conveniently illustrated, but the learning processing unit 40 may be installed in a separate system (machine learning system) from the information processing system 100. The machine learning system is implemented by a server device such as a web server, and establishes a generative model M through the learning process described above. The generative model M established by the machine learning system is transferred to the information processing system 100.
[0099] (7) In the above-described forms, examples were given in which the acoustic signal S represents a musical sound, but the type of sound that the acoustic signal S represents is arbitrary. For example, the acoustic signal S may be a signal that represents speech. In a configuration in which the acoustic signal S represents speech, examples of acoustic processing specified by the condition token X include speech recognition processing, attribute estimation processing, or speech identification processing. Speech recognition processing is an acoustic processing that estimates a string of characters corresponding to the speech represented by the acoustic signal S.
[0100] Attribute estimation processing is an acoustic processing method that estimates the attributes of the speaker of the sound represented by the acoustic signal S. For example, attributes such as the speaker's gender (male / female) or age (child / adult / elderly) can be estimated. It is also possible to estimate the speaker's place of origin based on the dialect of the sound represented by the acoustic signal S.
[0101] Speech recognition processing is an acoustic processing method that estimates the type of sound represented by an acoustic signal S. For example, it estimates the type of sound represented by the acoustic signal S, such as human speech, animal sounds, or natural sounds like waves or wind.
[0102] (8) For example, the information processing system 100 in each of the above-described forms may be realized by a server device that communicates with an information device such as a smartphone or tablet terminal. The information processing system 100 receives an instruction statement P and an acoustic signal S from the information device via a communication network. The information processing system 100 generates output data B by operating in the same manner as in each of the above-described forms and transmits the output data B to the information device. Therefore, the output control unit 23 is omitted from the information processing system 100.
[0103] In the configuration in which the condition token generation unit 31 is mounted on the information device, the input data acquisition unit 21 of the information processing system 100 receives the condition token X from the information device. As illustrated above, the acquisition of the condition token X by the input data acquisition unit 21 includes the operation of generating the condition token X from the instruction statement P and the operation of receiving the condition token X from an external device such as an information device.
[0104] Similarly, in a configuration in which the acoustic token generation unit 32 is mounted on an information device, the input data acquisition unit 21 of the information processing system 100 receives the acoustic token Y from the information device. As illustrated above, the acquisition of the acoustic token Y by the input data acquisition unit 21 includes the operation of generating the acoustic token Y from an acoustic signal S and the operation of receiving the acoustic token Y from an external device such as an information device.
[0105] (9) The functions of the information processing system 100 exemplified above are realized through the cooperation of one or more processors constituting the control device 11 and a program stored in the storage device 12, as described above. The program according to this disclosure can be provided in a form stored on a computer-readable recording medium and installed on a computer. The recording medium is, for example, a non-transitory recording medium, such as an optical recording medium (optical disc) like a CD-ROM, but also includes any known form of recording medium such as a semiconductor recording medium or a magnetic recording medium. A non-transitory recording medium includes any recording medium except for transient propagation signals, and volatile recording media are not excluded. Furthermore, in a configuration in which a distribution device distributes a program via a communication network, the storage medium in which the distribution device stores the program corresponds to the non-transitory recording medium described above.
[0106] F: Note From the forms exemplified above, the following configuration can be understood, for example.
[0107] An information processing method according to one aspect of this disclosure (Aspect 1) acquires input data including a condition token that specifies one of a plurality of different processes and an acoustic token that represents the frequency characteristics of an acoustic signal, and processes the input data with a machine learning-prepared generative model to generate output data that represents the result of performing the process specified by the condition token from among the plurality of processes on the acoustic signal represented by the acoustic token. In the above aspect, by processing input data including a condition token that specifies one of a plurality of processes and an acoustic token that represents the frequency characteristics of an acoustic signal with a machine learning-prepared generative model, output data that represents the result of performing the acoustic processing on the acoustic signal is generated. In other words, the generative model is reused for multiple processes. Therefore, compared to a configuration that requires a separate generative model for each process on an acoustic signal, a variety of processes on an acoustic signal can be realized with a simple configuration.
[0108] A "(condition / acoustic) token" is a unit of data that is processed by a generative model. Input data is composed of a time series of multiple tokens, including condition tokens and acoustic tokens.
[0109] A "generative model" is a statistical model that learns, through prior machine learning, the relationship between training input data, which includes conditional tokens and acoustic tokens, and training output data, which represents the results of processing acoustic signals. For example, a transformer, which is an encoder-decoder including an attention mechanism, is an example of a generative model.
[0110] In the specific example of Embodiment 1 (Embodiment 2), the acoustic signal represents a musical sound. According to the above embodiments, a variety of musical processing can be performed on the acoustic signal representing a musical sound. Note that "musical sound" refers to the sounds that make up music (for example, a specific song). For example, the sounds played by one or more instruments, the sounds sung by a singer, or a mixture of instrumental and vocal sounds are examples of "musical sounds".
[0111] In a specific example of Embodiment 2 (Embodiment 3), the plurality of processes include an instrument estimation process that estimates the instrument corresponding to the musical sound represented by the acoustic signal. According to the above embodiments, by specifying the instrument estimation process using a conditional token, the instrument corresponding to the musical sound represented by the acoustic signal can be estimated.
[0112] In a specific example of Embodiment 2 or Embodiment 3 (Embodiment 4), the plurality of processes include a transcription process that estimates the time series of notes corresponding to the musical sounds represented by the acoustic signal. According to the above embodiments, the time series of notes corresponding to the musical sounds represented by the acoustic signal can be estimated by specifying the transcription process using a conditional token.
[0113] In a specific example of Embodiment 4 (Embodiment 5), the condition token includes the designation of an instrument, and the transcription process is a process of estimating the time series of notes corresponding to the sounds played by the instrument designated by the condition token from among the musical sounds represented by the acoustic signal. According to the above embodiments, by designating an instrument with a condition token, it is possible to generate a musical score for that instrument from an acoustic signal of a musical piece composed of the sounds played by multiple instruments.
[0114] In any specific example of Embodiments 2 to 5 (Embodiment 6), the plurality of processes include a chord estimation process that estimates the time series of chords contained in the musical sound represented by the acoustic signal. According to the above embodiments, by specifying the chord estimation process using a conditional token, the time series of chords corresponding to the musical sound represented by the acoustic signal can be estimated.
[0115] In any specific example of Embodiments 2 to 6 (Embodiment 7), the plurality of processes include a key estimation process that estimates the key corresponding to the musical sound represented by the acoustic signal. According to the above embodiments, the key estimation process can be specified by a conditional token to estimate the key (tonality, key) corresponding to the musical sound represented by the acoustic signal.
[0116] In any specific example of Embodiments 2 to 7 (Embodiment 8), the plurality of processes include a beat estimation process that estimates the beats corresponding to the musical sounds represented by the acoustic signal. According to the above embodiments, the beats corresponding to the musical sounds represented by the acoustic signal can be estimated by specifying the beat estimation process using a conditional token. Note that "beat" refers to a point in time that marks a musical division in a piece of music. For example, beats or the timing of bar lines are examples of "beats".
[0117] In any specific example of Embodiments 2 to 8 (Embodiment 9), the acoustic token represents the frequency characteristics of the acoustic signal for each unit period of a reference length corresponding to the beat period of the musical sound, and the output data represents the result of performing the processing specified by the condition token on the acoustic signal, in units of the reference length. In the above embodiments, the temporal units are common between the input data (acoustic token) and the output data. Therefore, compared to forms in which the temporal units differ between the input data and the output data, machine learning of the generative model is made more efficient. Furthermore, since the reference length in the first embodiment is a time length corresponding to the beat period, it is suitable for processing acoustic signals representing musical sounds. Note that "beat period" means the temporal interval between consecutive beats in a musical piece. "Reference length corresponding to the beat period" is a time length equivalent to an integer multiple or positive fraction of the beat period.
[0118] In any specific example of Embodiments 1 to 8 (Embodiment 10), the acoustic token represents the frequency characteristics of the acoustic signal for each unit period of a predetermined reference length unrelated to the musical sound, and the output data represents the result of performing the processing specified by the condition token on the acoustic signal, with the reference length as the unit. In the above embodiments, the temporal unit is common between the input data (acoustic token) and the output data. Therefore, compared to forms in which the temporal units differ between the input data and the output data, machine learning of the generative model is made more efficient.
[0119] In any specific example of Embodiments 1 to 10 (Embodiment 11), the acoustic token represents the frequency characteristics of the spectrogram obtained by performing a constant-Q transform on the acoustic signal. According to the above embodiments, the spectrogram (constant-Q spectrogram) generated by the constant-Q transform on the acoustic signal is used as the acoustic token. By changing the range on the frequency axis extracted as the acoustic token from the constant-Q spectrogram of the acoustic signal, multiple acoustic tokens corresponding to different pitches can be generated. Therefore, compared to a method that generates multiple acoustic tokens corresponding to different pitches by, for example, pitch shifting of the acoustic signal, the processing load required to prepare a large amount of input data for training can be reduced.
[0120] In any specific example of Embodiments 1 to 11 (Embodiment 12), when acquiring the input data, a condition token is generated that specifies the process selected by the user from among the multiple processes. In the above embodiments, the process selected by the user from among the multiple processes is specified by the condition token. Therefore, it is possible to generate the result of executing the user's desired process on the acoustic signal.
[0121] In any specific example of Embodiments 1 to 11 (Embodiment 13), when acquiring the input data, one of the multiple processes is selected according to the feature quantities of the acoustic signal, and a condition token is generated that specifies the process. In the above embodiments, since the process selected according to the feature quantities of the acoustic signal is specified by the condition token, the user does not need to select one of the multiple processes. Therefore, the user's burden can be reduced.
[0122] An information processing method according to another aspect of the present disclosure (Aspect 14) acquires first input data including a first condition token that specifies a first process from among a plurality of different processes and an acoustic token that represents the frequency characteristics of an acoustic signal representing a musical sound, processes the first input data using a generation model to generate first output data representing the result of performing the first process on the acoustic signal represented by the acoustic token, acquires second input data including a second condition token that specifies a second process corresponding to the first output data from among the plurality of processes and the acoustic token, processes the second input data using the generation model to generate second output data representing the result of performing the second process on the acoustic signal represented by the acoustic token. In the above aspect, the first output data is generated by processing the first input data using the generation model, and the second output data is generated by processing the second input data including a second condition token corresponding to the first output data using the generation model. That is, the generation model can be shared for processes (first process, second process) that are sequentially executed on the acoustic signal S.
[0123] In a specific example of Embodiment 14 (Embodiment 15), the first process is an instrument estimation process that estimates the instrument corresponding to the musical sound represented by the acoustic signal, and the second process is a transcription process that estimates the time series of notes corresponding to the sounds played by the instrument estimated by the first process from among the musical sounds represented by the acoustic signal. According to the above embodiment, since the instrument estimation process and the transcription process are executed sequentially, a musical score corresponding to the acoustic signal can be generated for a specific instrument corresponding to the musical sound represented by the acoustic signal.
[0124] An information processing system according to one aspect of the present disclosure (Aspect 16) comprises an input data acquisition unit that acquires input data including a condition token that specifies one of a plurality of different processes and an acoustic token that represents the frequency characteristics of an acoustic signal, and an output data generation unit that processes the input data using a machine learning-based generative model to generate output data that represents the result of executing the process specified by the condition token among the plurality of processes on the acoustic signal represented by the acoustic token.
[0125] A program according to one aspect of the present disclosure (Aspect 17) causes a computer system to function as an input data acquisition unit that acquires input data including a condition token that specifies one of a plurality of different processes and an acoustic token that represents the frequency characteristics of an acoustic signal, and an output data generation unit that processes the input data using a machine learning-prepared generative model to generate output data that represents the result of executing the process specified by the condition token among the plurality of processes on the acoustic signal represented by the acoustic token. [Explanation of Symbols]
[0126] 100... Information processing system, 11... Control device, 12... Storage device, 13... Operating device, 14... Display device, 21... Input data acquisition unit, 22... Output data generation unit, 23... Output control unit, 31... Condition token generation unit, 32... Acoustic token generation unit, 40... Learning processing unit, 51... First input data acquisition unit, 52... First output data generation unit, 53... Second input data acquisition unit, 54... Second output data generation unit, 55... Output control unit.
Claims
1. The system obtains input data that includes a conditional token specifying one of several different processes, and an acoustic token representing the frequency characteristics of an acoustic signal. By processing the input data using a machine learning-based generative model, output data is generated that represents the result of performing the process specified by the condition token from among the multiple processes on the acoustic signal represented by the acoustic token. Information processing methods implemented by computer systems.
2. The aforementioned acoustic signal represents a musical sound. The information processing method of claim 1.
3. The aforementioned plurality of processes include an instrument estimation process that estimates the instrument corresponding to the musical sound represented by the acoustic signal. The information processing method of claim 2.
4. The aforementioned plurality of processes include a transcription process that estimates the time series of musical notes corresponding to the musical sounds represented by the acoustic signal. The information processing method of claim 2.
5. The aforementioned condition token includes the designation of an instrument, The aforementioned transcription process is a process of estimating the time series of notes corresponding to the sounds played by the instrument specified by the condition token, among the musical sounds represented by the acoustic signal. The information processing method of claim 4.
6. The aforementioned plurality of processes include a chord estimation process that estimates the time series of chords contained in the musical sound represented by the acoustic signal. The information processing method of claim 2.
7. The aforementioned plurality of processes include a key estimation process that estimates the key corresponding to the musical sound represented by the acoustic signal. The information processing method of claim 2.
8. The aforementioned plurality of processes include a rhythm estimation process that estimates the rhythms corresponding to the musical sounds represented by the acoustic signal. The information processing method of claim 2.
9. The aforementioned acoustic token represents the frequency characteristics of the acoustic signal for each unit period of a reference length corresponding to the beat period of the musical sound, The output data represents the result of performing the process specified by the condition token on the acoustic signal, in units of the reference length. The information processing method of claim 2.
10. The aforementioned acoustic token represents the frequency characteristics of the acoustic signal for each unit period of a predetermined reference length, which is independent of the musical sound. The output data represents the result of performing the process specified by the condition token on the acoustic signal, in units of the reference length. The information processing method of claim 1.
11. The acoustic token represents the frequency characteristics of the spectrogram obtained by performing a constant Q transform on the acoustic signal. The information processing method of claim 1.
12. In acquiring the aforementioned input data, The user selects one of the aforementioned processes to generate the conditional token. The information processing method of claim 1.
13. In acquiring the aforementioned input data, Select one of the above-mentioned processes according to the characteristic quantities of the acoustic signal. Generate the condition token that specifies the process. The information processing method of claim 1.
14. First input data is obtained, which includes a first condition token that specifies the first of several different processes, and an acoustic token that represents the frequency characteristics of an acoustic signal representing a musical sound. By processing the first input data using the generation model, first output data is generated that represents the result of performing the first processing on the acoustic signal represented by the acoustic token. A second input data is obtained which includes a second condition token that specifies a second process corresponding to the first output data among the plurality of processes, and the acoustic token. By processing the second input data using the generation model, a second output data is generated that represents the result of performing the second processing on the acoustic signal represented by the acoustic token. Information processing methods implemented by computer systems.
15. The first process is an instrument estimation process that estimates the instrument corresponding to the musical sound represented by the acoustic signal, The second process is a transcription process that estimates the time series of musical notes corresponding to the sounds played by the instruments estimated by the first process, from among the musical sounds represented by the acoustic signal. The information processing method of claim 14.
16. An input data acquisition unit acquires input data that includes a condition token that specifies one of several different processes and an acoustic token that represents the frequency characteristics of an acoustic signal. An output data generation unit processes the input data using a machine learning-based generative model to generate output data that represents the result of performing a process specified by the condition token from among the multiple processes on the acoustic signal represented by the acoustic token. An information processing system equipped with the following features.
17. An input data acquisition unit that acquires input data including a condition token that specifies one of several different processes and an acoustic token that represents the frequency characteristics of an acoustic signal, and An output data generation unit processes the input data using a machine learning-based generative model to generate output data that represents the result of performing a process specified by the condition token from among the multiple processes on the acoustic signal represented by the acoustic token. A program that makes a computer system function.