Information processing system, information processing method, and information processing program
Patent Information
- Application Number
- JP2025563261
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2023-12-13
- Filing Date
- 2024-07-23
- Publication Date
- 2025-06-19
Abstract
Description
Information processing system, information processing method, and information processing program
[0001] The present disclosure relates to an information processing system, an information processing method, and an information processing program.
[0002] In recent years, technology that automatically generates transcription data from audio such as music has been attracting attention. This technology, also known as Automatic Music Transcription (AMT), is an important technology in music information processing.
[0003] A technique related to AMT is known that uses a machine learning model as training data, which is a set of annotation data corresponding to audio and audio data, and enables the automatic generation of highly accurate transcription data for instruments with an abundance of such training data (e.g., Non-Patent Documents 1 to 4).
[0004] J. Gardner et al. “MT3: multi-task multitrack music transcription”in Proceedings of ICLR, 2021S. Ian et al. “Scaling polyphonic transcription with mixtures of monophonic transcriptions” in Proceedings of ISMIR, 2022R. M. Bittner et al. “A lightweight instrument-agnostic model for polyphonic note transcription and multipitch estimation” in Proceedings of ICASSP, 2022Y.-T. Wu et al. “Multi-instrument automatic music transcription with self-attention-based instance segmentation”in Proceedings of ICML, 2022
[0005] However, conventional techniques have not been able to automatically generate highly accurate transcription data when such training data is scarce or unavailable.
[0006] Therefore, the present disclosure proposes an information processing system, an information processing method, and an information processing program that enable the automatic generation of versatile and highly accurate transcription data.
[0007] In order to solve the above problem, one form of information processing system according to the present disclosure includes an acquisition unit that acquires audio data to be processed, and a generation unit that generates corresponding transcription data by inputting the audio data acquired by the acquisition unit into a trained machine learning model that has been trained using audio sample data.
[0008] FIG. 1 is a diagram showing an overview of an information processing system according to an embodiment. FIG. 2 is a diagram showing an example of MIDI data according to an embodiment. FIG. 3 is a diagram showing an overview of application processing using a machine learning model according to an embodiment. FIG. 4 is a diagram showing an example of experimental results using a proposed method according to an embodiment. FIG. 5 is a diagram showing an example of the configuration of a user terminal according to an embodiment. FIG. 6 is a diagram showing an example of the configuration of an information processing device according to an embodiment. FIG. 7 is a diagram showing an example of a MIDI data storage unit according to an embodiment. FIG. 8 is a diagram showing an example of a sample data storage unit according to an embodiment. FIG. 9 is a diagram showing an example of an actual data storage unit according to an embodiment. FIG. 10 is a flowchart showing the flow of learning processing in a control unit. FIG. 11 is a flowchart showing the flow of application processing in a control unit. FIG. 12 is a hardware configuration diagram showing an example of a computer that realizes the functions of an information processing device.
[0009] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.
[0010] The present disclosure will be described in the following order: 1. Embodiment 1-1. Overview of information processing device according to embodiment 1-2. Overview of information processing system according to embodiment 1-3. Configuration of information processing device according to embodiment 2. Other embodiments 3. Effects of information processing device according to the present disclosure 4. Hardware configuration
[0011] (1. Embodiment) (1-1. Overview of Information Processing Device According to Embodiment) First, an overview of an information processing device 100 according to an embodiment will be described. The information processing device 100 is an information processing device that aims to enable the automatic generation of versatile and highly accurate music transcription data, and may be any device as long as it is capable of realizing the processing according to the embodiment. The information processing device 100 is realized by, for example, a server device or a cloud system, and executes the information processing according to the embodiment.
[0012] For example, the information processing device 100 uses a trained machine learning model trained using audio sample data. Also, for example, the information processing device 100 performs various data processes on audio data transmitted from an external information processing device (such as the user terminal 10 described below) and provides transcription data of the audio desired by the user.
[0013] Furthermore, the information processing device 100 uses audio sample data as audio data for training the machine learning model. For example, the information processing device 100 uses a plurality of audio sample data called "One Shot" in the MIDI (Musical Instruments Digital Interface) standard.
[0014] "One Shot" is predetermined audio data, and is audio data of sounds divided into certain units, such as the sound "do" or the sound "re." For example, it is audio data of individual sounds, such as the sound "do" or the sound "re." The information processing device 100 then uses synthesized data obtained by synthesizing such sample data as audio data for training the machine learning model.
[0015] The information processing device 100 performs learning using two types of methods as a learning method for the machine learning model. In the first method, the information processing device 100 uses MIDI data. The MIDI data includes note information related to the onset time, offset time, pitch speed, etc. of the note, and the information processing device 100 acquires this note information and synthesizes it with sample data by rendering it.
[0016] That is, the information processing device 100 performs additive synthesis of the sample data in accordance with the MIDI data. As a result, audio data corresponding to the MIDI data is generated. That is, a pair of audio data and MIDI data is generated. This pair of audio data and MIDI data becomes training data according to the embodiment. Then, the information processing device 100 uses the synthesized data, which is the synthesis result, to train a machine learning model.
[0017] In the second method, the information processing device 100 acquires real data and trains a machine learning model to minimize deviation from the real data. By training the machine learning model to minimize deviation, it becomes possible to automatically generate highly accurate transcription data even when the audio data to be processed is real audio data.
[0018] By using these two types of learning methods, it becomes possible to automatically generate general-purpose and highly accurate transcription data even when such training data is scarce or unavailable.It also makes it possible to automatically generate general-purpose and highly accurate transcription data even when applied to domains where such training data is scarce or unavailable (such as instruments (including playing techniques), timbres, singers (including singing styles), genres, and languages).
[0019] (1-2. Overview of Information Processing System According to Embodiment) Next, an overview of an information processing system for explaining the technology of the present disclosure will be described using FIG. 1. In the following embodiment, an example will be described in which the audio data when the machine learning model is applied is from a piano, and the transcription data of the audio desired by user U1 is piano keystroke information; however, the audio data does not have to be limited to that from an instrument such as a piano. For example, the audio data may be singing data or the like.
[0020] FIG. 1 is a diagram showing an overview of an information processing system according to an embodiment. In FIG. 1, the machine learning model according to the embodiment will be described as machine learning model M1. The machine learning model M1 has an encoder 20 and a decoder 30. The encoder 20 performs learning in cooperation with a discriminator 40. The encoder 20 receives as input a feature Xs obtained by encoding synthetic data D1 and a feature Xr obtained by encoding real data D2.
[0021] The discriminator 40 determines whether the output information from the encoder 20 is synthetic data D1 or real data D2 (determination H1). For example, the discriminator 40 has three fully connected layers. For example, section selection is performed on the output information from the encoder 20, and flattened information is input to the discriminator 40.
[0022] The learning process of the machine learning model M1 will now be described. The information processing device 100 acquires MIDI data F1 and sample data F2. The three bars contained in the MIDI data F1 are information related to the tone color, and are information that defines details for audio playback, such as which note the tone color is caused by, when the note starts and ends, and at what pitch speed the note is played. Note that the information contained in the MIDI data F1 is just an example, and in reality it will contain complex information for performance.
[0023] Here, the MIDI data F1 will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of MIDI data according to an embodiment. The MIDI data shown in Fig. 2 is data based on axes of pitch (corresponding to the keyboard of a piano) and time. Alphanumeric characters such as "F3," "B3," "D4," and "F4" on the pitch axis are information indicating musical intervals.
[0024] In Figure 2, notes N1 through N4 are arranged in parallel at the beginning of the data, indicating that multiple intervals are played simultaneously. Specifically, this indicates that the intervals "G3#," "B3," "D4," and "D5" corresponding to notes N1 through N4 are played simultaneously. For example, on a piano, this indicates that the keys corresponding to "G3#," "B3," "D4," and "D5" are played simultaneously. Note that in Figure 2, only the first four bars of the data are labeled, but this is for convenience of explanation; other bars included in the data may also be labeled as appropriate.
[0025] Returning to the explanation of FIG. 1, the sample data F2 includes sample data for multiple audio streams. For example, it includes sample data F21 to F23. The sample data F2 also includes multiple types of audio data, namely, main audio Z1 and sub audio Z2. The sample data F2 includes the main audio Z1 and sub audio Z2 corresponding to the sample data F21 to F23, respectively.
[0026] Note that the information included in the sample data F2 is merely an example, and in reality, more sample data is included. For example, although not shown, sample data F24, sample data F25, etc. may be included. Alternatively, the sample data F2 may include only the main audio Z1 (without the sub audio Z2). Alternatively, the sample data F2 may include audio data of a different type from the main audio Z1 and the sub audio Z2. For example, the sample data F2 may include audio data of the sub audio type in addition to the sub audio Z2.
[0027] Furthermore, sample data F2 may include sample data of a main audio type other than main audio Z1, or may include sample data of a sub audio type other than sub audio Z2. Furthermore, main audio Z1 may be made up of a set of sample data of multiple main audio types, and sub audio Z2 may be made up of a set of sample data of multiple sub audio types.
[0028] The information processing device 100 may combine multiple types of audio data to generate new "One Shot" audio data. For example, the information processing device 100 may combine multiple types of audio data, such as main audio Z1 and sub audio Z2, to generate new "One Shot" audio data.
[0029] For example, the information processing device 100 may weight each of the main audio Z1 and the sub audio Z2 and generate new "One Shot" audio data based on the weighting. For example, because the note "C" on a piano and the note "C" on a guitar are different even though they are both "C," the information processing device 100 may weight each of them to generate new "One Shot" audio data for the note "C." This enables the information processing device 100 to generate new "One Shot" audio data that incorporates elements of both piano and guitar. As a result, it becomes possible to generate composite data D1 with a variety of timbres.
[0030] Furthermore, in order to provide diversity, the information processing device 100 may generate the synthetic data D1 while changing the weighting so that the timbre changes slightly over time in accordance with the music. For example, the information processing device 100 may generate the synthetic data D1 using completely different musical instruments while changing the weighting for each instrument so that the timbre changes over time.
[0031] Furthermore, the information processing device 100 is not limited to combining two types of audio data, and when there are multiple types of main audio and multiple types of sub audio, the information processing device 100 may appropriately combine these multiple types of audio data to generate new "One Shot" audio data. For example, the information processing device 100 may combine main audio Z1, sub audio Z2, and sub audio Z3 to generate new "One Shot" audio data.
[0032] Furthermore, the information processing device 100 may generate new "One Shot" audio data by combining sample data of the same instrument or the like across different audio data types, or may generate new "One Shot" audio data by combining sample data of different instruments or the like across different audio data types. For example, the information processing device 100 may combine piano sounds with piano sounds (e.g., combining the sound of a high-end grand piano of brand "XX" with the sound of a less high-end jazz piano of brand "XX"), piano sounds with guitar sounds, or piano sounds with vocals.
[0033] The information processing device 100 renders the sample data F2 and the MIDI data F1. Specifically, the information contained in the sample data F2 and the information contained in the MIDI data F1 are rendered to generate the composite data D1. For example, when the information processing device 100 obtains information from the MIDI data F1 that the tone of the first bar is a piano "C" and lasts for one second from "0:01 to 0:02" (the time from striking the key to releasing it is one second), the information processing device 100 identifies the audio corresponding to the tone of that bar from the sample data F2 and renders it.
[0034] In this way, by rendering the sample data F2 and the MIDI data F1, the audio of the sample data F2 can be played back in accordance with the MIDI data F1. As a result, it becomes possible to automatically generate a set of the MIDI data F1 and audio data (composite data D1) obtained by rendering the sample data F2 and the MIDI data F1.
[0035] Note that the information processing device 100 may perform a predetermined editing process R1 (preprocessing for training the machine learning model M1) when generating the composite data D1. For example, the information processing device 100 may perform editing process R1 such as compression or limiting. This is because if the rendered audio is used as is, the resulting audio may be a linear addition of the sounds of the notes in the MIDI data F1.
[0036] For this reason, editing processing R1 may be performed, such as compressing the dynamics so that the volume remains somewhat constant when a single note is played and when multiple notes are played simultaneously within a track.Also, editing processing R1 may be performed, such as limiting based on a threshold value (cutting off any part that exceeds a predetermined threshold as the upper limit of amplitude) so that the volume remains somewhat constant.
[0037] The information processing device 100 acquires the composite data D1 and the actual data D2. At this time, the information processing device 100 may acquire, for example, actual data D2 corresponding to the composite data D1 (e.g., actual data D2 linked to MIDI data F1 and determined in advance), or actual data D2 according to the actual use of the AMT (e.g., actual data D2 similar to the audio data to be processed).
[0038] In the latter case, for example, the machine learning model M1 may be tuned for each user by training using actual data D2 corresponding to the actual AMT application for each user. Therefore, for example, the information processing device 100 may acquire actual data D2 from the user U1 that corresponds to (is similar to, for example) real audio data that the user U1 wants to transcribe. Furthermore, for example, if the data acquired from the user U1 as actual data D2 corresponding to the real audio data that the user U1 wants to transcribe is insufficient for training, the information processing device 100 may search for actual data D2 corresponding to that data and acquire it from an external information processing device, etc.
[0039] Then, the information processing device 100 obtains a feature Xs obtained by encoding the composite data D1 and a feature Xr obtained by encoding the actual data D2. Specifically, the information processing device 100 encodes the composite data D1 to convert it into a feature Xs, and encodes the actual data D2 to convert it into a feature Xr, thereby obtaining the feature Xs and the feature Xr.
[0040] In the prior art, when the feature Xs is input to the encoder 20, the encoder 20 learns to output MIDI data F1, which is the original data of the feature Xs. However, if the audio data to be processed is real audio data, the accuracy of the automatic generation of transcription data may be low.
[0041] The information processing device 100 inputs the acquired feature amounts (feature amount Xs and feature amount Xr) to the encoder 20. Output information from the encoder 20 is input to the discriminator 40. The discriminator 40 determines whether the output information corresponding to the synthetic data D1 and the output information corresponding to the actual data D2 is the synthetic data D1 or the actual data D2 (determination H1). Learning is then performed based on the output information from the discriminator 40.
[0042] Note that a loss function such as BCE Loss (Binary Cross Entropy Loss) is used for the learning. The information processing device 100 performs two types of learning using such loss functions. The two types of learning are "Classification Loss" L1 and "Adversarial Loss" L2. First, for "Classification Loss" L1, the information processing device 100 calculates a simple loss between the discriminator 40 and the correct answer data, and uses the calculation result to update the parameters so as to improve the performance of the discriminator 40. In this case, the information processing device 100 uses information obtained from the real data D2 as correct answer data and updates the parameters based on the loss between the information obtained from the real data D2 and the information obtained from the synthetic data D1.
[0043] In this way, in "Classification Loss" L1, the encoder 20 is fixed and training is performed on the discriminator 40. In FIG. 1 , BCE(0, C(E(Xs))) is information corresponding to the synthetic data D1, and BCE(1, C(E(Xr))) is information corresponding to the real data D2. Here, the numbers "0" and "1" in parentheses of BCE indicate that training is performed to approach "0" or "1." For example, BCE(0, C(E(Xs))) indicates that training is performed so that the information obtained from C(E(Xs)) approaches "0," and BCE(1, C(E(Xr))) indicates that training is performed so that the information obtained from C(E(Xr)) approaches "1."
[0044] The information processing device 100 learns so that the information obtained from the synthetic data D1 via the encoder 20 and the discriminator 40 approaches "0", and learns so that the information obtained from the real data D2 via the encoder 20 and the discriminator 40 approaches "1".
[0045] On the other hand, in "Adversarial Loss," the information processing device 100 performs adversarial learning. Specifically, the information processing device 100 performs learning so that it becomes impossible to distinguish between the synthetic data D1 and the real data D2. More specifically, the information processing device 100 performs learning so that the discrepancy between the information corresponding to the synthetic data D1 and the information corresponding to the real data D2 becomes smaller.
[0046] For such adversarial learning, the information processing device 100 updates the parameters by backpropagating the loss to the encoder 20. That is, the information processing device 100 updates the parameters by performing adversarial learning by backpropagating the loss to the encoder 20 so that it becomes difficult for the discriminator 40 to distinguish between the synthetic data D1 and the real data D2. In this way, in "Adversarial Loss" L2, contrary to "Classification Loss" L1, the discriminator 40 is fixed and learning is performed on the encoder 20.
[0047] 1, BCE(0.5, C(E(Xs))) is information corresponding to synthetic data D1, and BCE(0.5, C(E(Xr))) is information corresponding to real data D2. The information processing device 100 performs learning by changing the learning conditions (changing to 0.5, which is halfway between 0 and 1) from BCE(0, C(E(Xs))) to BCE(0.5, C(E(Xs))) and from BCE(1, C(E(Xr))) to BCE(0.5, C(E(Xr))) so that synthetic data D1 and real data D2 cannot be distinguished from each other.
[0048] The information processing device 100 learns so that the information obtained from the synthetic data D1 via the encoder 20 and the discriminator 40 approaches "0.5", and learns so that the information obtained from the actual data D2 via the encoder 20 and the discriminator 40 approaches "0.5".
[0049] Furthermore, in addition to or instead of learning to make it impossible to distinguish between the synthetic data D1 and the real data D2, the information processing device 100 may also perform learning to confuse the synthetic data D1 with the real data D2. The information processing device 100 may perform learning to confuse the synthetic data D1 with the real data D2 by changing the learning conditions from BCE(0, C(E(Xs))) to BCE(1, C(E(Xs))) or from BCE(1, C(E(Xr))) to BCE(0, C(E(Xr))), or by changing the learning conditions from BCE(0.5, C(E(Xs))) to BCE(1, C(E(Xs))) or from BCE(0.5, C(E(Xs))) to BCE(0, C(E(Xr))).
[0050] The information processing device 100 performs learning so that information obtained from the synthetic data D1 via the encoder 20 and the discriminator 40 approaches "1," and so that information obtained from the real data D2 via the encoder 20 and the discriminator 40 approaches "0." Hereinafter, the learning method that makes it impossible to distinguish between the synthetic data D1 and the real data D2 will be referred to as "Domain Confusion," and the learning method that causes the synthetic data D1 to be mistaken for the real data D2 will be referred to as "Domain Adaptation."
[0051] Through the two types of learning described above ("Classification Loss" L1 and "Adversarial Loss" L2), highly accurate transcription data can be automatically generated. In this way, by incorporating the discriminator 40 after the encoder 20 and performing learning so that the synthetic data D1 and the real data D2, which were identified as different data at input, are no longer identified as different data at output, it is possible to accurately and appropriately process audio data even when the audio data to be processed is real audio data.
[0052] The information processing device 100 then generates transcription data via the encoder 20 and the decoder 30, as in the prior art. The information processing device 100 generates the transcription data by generating output information F3 (such as note information relating to pitch, time, pitch speed, etc.) such as MIDI data F1 output from the decoder 30 via the encoder 20 and the decoder 30. At this time, the information processing device 100 uses the input information of the encoder 20 and the output information F3 of the decoder 30 to learn the "Transcription Loss" L3, as in the prior art, and performs optimization for generating the transcription data.
[0053] The output information F3 output from the decoder 30 may be transcription data. That is, the machine learning model M1 may output transcription data directly, or may output information for generating transcription data (such as MIDI data F1). In the latter case, the information processing device 100 generates transcription data based on the information output from the machine learning model M1 (such as MIDI data F1).
[0054] The learning process of the machine learning model M1 has been described above. Here, an application process using the trained machine learning model M1 will be described with reference to FIG. 3 . FIG. 3 is a diagram illustrating an example of an application process using a machine learning model according to an embodiment. Before describing FIG. 3 , the user terminal 10 will first be described.
[0055] The user terminal 10 is an information processing device used by a user who desires transcription data corresponding to a specific audio. By providing specific audio data, the user obtains transcription data for that audio through the information processing according to the embodiment. For example, in the case of a piano, keystroke information (e.g., data based on the keys and time) is obtained as transcription data. Note that audio data is a sound wave signal, e.g., data based on amplitude and time. The user terminal 10 may be any device capable of implementing the processing according to the embodiment. The user terminal 10 may also be a device such as a smartphone, a tablet terminal, a notebook PC, a desktop PC, a mobile phone, or a PDA. FIG. 3 illustrates a case where the user terminal 10 is a smartphone.
[0056] The user terminal 10 is, for example, a smart device such as a smartphone or tablet, and is a mobile terminal device that can communicate with any server device via a wireless communication network such as 4G to 5G (Generations) or LTE (Long Term Evolution). The user terminal 10 may have a screen such as a liquid crystal display with touch panel functionality, and may accept various operations on displayed data such as content, such as tapping, sliding, and scrolling, performed by the user using a finger or stylus. In FIG. 3, the user terminal 10 is used by user U1.
[0057] 3, the information processing device 100 acquires audio data to be processed (step S1). For example, the information processing device 100 acquires audio data provided by a user U1 via a user terminal 10 as the audio data to be processed.
[0058] Furthermore, for example, the information processing device 100 may acquire, as the audio data to be processed, audio data input, specified, or selected by the user U1 via the user terminal 10 using a predetermined application or on the web.
[0059] Furthermore, for example, the information processing device 100 may acquire audio data being played in real time by a user U1 or the like as the audio data to be processed. For example, the information processing device 100 may acquire audio data being played in real time by a user U1 or the like and provided in real time by the user U1 via the user terminal 10 as the audio data to be processed. In this way, the information processing device 100 may acquire audio data that is actual data as the audio data to be processed.
[0060] The information processing device 100 then inputs the acquired audio data into the trained machine learning model M1 to generate transcription data corresponding to the audio data (step S2), and provides the transcription data of the audio desired by the user U1 (step S3).
[0061] The audio data to be processed may be a whole piece of music (a certain collection of music), or only a certain part, section, or track of a piece of music, or only a certain instrument part of a piece of music (for example, only the part for piano playing), and is not particularly limited and may be any kind of audio data.
[0062] By using the information processing according to the above embodiment, it is possible to automatically generate versatile and highly accurate transcription data even when applied to a domain in which training data, which is a set of annotation data corresponding to audio (corresponding to MIDI data) and audio data, is scarce or unavailable. In other words, by using the information processing according to the above embodiment, it is possible to realize annotation-free AMT.
[0063] Furthermore, by using "One Shot" audio data, it is possible to create simple and scalable composite data D1. For example, composite data D1 can be created by simply cutting and pasting "One Shot" sample data from various options for MIDI data F1.
[0064] Furthermore, the audio data for "One Shot" may be real data (here, for example, data recorded from an actual performance) or synthetic data (here, for example, data generated by a synthesizer, etc.), so the real data "One Shot" may be rendered to generate synthetic data D1, the synthetic data "One Shot" may be rendered to generate synthetic data D1, or the real data "One Shot" and the synthetic data "One Shot" may be combined and rendered to generate synthetic data D1, thereby making it possible to create synthetic data D1 that corresponds to a variety of tones.
[0065] Furthermore, by using "One Shot" audio data, it becomes possible to create highly accurate synthetic data D1 that matches the music of the era of desktop music (DTM). Since DTM is a form of music commonly heard in everyday life, matching it to DTM makes it possible to create audio data that can reproduce familiar sounds. This also makes it possible to promote improvements in accuracy from a practical standpoint.
[0066] (Information Processing Variation 1: Release Time) In the above embodiment, release time may be taken into consideration. For example, in the case of a piano, release time refers to the reverberation that remains after the keys are released. In the case of MIDI data F1, this corresponds to the amount of time it takes for the sound to decay after the offset ends, and differs depending on the instrument. Considering this type of release time is generally important in AMT.
[0067] The information processing device 100 may perform rendering taking into account the release time of the audio. For example, the information processing device 100 may perform rendering by determining an audio release time and applying it to the MIDI data F1. For example, the information processing device 100 may determine a release time randomly within a certain range based on information about the instrument type included in the MIDI data F1, and may perform rendering by fading out the sample data by the length of the release time.
[0068] (Information processing variation 2: synthesizer) In the above embodiment, an example was described in which the information processing device 100 trains the machine learning model M1 using synthetic data D1 generated by rendering MIDI data F1 and sample data F2 and additively synthesizing the sample data F2 along with the MIDI data F1. However, the information processing device 100 may also train the machine learning model M1 using synthetic data D1 generated by a synthesizer or the like.
[0069] For example, the information processing device 100 may train the machine learning model M1 using synthetic data D1 generated from the MIDI data F1 by a synthesizer or the like. In this way, the information processing device 100 may train the machine learning model M1 using synthetic data D1 generated from the MIDI data F1 by a synthesizer or the like, instead of rendering the audio data of "One Shot" included in the sample data F2.
[0070] (Experimental Results) Below, experimental results using the proposed method according to the embodiment will be described. Fig. 4 is a diagram showing an example of experimental results using the proposed method according to the embodiment. Note that Method T1 is a result using a conventional method, and Method T3 is a result using the proposed method. Method T2 differs from the proposed method in that it uses annotations, but is a result using a method that is similar in other respects.
[0071] First, we will explain the items in the horizontal columns of the chart shown in Figure 4. The "a" in "Real Data" indicates whether learning was performed using real data, and items in vertical columns with a check mark indicate that learning was performed using real data.
[0072] The "b" in "Real Data" indicates whether learning was performed using real data and annotation data, and items in the vertical column with a check mark indicate that learning was performed using real data and annotation data.
[0073] The "c" in "Real Data" indicates whether or not learning has been performed using more data, and for items in the vertical column marked with a check mark, this indicates that learning has been performed using more data. In other words, the "b" and "c" in "Real Data" indicate whether or not learning has been performed using a set of audio data and annotation data as training data.
[0074] In addition, the horizontal item "Guitarset" corresponds to the result of testing whether or not realistic guitar sounds can be transcribed, "Maestro" corresponds to the result of testing whether or not realistic piano sounds can be transcribed, "Monila" corresponds to the result of testing whether or not realistic vocal sounds can be transcribed, "Phenicx" corresponds to the result of testing whether or not realistic orchestral performances can be transcribed, and "Slakh" corresponds to the result of testing whether or not various realistic synthetic sounds can be transcribed.
[0075] "Fno," "F," and "Acc" are indices that indicate the accuracy of note estimation. Among these, "Fno" is generally the most important indicator and is an indicator based on whether the onset is within 50 ms. "F" is an indicator based on whether the offset is within 50 ms. "Acc" is an indicator based on the frame level.
[0076] Based on "Fno", "F", and "Acc", the start and end of the note, the accuracy of the pitch, etc. are determined. In addition, the values of "Fno", "F", and "Acc" are calculated based on the output MIDI data and the real MIDI data.
[0077] Next, the items in the vertical columns of the chart shown in Fig. 4 will be explained. "Bittner et al.", "Wu et al.", "Gardner et al.", and "Simon et al." correspond to results using conventional methods, and correspond to the above-mentioned Non-Patent Documents 1 to 4, respectively. Furthermore, "Synthetic-DC," "Synthetic-DA," "Synthetic-L," "Synthetic-M," and "Synthetic-S" correspond to results using the proposed method according to the embodiment.
[0078] Of these, "Synthetic-DC" corresponds to the result when learning is performed so that synthetic data D1 cannot be distinguished from real data D2. In other words, it corresponds to the result when "Domain Confusion" is performed. Note that "DC" in "Synthetic-DC" is an abbreviation for "Domain Confusion." "Synthetic-DC" corresponds to the result when "Domain Confusion" is performed using synthetic data D1 generated using "One Shot" audio data.
[0079] Similarly, "Synthetic-DA" corresponds to the result of learning to mistake synthetic data D1 for real data D2. In other words, it corresponds to the result of performing "Domain Adaptation." Note that "DA" in "Synthetic-DA" is an abbreviation for "Domain Adaptation." "Synthetic-DA" corresponds to the result of performing "Domain Adaptation" using synthetic data D1 generated using "One Shot" audio data.
[0080] On the other hand, "Synthetic-L", "Synthetic-M", and "Synthetic-S" correspond to the results obtained when learning is performed using synthetic data D1 generated using "One Shot" audio data without performing "Domain Adaptation" or "Domain Confusion" (i.e., without performing the second learning step according to the above embodiment).
[0081] The "L," "M," and "S" in "Synthetic-L," "Synthetic-M," and "Synthetic-S" stand for "Large," "Medium," and "Small," respectively, and the data sizes used in the synthesized data D1 are different ("Synthetic-L" has the largest data size, and "Synthetic-S" has the smallest). For example, the data sizes of the MIDI data F1 and sample data F2 used in the synthesized data D1 are different.
[0082] "Synthetic-L" is synthetic data D1 generated using a larger number of sample data F2 (e.g., the largest number of sample data F2 in terms of the number and variety of "One Shots") to match more complex MIDI data F1. Because "Synthetic-L" has the largest data size, it is possible to express more detailed sounds and produce more realistic timbres. For this reason, experimental results show that it tends to have the highest accuracy among "Synthetic-L," "Synthetic-M," and "Synthetic-S."
[0083] Furthermore, "Real Mix" corresponds to the results of training a model similar to the proposed method according to the embodiment by mixing real audio data. Specifically, it corresponds to the results of training and evaluation using all of the above-mentioned "Guitarset," "Maestro," "Monila," "Phenicx," and "Slakh." For example, in the case of "Guitarset," it corresponds to the results of training using all of "Guitarset," "Maestro," "Monila," "Phenicx," and "Slakh" and evaluating "Guitarset."
[0084] Furthermore, "Real Omit" corresponds to the results of evaluation performed on "Real Mix" by excluding the dataset of the target domain. For example, in the case of "Guitarset," it corresponds to the results of evaluation of "Guitarset" after training using all of "Maestro," "Monila," "Phenicx," and "Slakh" except for "Guitarset." "Real Mix" and "Real Omit" differ from the proposed method according to the embodiment in that they are trained using annotation data. "Real Mix" and "Real Omit" are evaluation results for comparison with the proposed method according to the embodiment.
[0085] (1-3. Configuration of user terminal according to embodiment) Next, the configuration of the user terminal 10 according to the embodiment will be described using Fig. 5. Fig. 5 is a diagram showing an example configuration of the user terminal 10 according to the embodiment. As shown in Fig. 5, the user terminal has a communication unit 11, an input unit 12, an output unit 13, and a control unit 14.
[0086] The communication unit 11 is realized by, for example, a network interface card (NIC), etc. The communication unit 11 is connected to a predetermined network N by wire or wirelessly, and transmits and receives information to and from the information processing device 100, etc., via the predetermined network N.
[0087] The input unit 12 accepts various operations from a user. In FIG. 3, the input unit 12 accepts various operations from a user U1. For example, the input unit 12 may accept various operations from a user via a display screen using a touch panel function. The input unit 12 may also accept various operations from buttons provided on the user terminal 10 or a keyboard or mouse connected to the user terminal 10.
[0088] The output unit 13 is a display screen of a tablet terminal or the like realized by, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display, and is a display device for displaying various information. For example, the output unit 13 displays information transmitted from the information processing device 100.
[0089] The control unit 14 is, for example, a controller, and is realized by a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs stored in a storage device within the user terminal 10 using RAM (Random Access Memory) as a work area. For example, these various programs include application programs installed on the user terminal 10. For example, these various programs include an application program that converts audio played by the user into audio data. The control unit 14 is also realized by an integrated circuit, such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).
[0090] As shown in FIG. 5, the control unit 14 has a receiving unit 141 and a transmitting unit 142, and realizes or executes the information processing operations described below.
[0091] The receiving unit 141 receives various information from other information processing devices such as the information processing device 100. For example, the receiving unit 141 receives transcription data generated in response to predetermined audio data. For example, the receiving unit 141 receives transcription data generated in response to audio data provided by a user.
[0092] For example, the receiving unit 141 receives transcription data generated in response to audio data input, specified, or selected by the user. For example, the receiving unit 141 receives transcription data generated in response to audio data of audio being performed by the user. For example, the receiving unit 141 receives transcription data generated in response to audio data of audio being performed by someone other than the user.
[0093] The transmitting unit 142 transmits various information to other information processing devices such as the information processing device 100. For example, the transmitting unit 142 transmits information indicating that a user desires transcription data corresponding to predetermined audio. For example, the transmitting unit 142 transmits predetermined audio data. For example, the transmitting unit 142 transmits audio data input, specified, or selected by the user.
[0094] Further, for example, the transmitting unit 142 transmits information indicating that the user desires transcription data corresponding to the audio being played by the user. For example, the transmitting unit 142 transmits the audio data being played by the user. Further, for example, the transmitting unit 142 transmits information indicating that the user desires transcription data corresponding to the audio being played by someone other than the user. For example, the transmitting unit 142 transmits the audio data being played by someone other than the user.
[0095] (1-4. Configuration of Information Processing Apparatus According to Embodiment) Next, the configuration of the information processing apparatus 100 according to the embodiment will be described with reference to FIG. 6. FIG. 6 is a diagram showing an example of the configuration of the information processing apparatus 100 according to the embodiment. As shown in FIG. 6, the information processing apparatus 100 has a communication unit 110, a storage unit 120, and a control unit 130. Note that the information processing apparatus 100 may also have an input unit (e.g., a keyboard or a mouse) that accepts various operations from an administrator of the information processing apparatus 100, and a display unit (e.g., a liquid crystal display) that displays various information.
[0096] The communication unit 110 is realized by, for example, a NIC etc. The communication unit 110 is connected to the network N by wire or wirelessly, and transmits and receives information to and from the user terminal 10 etc. via the network N.
[0097] The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. As shown in Fig. 6, the storage unit 120 has a MIDI data storage unit 121, a sample data storage unit 122, and an actual data storage unit 123.
[0098] The MIDI data storage unit 121 stores information related to MIDI data (corresponding to MIDI data F1 in FIG. 1). An example of the MIDI data storage unit 121 according to the embodiment is shown in FIG. 7. As shown in FIG. 7, the MIDI data storage unit 121 has items such as "MIDI data ID" and "MIDI data."
[0099] "MIDI data ID" indicates identification information for identifying MIDI data. "MIDI data" indicates MIDI data. For example, information included in the MIDI data is stored in "MIDI data." For example, using the example in FIG. 2, information such as "note N1 (pitch: G3#, time: 2.5 seconds to 3.4 seconds), note N2 (pitch: B3, time: 2.5 seconds to 3.5 seconds), note N3 (pitch: D4, time: 2.5 seconds to 3.6 seconds), note N4 (pitch: D5, time: 2.5 seconds to 3.0 seconds)" is stored.
[0100] The sample data storage unit 122 stores information about sample data (corresponding to "One Shot" audio data in FIG. 1). FIG. 8 shows an example of the sample data storage unit 122 according to the embodiment. As shown in FIG. 8, the sample data storage unit 122 has items such as "sample data ID" and "sample data."
[0101] "Sample data ID" indicates identification information for identifying sample data. "Sample data" indicates sample data. In the example shown in FIG. 8, conceptual information such as "sample data #1" and "sample data #2" is stored in "sample data," but in reality, sound wave data (e.g., data based on amplitude and time) is stored. For example, using the example of FIG. 1, sound wave data corresponding to sample data F21 may be stored, or sound wave data corresponding to main audio Z1 of sample data F21 may be stored.
[0102] The actual data storage unit 123 stores information about actual data used for training the machine learning model M1 (corresponding to actual data D2 in FIG. 1). An example of the actual data storage unit 123 according to the embodiment is shown in FIG. 9. As shown in FIG. 9, the actual data storage unit 123 has fields such as "actual data ID" and "actual data."
[0103] "Actual data ID" indicates identification information for identifying actual data. "Actual data" indicates actual data. In the example shown in FIG. 9, conceptual information such as "actual data #1" and "actual data #2" is stored in "actual data", but in reality, sound wave data is stored. For example, using the example of FIG. 1, sound wave data corresponding to actual data D2 may be stored, or a feature Xr obtained by encoding may be stored.
[0104] The control unit 130 is a controller, and is realized by, for example, a CPU or an MPU executing various programs stored in a storage device inside the information processing device 100 using RAM as a work area. The control unit 130 is also realized by, for example, an integrated circuit such as an ASIC or an FPGA.
[0105] 6, the control unit 130 has an acquisition unit 131, a rendering unit 132, an editing unit 133, a conversion unit 134, a learning unit 135, and a generation unit 136, and realizes or executes the information processing functions described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in FIG. 6, and other configurations may be used as long as they perform the information processing described below.
[0106] The acquisition unit 131 acquires various types of information, such as MIDI data (corresponding to MIDI data F1 in FIG. 1) and sample data (corresponding to sample data F2 in FIG. 1).
[0107] Also, for example, the acquisition unit 131 acquires composite data (corresponding to composite data D1 in FIG. 1 ) and actual data (corresponding to actual data D2 in FIG. 1 ). For example, the acquisition unit 131 acquires actual data according to the actual use of the AMT.
[0108] For example, the acquisition unit 131 acquires audio data to be processed. For example, the acquisition unit 131 acquires information indicating that the user desires transcription data corresponding to the audio data to be processed.
[0109] The rendering unit 132 renders the MIDI data (corresponding to MIDI data F1 in FIG. 1 ) and sample data (corresponding to sample data F2 in FIG. 1 ) acquired by the acquisition unit 131. Specifically, the rendering unit 132 renders the MIDI data and the audio data of "One Shot."
[0110] 1, the rendering unit 132 renders MIDI data F1 and information included in sample data F2 that corresponds to the note information included in the MIDI data F1. Specifically, the rendering unit 132 renders sample data such as sample data F21 to F23 included in the sample data F2. In other words, the rendering unit 132 renders sample data that combines main audio Z1 and sub audio Z2.
[0111] Furthermore, for example, the rendering unit 132 may render the sample data F21 with the main audio Z1 or the sub audio Z2 of the sample data F21. That is, the rendering unit 132 may render the sample data F21 with either the main audio Z1 or the sub audio Z2.
[0112] Furthermore, for example, the rendering unit 132 generates composite data (corresponding to composite data D1 in FIG. 1 ) that is audio data corresponding to the MIDI data by rendering the MIDI data and the sample data. In this way, the rendering unit 132 generates a set of MIDI data and audio data corresponding to the MIDI data.
[0113] The editing unit 133 edits the composite data generated by the rendering unit 132 (corresponding to editing process R1 in FIG. 1 ). Specifically, the editing unit 133 performs preprocessing for training the machine learning model M1. For example, the editing unit 133 compresses the dynamics of the composite data so that the volume is constant. Furthermore, for example, the editing unit 133 limits the dynamics of the composite data based on a threshold. Then, the editing unit 133 uses the edited composite data as composite data for training the machine learning model M1 (corresponding to composite data D1 in FIG. 1 ).
[0114] The conversion unit 134 encodes the composite data (corresponding to composite data D1 in Figure 1) acquired by the acquisition unit 131, generated by the rendering unit 132, or edited by the editing unit 133, and converts it into a feature (corresponding to feature Xs in Figure 1).
[0115] The conversion unit 134 also encodes the actual data (corresponding to the actual data D2 in FIG. 1) acquired by the acquisition unit 131 and converts it into a feature (corresponding to the feature Xr in FIG. 1).
[0116] The learning unit 135 trains the machine learning model M1. Specifically, the learning unit 135 trains the machine learning model M1 based on two types of feature amounts converted by the conversion unit 134. For example, the learning unit 135 trains the machine learning model M1 based on synthesized data (corresponding to synthesized data D1 in FIG. 1 ) obtained by combining multiple pieces of audio sample data (corresponding to the audio data of "One Shot" in FIG. 1 ). The learning unit 135 performs the following two types of learning processing.
[0117] As a first learning process by the learning unit 135, for example, the learning unit 135 trains the machine learning model M1 based on composite data (corresponding to composite data D1 in Figure 1) obtained by combining multiple corresponding audio sample data (corresponding to the audio data of "One Shot" in Figure 1) in accordance with MIDI data (corresponding to MIDI data F1 in Figure 1).
[0118] For example, the learning unit 135 trains the machine learning model M1 based on synthetic data (corresponding to synthetic data D1 in FIG. 1) obtained by synthesizing sample data (corresponding to sample data F21 to F23, etc. in FIG. 1) in which each sample data includes at least the types of main audio (corresponding to main audio Z1 in FIG. 1) and sub-audio (corresponding to sub-audio Z2 in FIG. 1).
[0119] Furthermore, for example, the learning unit 135 trains the machine learning model M1 based on synthetic data (corresponding to synthetic data D1 in FIG. 1) obtained by synthesizing sample data (corresponding to sample data such as main audio Z1 of sample data F21 and sub-audio Z2 of sample data F21 in FIG. 1), each of which is either sample data of the main audio type (corresponding to main audio Z1 in FIG. 1) or sample data of the sub-audio type (corresponding to sub-audio Z2 in FIG. 1).
[0120] Using the example of Figure 1, for example, the learning unit 135 trains the machine learning model M1 based on composite data D1 obtained by combining sample data of main audio Z1 of sample data F21 and sample data of sub-audio Z2 of sample data F22.
[0121] Furthermore, for example, the learning unit 135 trains the machine learning model M1 based on synthetic data (corresponding to synthetic data D1 in FIG. 1 ) obtained by combining sample data weighted for each type. For example, the learning unit 135 trains the machine learning model M1 based on synthetic data (corresponding to synthetic data D1 in FIG. 1 ) obtained by combining sample data weighted for each type depending on whether the sample data is main audio (corresponding to main audio Z1 in FIG. 1 ) or sub audio (corresponding to sub audio Z2 in FIG. 1 ).
[0122] Furthermore, for example, the learning unit 135 trains the machine learning model M1 based on synthetic data (corresponding to synthetic data D1 in FIG. 1 ) obtained by synthesizing sample data in which different instruments are used for each type of sample. For example, the learning unit 135 trains the machine learning model M1 based on synthetic data (corresponding to synthetic data D1 in FIG. 1 ) obtained by synthesizing sample data in which different instruments are used for main audio (corresponding to main audio Z1 in FIG. 1 ) and sub-audio (corresponding to sub-audio Z2 in FIG. 1 ) for each type of sample.
[0123] Furthermore, for example, the learning unit 135 trains the machine learning model M1 based on the composite data to which a predetermined editing process has been applied (corresponding to the composite data D1 in FIG. 1 ). For example, the learning unit 135 trains the machine learning model M1 based on the composite data edited by the editing unit 133 (corresponding to the composite data D1 in FIG. 1 ).
[0124] As a second learning process by the learning unit 135, for example, the learning unit 135 trains the machine learning model M1 so as to reduce the deviation between the synthetic data (corresponding to synthetic data D1 in Figure 1) and the actual data (corresponding to actual data D2 in Figure 1).
[0125] For example, the learning unit 135 trains a machine learning model M1 that includes at least an encoder (encoder 20 in FIG. 1), a decoder (decoder 30 in FIG. 1), and a discriminator (discriminator 40 in FIG. 1).
[0126] Furthermore, for example, the learning unit 135 trains the machine learning model M1 using a predetermined loss function (such as BCE Loss). For example, the learning unit 135 trains the machine learning model M1 using "Classification Loss" ("Classification Loss" L1 in FIG. 1).
[0127] For example, when a pair of synthetic data (corresponding to synthetic data D1 in FIG. 1) and real data (corresponding to real data D2 in FIG. 1) is input to an encoder (encoder 20 in FIG. 1), the learning unit 135 trains the discriminator (discriminator 40 in FIG. 1) so that the discriminator can distinguish between the synthetic data (corresponding to synthetic data D1 in FIG. 1) and the real data (corresponding to real data D2 in FIG. 1).
[0128] In this way, in "Classification Loss" ("Classification Loss" L1 in Figure 1), the learning unit 135 regards the difference between the result obtained from synthetic data (corresponding to synthetic data D1 in Figure 1) via an encoder (encoder 20 in Figure 1) and the result obtained from real data (corresponding to real data D2 in Figure 1) via an encoder (encoder 20 in Figure 1) as an error, and trains the machine learning model M1 so as to maximize this error (loss).
[0129] The learning of the machine learning model M1 using such a "Classification Loss" ("Classification Loss" L1 in FIG. 1) will be referred to as "first learning" below as appropriate.
[0130] Furthermore, for example, the learning unit 135 trains the machine learning model M1 using "Adversarial Loss" ("Adversarial Loss" L2 in FIG. 1). For example, when a pair of synthetic data (corresponding to synthetic data D1 in FIG. 1) and real data (corresponding to real data D2 in FIG. 1) is input to the encoder (encoder 20 in FIG. 1), the learning unit 135 trains the encoder (encoder 20 in FIG. 1) so that the discriminator (discriminator 40 in FIG. 1) cannot distinguish between the synthetic data (corresponding to synthetic data D1 in FIG. 1) and the real data (corresponding to real data D2 in FIG. 1).
[0131] That is, for example, after the first learning, the learning unit 135 trains the encoder (encoder 20 in FIG. 1 ) in an antagonistic manner to the first learning so that the discriminator (discriminator 40 in FIG. 1 ) is no longer able to distinguish between synthetic data (corresponding to synthetic data D1 in FIG. 1 ) and real data (corresponding to real data D2 in FIG. 1 ).
[0132] In such an "Adversarial Loss" ("Adversarial Loss" L2 in FIG. 1), the learning unit 135 regards the difference between the result obtained from the synthetic data (corresponding to the synthetic data D1 in FIG. 1) via the encoder (encoder 20 in FIG. 1) and the result obtained from the real data (corresponding to the real data D2 in FIG. 1) via the encoder (encoder 20 in FIG. 1) as an error, and trains the machine learning model M1 so as to minimize this error.
[0133] The learning of the machine learning model M1 using such an "Adversarial Loss" ("Adversarial Loss" L2 in FIG. 1) will be referred to as "second learning" below as appropriate.
[0134] Also, for example, when a pair of synthetic data (corresponding to synthetic data D1 in FIG. 1) and real data (corresponding to real data D2 in FIG. 1) is input to the encoder (encoder 20 in FIG. 1), the learning unit 135 trains the encoder (encoder 20 in FIG. 1) so that the discriminator (discriminator 40 in FIG. 1) mistakes the synthetic data (corresponding to synthetic data D1 in FIG. 1) for real data (corresponding to real data D2 in FIG. 1).
[0135] That is, for example, after the first learning, the learning unit 135 trains the encoder (encoder 20 in FIG. 1) in an antagonistic manner to the first learning so that the discriminator (discriminator 40 in FIG. 1) mistakes the synthetic data (corresponding to synthetic data D1 in FIG. 1) for the real data (corresponding to real data D2 in FIG. 1).
[0136] In such an "Adversarial Loss" ("Adversarial Loss" L2 in FIG. 1), the learning unit 135 regards the difference between the result obtained from the synthetic data (corresponding to the synthetic data D1 in FIG. 1) via the encoder (encoder 20 in FIG. 1) and the result obtained from the real data (corresponding to the real data D2 in FIG. 1) via the encoder (encoder 20 in FIG. 1) as an error, and trains the machine learning model M1 so as to maximize this error in the reverse direction.
[0137] The learning of the machine learning model M1 using such an "Adversarial Loss" ("Adversarial Loss" L2 in FIG. 1) will be referred to as "third learning" below as appropriate.
[0138] The generation unit 136 generates corresponding transcription data by inputting the audio data to be processed acquired by the acquisition unit 131 to the machine learning model M1. For example, the generation unit 136 generates corresponding transcription data by inputting the audio data to be processed acquired by the acquisition unit 131 to the trained machine learning model M1 that has been trained using audio sample data.
[0139] Next, the processing of each unit constituting the information processing device 100 will be described in detail along the flow with reference to Figures 10 and 11. Figure 10 is a flowchart showing the flow of the learning processing in the control unit 130.
[0140] The information processing device 100 determines whether or not MIDI data has been acquired (step S11).
[0141] When it is determined that MIDI data has been acquired (step S11; YES), the information processing apparatus 100 generates synthesized data by rendering the sample data in accordance with the acquired MIDI data (step S12).
[0142] On the other hand, when it is determined that MIDI data has not been acquired (step S11; NO), the information processing apparatus 100 waits until MIDI data is acquired.
[0143] The information processing device 100 acquires the composite data and the actual data (step S13). For example, the information processing device 100 acquires the composite data generated in step S12. Then, the information processing device 100 converts the acquired composite data and actual data into feature quantities (step S14).
[0144] The information processing device 100 inputs the feature amounts of the synthetic data and the feature amounts of the real data to the encoder, and further performs first learning via the discriminator (step S15).
[0145] The information processing device 100 inputs the feature amounts of the synthetic data and the feature amounts of the real data to the encoder, and further performs second learning or third learning via the discriminator (step S16).
[0146] Next, the information processing device 100 performs application processing using the trained machine learning model M1. This processing will be described with reference to Fig. 11. Fig. 11 is a flowchart showing the flow of the application processing in the control unit 130.
[0147] The information processing device 100 determines whether or not audio data to be processed has been acquired (step S21).
[0148] When the information processing device 100 determines that it has acquired the audio data to be processed (step S21; YES), it inputs the acquired audio data to be processed into a predetermined machine learning model (corresponding to the trained machine learning model M1) (step S22) to generate transcription data corresponding to the audio data to be processed.
[0149] On the other hand, when it is determined that the audio data to be processed has not been acquired (step S21; NO), the information processing device 100 waits until the audio data to be processed is acquired.
[0150] Then, the information processing device 100 provides the generated transcription data (step S23).
[0151] (2. Other Embodiments) The processing according to each of the above-described embodiments may be implemented in various different forms other than the above-described embodiments.
[0152] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. Furthermore, the information, including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings, can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.
[0153] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0154] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.
[0155] Furthermore, the effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0156] (3. Effects of the Information Processing System According to the Present Disclosure) As described above, the information processing system according to the present disclosure (information processing system 1 in the embodiment) includes an acquisition unit (acquisition unit 131 in the embodiment) and a generation unit (generation unit 136 in the embodiment). The acquisition unit acquires audio data to be processed. The generation unit generates corresponding transcription data by inputting the audio data acquired by the acquisition unit into a trained machine learning model trained using audio sample data.
[0157] In this way, the information processing system according to the present disclosure can enable the automatic generation of versatile and highly accurate transcription data corresponding to the audio data to be processed.
[0158] The information processing system further includes a learning unit (in this embodiment, the learning unit 135). The learning unit trains the machine learning model based on synthesized data obtained by synthesizing a plurality of audio sample data.
[0159] In this way, the information processing system can use multiple audio sample data to enable training of a machine learning model for automatic generation of transcription data.
[0160] The learning unit also trains the machine learning model based on synthesized data obtained by synthesizing multiple pieces of audio sample data corresponding to the MIDI data.
[0161] In this way, the information processing system can generate synthetic data from multiple audio sample data using MIDI data, thereby enabling the training of a machine learning model for the automatic generation of transcription data.
[0162] In addition, the learning unit trains the machine learning model so as to reduce the deviation between the synthetic data and the real data.
[0163] In this way, the information processing system can automatically generate versatile and highly accurate transcription data regardless of whether the audio data to be processed is synthesized data or real data.
[0164] The learning unit also trains a machine learning model including at least an encoder, a decoder, and a discriminator.
[0165] In this way, by incorporating a discriminator in addition to an encoder and decoder, the information processing system can enable learning of a machine learning model that reduces the deviation between synthetic data and real data.
[0166] Furthermore, when a set of synthetic data and real data is input to the encoder, the learning unit performs first learning to train the discriminator so that the discriminator can distinguish between the synthetic data and the real data.
[0167] In this way, the information processing system can appropriately identify whether the input information input to the encoder is synthetic data or real data.
[0168] Furthermore, after the first learning, the learning unit performs second learning, which is antagonistic to the first learning, to train the encoder so that the discriminator cannot distinguish between synthetic data and real data.
[0169] In this way, the information processing system trains the encoder so that the discriminator cannot distinguish between synthetic data and real data, thereby enabling the automatic generation of versatile and highly accurate transcription data regardless of whether the audio data to be processed is synthetic data or real data.
[0170] Furthermore, after the first learning, the learning unit performs third learning, which is antagonistic to the first learning, to train the encoder so that the discriminator mistakes the synthetic data for the real data.
[0171] In this way, the information processing system can enable the automatic generation of versatile and highly accurate transcription data whether the audio data to be processed is synthetic data or real data by training the encoder so that the discriminator will not mistake synthetic data for real data. Furthermore, if the information processing system also performs machine learning to train the encoder so that the discriminator will not be able to distinguish between synthetic data and real data, the automatic generation of even more versatile and highly accurate transcription data can be enabled.
[0172] The learning unit also trains the machine learning model using a loss function of BCE Loss (Binary Cross Entropy Loss).
[0173] In this way, by using the BCE Loss loss function, the information processing system can improve the accuracy of identifying whether the data is synthetic or real, and can also enable the automatic generation of versatile and highly accurate transcription data regardless of whether the audio data to be processed is synthetic or real.
[0174] The learning unit also trains the machine learning model based on composite data obtained by combining sample data in which each sample data includes at least the main audio and sub audio types.
[0175] In this way, the information processing system generates synthetic data from sample data containing multiple types of audio data, making it possible to automatically generate versatile and highly accurate transcription data even when the audio data to be processed is complex music consisting of multiple tones.
[0176] The learning unit also trains the machine learning model based on composite data obtained by combining sample data weighted for each type.
[0177] In this way, the information processing system can generate synthetic data from sample data containing multiple types of audio data, weighted by type, thereby enabling the training of a machine learning model that is personalized for each user or for each piece of audio data to be processed.
[0178] The learning unit also trains the machine learning model based on composite data obtained by combining sample data of different instruments for each type.
[0179] In this way, the information processing system generates synthetic data from sample data containing multiple types of audio data played using different instruments for each type, enabling the automatic generation of versatile and highly accurate transcription data, which is possible only with AI.
[0180] The learning unit also trains the machine learning model based on the composite data to which the editing process has been applied based on the composite data.
[0181] In this way, the information processing system can use synthetic data that is closer to more realistic audio data as training data for the machine learning model, making it possible to automatically generate more versatile and accurate transcription data when the audio data to be processed is real data.
[0182] The acquisition unit also acquires audio data, which is actual data, as the audio data to be processed.
[0183] In this way, the information processing system can automatically generate versatile and highly accurate transcription data that corresponds to the actual audio data.
[0184] (4. Hardware Configuration) Information devices such as the information processing device 100 and user terminal 10 according to each of the above-described embodiments are realized by a computer 1000 having a configuration such as that shown in FIG. 12 . The following description will be given using the information processing device 100 according to the embodiment as an example. FIG. 12 is a hardware configuration diagram showing an example of a computer 1000 that realizes the functions of the information processing device 100. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.
[0185] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.
[0186] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .
[0187] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records a conversion program according to the present disclosure, which is an example of program data 1450.
[0188] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.
[0189] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disk), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.
[0190] For example, when the computer 1000 functions as the information processing device 100 according to the embodiment, the CPU 1100 of the computer 1000 executes an information processing program loaded onto the RAM 1200 to realize functions of the control unit 130 and the like. The information processing program according to the present disclosure and data in the storage unit 120 are stored in the HDD 1400. The CPU 1100 reads and executes program data 1450 from the HDD 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.
[0191] The present technology may also be configured as follows. (1) An information processing system comprising: an acquisition unit that acquires audio data to be processed; and a generation unit that generates corresponding transcription data by inputting the audio data acquired by the acquisition unit into a trained machine learning model that has been trained using audio sample data. (2) A training unit that trains the machine learning model based on synthetic data obtained by synthesizing the sample data of a plurality of audio pieces. (3) The information processing system described in (2), in which the training unit trains the machine learning model based on synthetic data obtained by synthesizing the sample data of a plurality of corresponding audio pieces in accordance with MIDI data. (4) The information processing system described in (2), in which the training unit trains the machine learning model to reduce a discrepancy between the synthetic data and actual data. (5) The information processing system described in (4), in which the training unit trains the machine learning model including at least an encoder, a decoder, and a discriminator. (6) The information processing system according to (5), wherein the learning unit performs first learning to train the discriminator so that, when a pair of the synthetic data and the real data is input to the encoder, the discriminator can distinguish between the synthetic data and the real data. (7) The information processing system according to (6), wherein, after the first learning, the learning unit performs second learning to train the encoder in an adversarial manner to the first learning so that the discriminator cannot distinguish between the synthetic data and the real data. (8) The information processing system according to (6), wherein, after the first learning, the learning unit performs third learning to train the encoder in an adversarial manner to the first learning so that the discriminator mistakes the synthetic data for the real data. (9) The information processing system according to any one of (4) to (8), wherein the learning unit trains the machine learning model using a binary cross entropy loss (BCE loss) loss function.(10) The information processing system according to any one of (2) to (9), wherein the learning unit trains the machine learning model based on synthetic data obtained by synthesizing the sample data, each of which includes at least main audio and sub audio types. (11) The information processing system according to (10), wherein the learning unit trains the machine learning model based on synthetic data obtained by synthesizing the sample data weighted for each type. (12) The information processing system according to (10) or (11), wherein the learning unit trains the machine learning model based on synthetic data obtained by synthesizing the sample data for which different instruments are used for each type. (13) The information processing system according to any one of (2) to (12), wherein the learning unit trains the machine learning model based on synthesized data obtained by applying an editing process to the synthesized data. (14) The information processing system according to any one of (1) to (13), wherein the acquisition unit acquires actual audio data as the audio data to be processed. (15) An information processing method including: an information processing device acquiring audio data to be processed; and generating corresponding transcription data by inputting the acquired audio data into a trained machine learning model trained using audio sample data. (16) A program for causing a computer to function as: an acquisition unit that acquires the audio data to be processed; and a generation unit that generates corresponding transcription data by inputting the audio data acquired by the acquisition unit into a trained machine learning model trained using audio sample data.
[0192] REFERENCE SIGNS LIST 1 Information processing system 10 User terminal 11 Communication unit 12 Input unit 13 Output unit 14 Control unit 20 Encoder 30 Decoder 40 Discriminator 100 Information processing device 110 Communication unit 120 Storage unit 121 MIDI data storage unit 122 Sample data storage unit 123 Actual data storage unit 130 Control unit 131 Acquisition unit 132 Rendering unit 133 Editing unit 134 Conversion unit 135 Learning unit 136 Generation unit 141 Reception unit 142 Transmission unit N Network
Claims
1. An information processing system comprising: an acquisition unit that acquires audio data to be processed; and a generation unit that generates corresponding transcription data by inputting the audio data acquired by the acquisition unit into a trained machine learning model trained using audio sample data.
2. The information processing system according to claim 1, further comprising a learning unit that trains the machine learning model based on synthetic data obtained by synthesizing the sample data of a plurality of audios.
3. The information processing system according to claim 2, wherein the learning unit trains the machine learning model based on composite data obtained by synthesizing the sample data of multiple corresponding audios in accordance with MIDI data.
4. The information processing system according to claim 2, wherein the learning unit trains the machine learning model so as to reduce the deviation between the synthetic data and the actual data.
5. The information processing system according to claim 4, wherein the learning unit trains the machine learning model including at least an encoder, a decoder, and a discriminator.
6. The information processing system according to claim 5, wherein the learning unit performs a first learning process to train the discriminator so that the discriminator can distinguish between the synthetic data and the actual data when a pair of the synthetic data and the actual data is input to the encoder.
7. The information processing system according to claim 6, wherein the learning unit performs a second learning process after the first learning process, in an antagonistic manner to the first learning process, to train the encoder so that the discriminator is unable to distinguish between the synthetic data and the real data.
8. The information processing system according to claim 6, wherein the learning unit performs a third learning process after the first learning process, in an antagonistic manner to the first learning process, to train the encoder so that the discriminator mistakes the synthetic data for the real data.
9. The information processing system according to claim 4, wherein the learning unit trains the machine learning model using a loss function of BCE Loss (Binary Cross Entropy Loss).
10. The information processing system according to claim 2, wherein the learning unit trains the machine learning model based on composite data obtained by synthesizing the sample data, each sample data including at least the types of main audio and sub audio.
11. The information processing system according to claim 10, wherein the learning unit trains the machine learning model based on composite data obtained by combining the sample data weighted for each type.
12. The information processing system according to claim 10, wherein the learning unit trains the machine learning model based on composite data obtained by combining the sample data in which different instruments are used for each type.
13. The information processing system according to claim 2, wherein the learning unit trains the machine learning model based on the synthetic data after an editing process has been applied based on the synthetic data.
14. The information processing system according to claim 1, wherein the acquisition unit acquires audio data that is actual data as the audio data to be processed.
15. An information processing method comprising: an information processing device acquiring audio data to be processed; and generating corresponding transcription data by inputting the acquired audio data into a trained machine learning model trained using audio sample data.
16. An information processing program for causing a computer to function as: an acquisition unit that acquires audio data to be processed; and a generation unit that generates corresponding transcription data by inputting the audio data acquired by the acquisition unit into a trained machine learning model trained using audio sample data.