Information processing method

A generative model trained by machine learning generates musically consistent second parts for a musical piece, addressing the challenge of creating additional parts with specialized knowledge and effort, offering a user-friendly solution for musical composition.

JP2025079055APending Publication Date: 2025-05-21YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023191467
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-21

AI Technical Summary

Technical Problem

Creating additional parts of a musical piece other than the existing parts requires specialized knowledge and significant effort from the creator.

Method used

An information processing method that utilizes a generative model trained by machine learning to generate performance information for a second part of a musical piece corresponding to a different instrument from the first part, using a computer system to process control data and acquire first performance information.

Benefits of technology

Reduces the burden on users by automatically generating musically consistent second parts based on user-specified instrument preferences, allowing for efficient creation of new musical parts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025079055000001_ABST
    Figure 2025079055000001_ABST
Patent Text Reader

Abstract

To provide an information processing method which reduces a load for making a part other than an existing part of a music piece.SOLUTION: An information processing system 100 includes: an information acquisition section 22 for acquiring performance information X representing a first part corresponding to one or more music instruments in a music piece; and an information generation section 23 for processing control data C including the performance information X with a generative model G which has been trained by machine learning, so as to generate performance information Y representing a second part corresponding to a music instrument different from the one or more music instruments in the music piece.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to techniques for generating information representing music. [Background technology]

[0002] Technologies for performing various processes related to music composed of multiple parts have been proposed. For example, Patent Literature 1 discloses a technology for generating a trained model that selects a part to be played by a specific instrument from multiple parts included in music data by machine learning using features that affect the performance of the instrument. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2019-159146 A Summary of the Invention [Problem to be solved by the invention]

[0004] In a situation where a part of a piece of music corresponding to a specific instrument has been created, creating other parts of the piece of music requires specialized knowledge of music and a great deal of effort on the part of the creator. In consideration of the above circumstances, one aspect of the present disclosure aims to reduce the burden of creating parts other than the existing parts of a piece of music. [Means for solving the problem]

[0005] In order to solve the above problems, an information processing method according to one embodiment of the present disclosure obtains first performance information representing a first part in a musical piece corresponding to one or more instruments, and processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second part in the musical piece corresponding to an instrument different from the one or more instruments.

[0006] An information processing system according to one embodiment of the present disclosure includes an information acquisition unit that acquires first performance information representing a first part in a musical piece corresponding to one or more instruments, and an information generation unit that processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second part in the musical piece corresponding to an instrument different from the one or more instruments.

[0007] A program according to one embodiment of the present disclosure causes a computer system to function as an information acquisition unit that acquires first performance information representing a first part in a musical piece corresponding to one or more instruments, and an information generation unit that processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second part in the musical piece corresponding to an instrument different from the one or more instruments. [Brief description of the drawings]

[0008] [Figure 1] 1 is a block diagram illustrating a configuration of an information processing system according to a first embodiment. [Diagram 2] FIG. 2 is an explanatory diagram regarding the arrangement of a target piece of music by the information processing system. [Diagram 3] FIG. 2 is a block diagram illustrating an example of a functional configuration of an information processing system. [Figure 4] FIG. [Diagram 5] FIG. 2 is an explanatory diagram of tokens constituting performance information. [Figure 6] FIG. 4 is a schematic diagram of control data. [Figure 7] 4 is a schematic diagram of performance information generated from control data. FIG. [Figure 8] 13 is a flowchart of an arrangement process. [Figure 9] FIG. 11 is a block diagram illustrating a configuration of a machine learning system according to a second embodiment. [Figure 10] FIG. 1 is a block diagram illustrating an example of the functional configuration of a machine learning system. [Figure 11]13 is a flowchart of a training data generation process. [Figure 12] 13 is a flowchart of a training process. [Figure 13] FIG. 13 is an explanatory diagram regarding the arrangement of a target piece of music in the third embodiment. [Figure 14] FIG. 13 is a block diagram illustrating an example of a functional configuration of an information processing system according to a third embodiment. [Figure 15] FIG. 13 is a block diagram illustrating an example of a functional configuration of an information processing system according to a fourth embodiment. [Figure 16] 13 is an explanatory diagram regarding an arrangement of a target piece of music in a modified example. FIG. [Figure 17] FIG. 11 is a block diagram illustrating a functional configuration of an information processing system according to a modified example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] A: First embodiment 1 is a block diagram illustrating an example of the configuration of an information processing system 100 in the first embodiment. The information processing system 100 is a computer system that arranges existing music (hereinafter referred to as "target music"), and is realized by an information device such as a smartphone, a tablet terminal, or a personal computer.

[0010] FIG. 2 is an explanatory diagram of an overview of processing by the information processing system 100. The information processing system 100 generates music data Dy from music data Dx. The music data Dx is time-series data representing a plurality of first parts constituting a target music piece. Specifically, the music data Dx represents a time series of notes constituting each of the plurality of first parts. Each of the plurality of first parts is an existing performance part corresponding to a different type of musical instrument. Note that a "performance part" is a part (vocal part) constituting a music piece, and means, for example, a set of one or more performers (or a set of one or more musical instruments) who play a common note. For example, the performance parts are distinguished by musical instrument, by range, or by musical role (e.g., main melody / sub-melody).

[0011] The musical piece data Dy is time-series data representing a target musical piece in which a new second part has been added to a plurality of first parts. The second part is a performance part corresponding to a different type of instrument from the plurality of first parts, and constitutes the target musical piece by being performed in parallel with the plurality of first parts. Therefore, the second part is in musical harmony with the plurality of first parts. As described above, the information processing system 100 arranges an existing target musical piece composed of a plurality of first parts into a target musical piece to which a new second part has been added.

[0012] The music data Dy represents a time series of notes that make up each of the first and second parts of a target music piece. Specifically, the music data Dy includes existing music data Dx that represents a plurality of first parts, and new music data Da that represents the second part. That is, the information processing system 100 generates music data Da of the second part from the existing music data Dx. The music data Da represents a time series of notes that make up the second part.

[0013] The music data Dx and the music data Dy specify the pitch and the sounding period for each note constituting each of the multiple performance parts. The pitch is one of multiple discretely set scale notes (e.g., note number). The sounding period is specified, for example, by the start point and duration (or end point) of the note. The music data Dx and the music data Dy are, for example, music files that comply with the MIDI (Musical Instrument Digital Interface) standard. Note that the music data Dx is an example of "first music data," and the music data Da or the music data Dy is an example of "second music data."

[0014] 1, the information processing system 100 includes a control device 11, a storage device 12, a communication device 13, an operation device 14, a sound source device 15, and a sound emission device 16. The information processing system 100 may be realized as a single device, or may be realized as a plurality of devices configured separately from each other.

[0015] The control device 11 is composed of one or more processors that control each element of the information processing system 100. For example, the control device 11 is composed of one or more types of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a sound processing unit (SPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC). The communication device 13 communicates with an external device via a communication network such as the Internet.

[0016] The storage device 12 is one or more memories that store the programs executed by the control device 11 and various data used by the control device 11. For example, the storage device 12 stores music data Dx of a target music piece. The storage device 12 is configured with a known recording medium such as a magnetic recording medium or a semiconductor recording medium. The storage device 12 may be configured with a combination of multiple types of recording media. In addition, a portable recording medium that is detachable from the information processing system 100, or a recording medium (e.g., cloud storage) to which the control device 11 can write or read via a communication network may be used as the storage device 12.

[0017] The operation device 14 is an input device that accepts instructions from a user. The operation device 14 is, for example, an operator operated by the user, or a touch panel that detects contact by the user. The user can specify the type of instrument corresponding to the second part by operating the operation device 14. For example, the user operates the operation device 14 to select a desired type of instrument from a plurality of candidates prepared in advance. Note that the operation device 14, which is separate from the information processing system 100, may be connected to the information processing system 100 by wire or wirelessly.

[0018] The sound source device 15 generates an audio signal representing a waveform of a musical tone designated by the music data Dx or Dy. The function of the sound source device 15 may be realized by the control device 11 executing a program.

[0019] The sound emitting device 16 reproduces sound waves under the control of the control device 11. The sound emitting device 16 is an output device such as a speaker or a headphone. Specifically, the sound emitting device 16 reproduces musical tones of a target piece of music represented by an acoustic signal generated by the sound source device 15. Note that the sound source device 15 or the sound emitting device 16, which are separate from the information processing system 100, may be connected to the information processing system 100 by wire or wirelessly.

[0020] 3 is a block diagram illustrating an example of a functional configuration of the information processing system 100. The control device 11 executes a program stored in the storage device 12 to realize a plurality of functions (a pre-processing unit 21, an information acquisition unit 22, an information generation unit 23, and a post-processing unit 24) for generating music data Dy of an arranged target music piece from music data Dx of an existing target music piece.

[0021] The pre-processing unit 21 generates performance information X from the music data Dx. The performance information X is time-series data representing multiple first parts of a target music piece. The performance information X is also expressed as intermediate or alternative data having a different format from the music data Dx. The performance information X is an example of "first performance information."

[0022] FIG. 4 is a schematic diagram of performance information X. The performance information X is a time series of multiple tokens T that represent each first part of a target piece of music. A token T is a unit that constitutes the performance information X. The performance information X is described, for example, by a series of character strings. In FIG. 4, the performance information X is illustrated as a character string spanning multiple lines. Each token T is separated by a blank (for example, a half-width space).

[0023] 5 is an explanatory diagram of the tokens T. The multiple tokens T constituting the performance information X include a note token Ta, a beat token Tb, and an auxiliary token Tc.

[0024] The note token Ta is a token T that represents a note of the target music piece. Specifically, the note token Ta is divided into a performance token Ta1 and a note value token Ta2. The performance token Ta1 specifies the type of instrument used to play the note and the pitch of the note (e.g., note number). For example, a performance token Ta1 written as "acg_47" means a note with a pitch of 47 to be played by an acoustic guitar (acg: acoustic guitar). Also, a performance token Ta1 written as "acp_65" means a note with a pitch of 65 to be played by an acoustic piano (acp: acoustic piano). The performance token Ta1 is also expressed as identification information that specifies the type of instrument.

[0025] The note value token Ta2 specifies the note value (duration) of the note. For example, the note value token Ta2 written as "len_L" (L is a natural number) means a time length equivalent to L unit times. The unit time is a time length according to the beat interval of the target music piece. Specifically, the unit time is a time length equivalent to, for example, 1 / 12 beat of the target music piece. The combination of the performance token Ta1 and note value token Ta2 expresses one note with a specified instrument, pitch, and note value.

[0026] The beat token Tb is a token T that represents the beat of the target music piece. Specifically, a specific point in time in the target music piece is represented by the beat token Tb. Specifically, the beat token Tb is divided into a bar token Tb1 (bar), a beat token Tb2 (beat), and a position token Tb3 (pos). The beat token Tb that represents a specific point in time in the target music piece is placed at a position corresponding to that point in the time series of tokens T in the performance information X.

[0027] The bar token Tb1 is a token T that signifies a bar line of the target music (i.e., the first beat of each bar). The beat token Tb2 is a token T that signifies each beat of the target music. The position token Tb3 is a token T that expresses a specific point in time in the target music. For example, a position token Tb3 written as "pos_K" (K=1 to 11) signifies a point in time when a time equivalent to K unit times has elapsed from the bar line (bar token Tb1) or beat (beat token Tb2) immediately preceding the position token Tb3.

[0028] The position token Tb3 is used to express the position of a note, for example. Specifically, the position of one note is expressed by the position token Tb3 arranged immediately before the note token Ta (performance token Ta1 and note value token Ta2) that represents the note. For example, the notation "pos_9 acp_65 len_27" in the performance information X means that a note with a pitch of 65 and a note value of 27 played by an acoustic piano (acp) starts at a point (pos_9) when nine unit times have elapsed from the previous beat. Note that if the start point of the note coincides with a bar line (bar token Tb1) or a beat (beat token Tb2), the position token Tb3 is omitted. As shown in the above example, the sounding conditions (pitch, note value, and position) of one note are specified by the performance token Ta1, note value token Ta2, and position token Tb3.

[0029] The auxiliary token Tc is a token T that represents various information related to the performance of the target song. Specifically, the auxiliary token Tc is divided into a section token Tc1 (section), a key token Tc2 (key), and a tempo token Tc3 (tempo). The section token Tc1 is a token T that indicates the start point of each structural section of the target song. The structural section is a section obtained by dividing the target song on the time axis according to its musical meaning. For example, each section such as an intro, an A melody, a B melody (bridge), a chorus, and an outro is exemplified as a structural section. The section token Tc1 is placed at a position in the performance information X that corresponds to the start point of a structural section in the target song.

[0030] The key token Tc2 is a token T that represents the key of the target song. The numerical value specified by the key token Tc2 is a number for identifying the key. The key token Tc2 is placed at a position in the performance information X that corresponds to a time point at which the key changes in the target song. The speed token Tc3 is a token T that represents the tempo of the target song. The numerical value specified by the speed token Tc3 indicates the tempo (BPM: Beats Per Minute). The speed token Tc3 is placed at a position in the performance information X that corresponds to a time point at which the tempo changes in the target song.

[0031] The above are specific examples of the tokens T that make up the performance information X. The pre-processing unit 21 in Fig. 3 converts the music data Dx into the performance information X under the rules exemplified above. According to the first embodiment, the widely used existing music data Dx can be used to generate music data Dy (performance information Y described later).

[0032] The information acquiring unit 22 acquires performance information X. The information acquiring unit 22 in the first embodiment generates control data C including the performance information X and instrument information Q. The instrument information Q is information specifying the type of instrument corresponding to the second part. Specifically, the type of instrument specified by the user by operating the operation device 14 is specified by the instrument information Q.

[0033] Fig. 6 is a schematic diagram of control data C. The information acquisition unit 22 generates the control data C by adding instrument information Q to the beginning of performance information X. The instrument information Q is added to the performance information X as, for example, one token T. In the following description, it is assumed that a drum set (drums) consisting of multiple types of percussion instruments is specified as the instrument of the second part, as shown in the example of Fig. 6. The instrument information Q may be added to any position in the performance information X.

[0034] The information generating unit 23 in FIG. 3 generates performance information Y from the control data C. The performance information Y is time-series data representing the second part of the target song after arrangement. FIG. 7 is a schematic diagram of the performance information Y. The performance information Y is a time-series of multiple tokens T representing the second part of the target song. Like the performance information X, the performance information Y is described, for example, as a series of character strings. Note that the performance information Y is an example of "second performance information."

[0035] The types and meanings of each token T constituting the performance information Y are the same as those of each token T in the performance information X described above with reference to Fig. 5. That is, the multiple tokens T constituting the performance information Y include note tokens Ta (Ta1, Ta2), beat tokens Tb (Tb1, Tb2, Tb3), and auxiliary tokens Tc (Tc1, Tc2, Tc3), as in the example of Fig. 5.

[0036] In the first embodiment, as described above, it is assumed that a drum set consisting of multiple types of percussion instruments is specified as the second part. Pitch and value are not considered for the performance sounds of the drum set. Therefore, in the performance information Y shown in FIG. 7, the performance token Ta1 does not include a pitch specification, and the value token Ta2 is not used.

[0037] In addition, the performance token Ta1 of the drum set specifies the type of each percussion instrument that constitutes the drum set. For example, the performance token Ta1 written as "bass" means the performance of a bass drum, the performance token Ta1 written as "hhcls" means the performance of a hi-hat, and the performance token Ta1 written as "snare" means the performance of a snare drum. If an instrument other than the drum set is specified for the second part, one note is expressed by a combination of the performance token Ta1 and the note value token Ta2, as shown in the example of FIG. 5.

[0038] A trained generative model G is used for generating the performance information Y by the information generating unit 23. The generative model G is a statistical model that has learned the relationship between the control data C and the performance information Y by prior machine learning. The information generating unit 23 processes the control data C by the trained generative model G to generate the performance information Y.

[0039] The generative model G is realized by a combination of a program that causes the control device 11 to execute a calculation to generate performance information Y from control data C, and a number of variables (e.g., bias and weighting values) that are applied to the calculation. The numerical values ​​of each of the multiple variables are set in advance by machine learning.

[0040] For example, a transformer, which is an encoder-decoder model including a self-attention mechanism (specifically, a multi-head attention mechanism), is used as the generative model G. The transformer is disclosed, for example, in Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin, "Attention Is All You Need", 31st Conference on Neural Information Processing Systems (NIPS 2017). In addition, MEGA (Moving Average Equipped Gated Attention) may be adopted as the generative model G. MEGA is disclosed in Ma, C. Zhou, X. Kong, J. He, L. Gui, G. Neubig, J. May, and L. Zettlemoyer. "Mega: moving average equipped gated attention", arXiv:2209.10655, 2022.

[0041] The format of the performance information X and the performance information Y, which are configured as a time series of multiple tokens T, is suitable for processing by the generation model G. As described above, according to the first embodiment, by utilizing the generation model G suitable for processing a time series of multiple tokens T, an effective second part that is musically consistent with the first part can be generated.

[0042] The post-processing unit 24 in Fig. 3 generates the music data Dy from the performance information Y. Specifically, the post-processing unit 24 generates the second part of the music data Da from the performance information Y under the rules described above with reference to Fig. 5, and generates the music data Dy by integrating the pre-arrangement music data Dx and the second part of the music data Da. Since the music data Dy is generated from the performance information Y in the above manner, the musical tones of the second part can be generated by an existing sound source device 15 capable of processing the music data Dy.

[0043] 8 is a flowchart of the process (hereinafter referred to as "arrangement process") in which the control device 11 generates music data Dy from music data Dx. For example, the arrangement process is started in response to an operation of the operation device 14 by the user.

[0044] When the arrangement process is started, the control device 11 (pre-processing unit 21) generates performance information X from the music piece data Dx (Sa1). The control device 11 (information acquisition unit 22) generates instrument information Q that specifies the type of instrument for the second part (Sa2). The control device 11 (information acquisition unit 22) acquires the performance information X and generates control data C including the instrument information Q and the performance information X (Sa3).

[0045] The control device 11 (information generation unit 23) processes the control data C using the generation model G to generate performance information Y (Sa4). The control device 11 (post-processing unit 24) generates song data Da of the second part from the performance information Y (Sa5). The control device 11 (post-processing unit 24) combines the song data Dx and the song data Da to generate song data Dy of the arranged target song (Sa6).

[0046] As described above, in the first embodiment, the performance information Y of the second part corresponding to an instrument different from the instrument of the first part is generated from the performance information X representing the existing first part of the target song. Therefore, the burden on the user who creates a second part other than the existing first part for the target song can be reduced. In other words, according to the first embodiment, it is possible to provide the user with a unique customer experience of being able to obtain a new second part corresponding to the existing first part.

[0047] In particular, in the first embodiment, performance information Y can be generated for the second part corresponding to the type of instrument specified by the instrument information Q. Since the instrument information Q specifies the instrument selected by the user, the second part can be generated according to the user's intention or preference. In addition, by changing the type of instrument specified by the instrument information Q, second parts for various instruments can be generated using a single generative model G.

[0048] B: Second embodiment A second embodiment will be described. Note that, for elements having the same functions as those in the first embodiment in each of the following exemplary aspects, the same reference numerals as those in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.

[0049] 9 is a block diagram illustrating the configuration of a machine learning system 200 in the second embodiment. The machine learning system 200 is a computer system that establishes the above-mentioned generative model G used by the information processing system 100 through machine learning. The machine learning system 200 includes a control device 31, a storage device 32, and a communication device 33. Note that the machine learning system 200 may be realized as a single device, or may be realized as a plurality of devices configured separately from each other.

[0050] The control device 31 is composed of one or more processors that control each element of the machine learning system 200. For example, the control device 31 is composed of one or more types of processors such as a CPU, a GPU, an SPU, a DSP, an FPGA, or an ASIC.

[0051] The communication device 33 communicates with an external device via a communication network such as the Internet. For example, the communication device 33 communicates with the information processing system 100. The generative model G established by the machine learning system 200 is provided to the information processing system 100 by the communication device 33.

[0052] The storage device 32 is one or more memories that store the programs executed by the control device 31 and various data used by the control device 31. The storage device 32 is configured with a known recording medium such as a magnetic recording medium or a semiconductor recording medium. The storage device 32 may be configured with a combination of multiple types of recording media. In addition, a portable recording medium that is detachable from the information processing system 100, or a recording medium (e.g., cloud storage) to which the control device 31 can write or read via a communication network may be used as the storage device 32.

[0053] The storage device 32 stores a plurality of pieces of music data R used for machine learning of the generative model G. Each of the plurality of pieces of music data R is time-series data representing a piece of music for machine learning (hereinafter referred to as a "reference piece of music"). Each reference piece of music is composed of H performance parts (H is a natural number of 2 or more). The music data R represents a time series of notes for each of the H performance parts. The music data R also specifies the type of instrument for each of the H performance parts. The music data R is, for example, a music file that complies with the MIDI (Musical Instrument Digital Interface) standard. For example, a large number of karaoke data created in the past is used as the music data R. Therefore, a large number of pieces of music data R can be easily prepared. The number H of performance parts of the reference piece of music differs for each reference piece of music.

[0054] 10 is a block diagram illustrating an example of a functional configuration of the machine learning system 200. The control device 31 executes a program stored in the storage device 32 to realize a plurality of functions (a training data generation unit 41, a training processing unit 42) for establishing the generative model G.

[0055] The training data generation unit 41 generates multiple pieces of training data Z to be used for machine learning of the generative model G. Each of the multiple pieces of training data Z is composed of a combination of training control data Ct and training performance information Yt. The control data Ct is data in the same format as the above-mentioned control data C. The performance information Yt is data in the same format as the above-mentioned performance information Y. The training data generation unit 41 generates multiple pieces of training data Z from multiple pieces of music data R stored in the storage device 32. The training processing unit 42 establishes the generative model G by machine learning using the multiple pieces of training data Z.

[0056] 11 is a flowchart of a process (hereinafter referred to as "training data generation process") in which the control device 31 (training data generation unit 41) generates multiple pieces of training data Z. For example, the training data generation process is started in response to an instruction from an administrator of the machine learning system 200. When the training data generation process is started, the control device 31 selects one of multiple pieces of music data R stored in the storage device 32 (hereinafter referred to as "selected music data R") (Sb1).

[0057] The control device 31 selects one of the H performance parts of the reference song represented by the selected song data R as the second part (Sb2), and generates song data Da for the second part from the selected song data R (Sb3). For example, a part of the selected song data R that corresponds to the second part is extracted as the song data Da. The control device 31 generates performance information Yt from the song data Da for the second part (Sb4). The process of generating the performance information Yt from the song data Da is similar to the process of the pre-processing unit 21 generating performance information X from the song data Dx. The control device 31 generates instrument information Q indicating the type of instrument specified by the selected song data R for the second part (Sb5).

[0058] The control device 31 selects a plurality of first parts from (H-1) performance parts excluding the second parts among the H performance parts of the reference music piece represented by the selected music piece data R (Sb6). For example, (H-1) or less first parts are randomly selected from the (H-1) performance parts. All combinations of a predetermined number of first parts selected from the (H-1) performance parts may be used in order. The control device 31 generates music piece data Dx representing a plurality of first parts from the selected music piece data R (Sb7). For example, parts of the selected music piece data R corresponding to the plurality of first parts are extracted as the music piece data Dx. The control device 31 generates performance information X from the music piece data Dx (Sb8). The generation of the performance information X is similar to the process in which the pre-processing unit 21 generates the performance information X from the music piece data Dx.

[0059] The control device 31 generates training control data Ct including performance information X corresponding to a plurality of first parts and instrument information Q specifying an instrument for the second part (Sb9). The control device 31 then generates training data Z by associating the control data Ct with the performance information Yt, and stores the training data Z in the storage device 32 (Sb10).

[0060] The control device 31 judges whether a predetermined end condition is satisfied (Sb11). The end condition is, for example, that the number of the performance parts selected as the second parts among the H performance parts of the reference music piece reaches a predetermined value.

[0061] If the end condition is not met (Sb11: NO), the control device 31 moves the process to step Sb2. That is, the control device 31 selects an unselected performance part from among the H performance parts of the reference music piece represented by the selected music piece data R as the second part (Sb2). As described above, the generation of performance information Yt (Sb2-Sb4), musical instrument information Q (Sb5), control data Ct (Sb6-Sb9), and training data Z (Sb10) are repeated until the end condition is met (Sb11: YES). That is, from the music piece data R of one reference music piece, multiple training data Z are generated with different types of musical instruments for the second part.

[0062] If the end condition is met (Sb11: YES), the control device 31 judges whether or not the above processes (Sb2 to Sb10) have been performed for all music data R stored in the storage device 32 (Sb12). If there is unprocessed music data R (Sb12: NO), the control device 31 transitions the process to step Sb1. That is, the control device 31 selects the unprocessed music data R as new selected music data R (Sb1). As described above, the generation of multiple training data Z (Sb2 to Sb10) is repeated for each of the multiple music data R. On the other hand, if all music data R has been processed (Sb12: YES), the control device 31 ends the training data generation process.

[0063] However, in a form in which the performance information X and the performance information Yt are prepared separately in the process of generating the training data Z, it is necessary to recognize and manage the temporal correspondence between the performance information X and the performance information Y. In the first embodiment, since the performance information X and the performance information Yt are generated from the existing music piece data R, it is not necessary to recognize and manage the temporal correspondence between the performance information X and the performance information Yt. Therefore, the load required for generating the training data Z can be reduced.

[0064] 12 is a flowchart of a process (hereinafter referred to as the "training process") in which the control device 31 (training processing unit 42) establishes a generative model G through machine learning using a plurality of training data Z. After the training data generation process is executed, the training process is started, for example, in response to an instruction from an administrator of the machine learning system 200. The training process is an example of a method for generating a generative model G.

[0065] When the training process is started, the control device 31 selects one of the multiple training data Z (hereinafter referred to as "selected training data Z") stored in the storage device 32 (Sc1). As illustrated in FIG. 10, the control device 31 generates performance information Y by processing control data Ct of the selected training data Z using an initial or provisional generation model G (hereinafter referred to as "provisional model G0") (Sc2). The control device 31 calculates a loss function that represents the error between the performance information Y generated by the provisional model G0 and the performance information Yt of the selected training data Z (Sc3). The control device 31 updates multiple variables of the provisional model G0 so that the loss function is reduced (ideally minimized) (Sc4).

[0066] The control device 31 determines whether a predetermined termination condition is satisfied (Sc5). The termination condition is, for example, that the loss function falls below a predetermined threshold, or that the amount of change in the loss function falls below a predetermined threshold. If the termination condition is not satisfied (Sc5: NO), the control device 31 selects the unselected training data Z stored in the storage device 32 as new selected training data Z (Sc1). That is, the process of updating multiple variables of the provisional model G0 (Sc2 to Sc4) is repeated until the termination condition is satisfied (Sc5: YES).

[0067] When the termination condition is met (Sc5: YES), the control device 31 ends the training process. The provisional model G0 at the time when the termination condition is met is determined as the trained generative model G.

[0068] As can be understood from the above explanation, the generative model G learns the latent relationship between the control data Ct and the performance information Yt in multiple training data Z. Therefore, the trained generative model G outputs the performance information Y that is statistically appropriate for the unknown control data C based on the above relationship.

[0069] The trained generative model G is transmitted from the communication device 33 to the information processing system 100. The control device 11 of the information processing system 100 receives the generative model G via the communication device 13, and stores the generative model G in the storage device 12. The generative model G provided by the above procedure is used in the musical arrangement process illustrated in FIG.

[0070] C: Third embodiment FIG. 13 is an explanatory diagram of the process by the information processing system 100 of the third embodiment. As in the first embodiment, the control device 11 generates the performance information Y of the second part (hereinafter referred to as "performance information Y1") from the control data C including the performance information X and the instrument information Q. In the third embodiment, the performance information Y2 of the third part is generated from new performance information X including the performance information X of the multiple first parts and the performance information Y1 of the second part. The third part is a performance part corresponding to a different type of instrument from the multiple first parts and the second parts, and constitutes the target musical piece after arrangement by being played in parallel with the multiple first parts and the second parts. Therefore, the third part is musically in harmony with each of the first parts and the second parts. As described above, in the third embodiment, the process of generating the performance information Y from the performance information X is cumulatively repeated, so that the performance parts constituting the target musical piece are sequentially increased.

[0071] 14 is a block diagram illustrating a functional configuration of the information processing system 100 in the third embodiment. The configuration and operation of the preprocessing unit 21 in the third embodiment are similar to those in the first embodiment. In the third embodiment, the user operates the operation device 14 to specify the types of musical instruments corresponding to the second part and the third part, respectively. That is, two different types of musical instruments are specified by the user.

[0072] The information acquisition unit 22 first generates control data C including performance information X representing a plurality of first parts and instrument information Q1 representing the type of instrument corresponding to the second part, as in the first embodiment. The information generation unit 23 processes the control data C using a generation model G to generate performance information Y1 for the second part.

[0073] After generating the performance information Y1, the information acquisition unit 22 generates new performance information X including the existing performance information X and the performance information Y1. That is, the performance information X of the target music piece composed of a plurality of first parts and second parts is generated. The information acquisition unit 22 generates control data C including the performance information X to which the second part is added and instrument information Q2 representing the type of instrument corresponding to the third part. The information generation unit 23 processes the control data C using the generation model G to generate performance information Y2 of the third part. As described above, the information generation unit 23 of the third embodiment generates performance information Y2 representing the third part by processing the control data C including the performance information X and the performance information Y1 using the generation model G. The performance information Y2 is an example of the "third performance information".

[0074] The post-processing unit 24 generates the second part of the song data Da1 from the performance information Y1, and generates the third part of the song data Da2 from the performance information Y2. The post-processing unit 24 generates the song data Dy by integrating the pre-arrangement song data Dx, the second part of the song data Da1, and the third part of the song data Da2.

[0075] The third embodiment also achieves the same effect as the first embodiment. In the third embodiment, the control data C including the existing performance information X and the generated performance information Y1 is processed by the generation model G to generate the performance information Y2 representing the third part. Therefore, the performance information Y2 of the third part that is musically consistent with both the first and second parts can be generated. That is, according to the third embodiment, the target music piece that is composed of multiple performance parts can be arranged. According to the third embodiment, the user can be provided with a unique customer experience in which multiple new performance parts (second and third parts) corresponding to the existing first part can be obtained.

[0076] The number of times that the information generating unit 23 repeats the generation of the performance information Y is arbitrary. In the (n+1)th processing by the information generating unit 23 (n is a natural number), performance information Yn+1 is generated from control data C that includes, as new performance information X, the performance information Yn generated in the immediately preceding nth processing and the performance information X used in the nth processing.

[0077] D: Fourth embodiment 15 is a block diagram illustrating a functional configuration of an information processing system 100 in the fourth embodiment. The control device 11 in the fourth embodiment functions as a musical score generator 25 in addition to the same elements as those in the first embodiment (a pre-processing unit 21, an information acquisition unit 22, an information generating unit 23, and a post-processing unit 24).

[0078] The score generation unit 25 generates score data E from the music piece data Dy. The score data E is data representing the score of the target music piece after editing, including the first part and the second part. For example, data in MusicXML format or PDF format is exemplified as the score data E. Any known technology may be used to generate the score data E. The score generation unit 25 displays the score represented by the score data E on a display device (not shown).

[0079] The fourth embodiment also achieves the same effects as the first embodiment. Moreover, in the fourth embodiment, the score data E of the edited target musical piece including the first and second parts is generated, so that the user can check the contents of the score of the arranged target musical piece.

[0080] The configuration of the third embodiment is similarly applied to the fourth embodiment. That is, the target music piece represented by the musical score data E may include a third part (and other parts) in addition to the first and second parts.

[0081] E: Variation Specific modified embodiments added to each of the above-mentioned embodiments are exemplified below. Two or more embodiments selected from the following examples may be appropriately combined as long as they are not mutually contradictory.

[0082] (1) In the above-mentioned embodiments, the control data C includes the performance information X and the instrument information Q, but the control data C may include other information. For example, the control data C may include condition information specifying musical conditions related to the second part. The condition information may specify conditions such as whether the constituent notes of the second part are single notes or chords, or whether the entire section or a part of the target musical piece is to be generated for the second part. The condition information may also include the difficulty of performance related to the second part. According to the above-mentioned embodiments, it is possible to generate a variety of target musical pieces that satisfy the conditions specified by the condition information.

[0083] (2) In each of the above-mentioned embodiments, the section of the target song that serves as the unit of processing by the information processing system 100 may be any section. For example, the above-mentioned arrangement process may be performed sequentially or in parallel for each of a plurality of sections (hereinafter referred to as "unit sections") obtained by dividing the target song on the time axis. The song data Dx and the song data Dy represent a string of notes corresponding to one unit section of the target song. Each unit section is, for example, a section of a time length equivalent to a predetermined number of bars in the target song (for example, 4 to 8 bars). The second part song data Da that covers all sections of the target song may be generated by a single arrangement process targeting all sections of the target song.

[0084] (3) In each of the above-mentioned embodiments, the song data Dy is generated by integrating the second part song data Da with the first part song data Dx, but the generation of the song data Dy may be omitted. For example, the second part song data Da may be generated as the final result (song data Dy). That is, the song data Dy is expressed as data representing the time series of notes that make up the second part. The song data Dy may include a part other than the second part (for example, the first part or the third part), or may be composed of only the second part.

[0085] In addition, the musical score represented by the musical score data E in the fourth embodiment may be a musical score including a first part and a second part (and other parts such as a third part), a musical score consisting of only the second part, and a musical score consisting of the second part and a third part (i.e., a new part).

[0086] (4) In the above-mentioned embodiments, the music data Da corresponding to the performance information Y is integrated with the existing music data Dx, but the method of generating the arranged music data Dy is not limited to the above. For example, the post-processing unit 24 may integrate the performance information X and the performance information Y and generate the music data Dy from the integrated performance information.

[0087] (5) In each of the above-mentioned embodiments, the performance information X represents multiple first parts, but the number of first parts represented by the performance information X (music data Dx) may be 1. The performance information X is comprehensively expressed as data representing first parts corresponding to one or more musical instruments.

[0088] In addition, in each of the above-mentioned embodiments, the performance information Y represents one second part, but the number of second parts represented by the performance information Y (music data Da) may be two or more. The performance information Y is comprehensively expressed as data representing second parts corresponding to one or more musical instruments.

[0089] (6) In each of the above-described embodiments, the control data C including the instrument information Q is processed by a single generative model G to generate the performance information Y of an instrument corresponding to the instrument information Q. However, multiple generative models G corresponding to different instruments may be selectively used to generate the performance information Y.

[0090] For example, by applying a plurality of training data Z, in which a performance part of a specific instrument is selected as the second part from among a large number of training data Z generated by the training data generation process, to the training process, a generation model G corresponding to the instrument is established. The generation model G of each instrument generates performance information Y of the performance part corresponding to the instrument.

[0091] The information generating unit 23 processes the control data C using a generation model G corresponding to an instrument selected by a user from among a plurality of generation models G corresponding to different instruments, thereby generating performance information Y of the instrument. In the above embodiment, the configuration for inputting the instrument information Q to the generation model G is omitted. In other words, the control data C does not need to include the instrument information Q.

[0092] According to each of the above-mentioned modes in which the musical instrument information Q is input to the generative model G, the performance information Y of various musical instruments can be generated using a single generative model G. In other words, there is an advantage that there is no need to prepare an individual generative model G for each musical instrument.

[0093] (7) In the third embodiment, the second and third parts of the target song are generated by cumulatively repeating the process of generating performance information Y from performance information X, but the method of generating a target song including the second and third parts is not limited to the above example. For example, as illustrated in Fig. 16, the control device 11 may sequentially or in parallel execute a process of generating performance information Y1 of the second part from performance information X and a process of generating performance information Y2 of the third part from performance information X, and generate the target song by integrating the performance information X (song data Dx), the performance information Y1 (song data Da1), and the performance information Y2 (song data Da2).

[0094] (8) In each of the above-described embodiments, a transformer is exemplified as the generative model G, but the configuration or type of the generative model G is arbitrary. For example, a deep neural network such as a recurrent neural network (RNN) or a long short-term memory (LSTM) may be used as the generative model G. The generative model G may be configured by combining multiple types of statistical models.

[0095] (9) In each of the above-described forms, the information processing system 100 and the machine learning system 200 are illustrated as separate systems, but the functions of the machine learning system 200 (the training data generation unit 41 and the training processing unit 42) may be incorporated in the information processing system 100.

[0096] (10) In each of the above-mentioned forms, the performance parts of the target song correspond to musical instruments, but some or all of the performance parts constituting the target song may be singing parts corresponding to singing voices. That is, the "musical instrument" in each of the above-mentioned forms may be replaced with a "singer". Also, the "musical instrument" in each of the above-mentioned forms may be replaced with a "sound source" including both an instrument and a singer. That is, each performance part of the target song corresponds to a different type of sound source.

[0097] (11) The information processing system 100 may be realized by a server device that communicates with a terminal device such as a smartphone, a tablet terminal, or a personal computer. For example, the control device 11 receives music data Dx and instrument information Q transmitted from the terminal device via the communication device 13, and generates music data Dy using the music data Dx and the instrument information Q. The control device 11 transmits the music data Dy to the terminal device via the communication device 13.

[0098] In addition, in a configuration in which the pre-processing unit 21 that generates performance information X from music piece data Dx is installed in the terminal device, the pre-processing unit 21 is omitted from the information processing system 100 as illustrated in Fig. 17. That is, the information acquiring unit 22 receives the performance information X and the musical instrument information Q transmitted from the terminal device via the communication device 13, and generates control data C including the performance information X and the musical instrument information Q.

[0099] In addition, in a configuration in which the post-processing unit 24 that generates the music piece data Da from the performance information Y is installed in the terminal device, the post-processing unit 24 is omitted from the information processing system 100, as illustrated in Fig. 17. That is, the information generating unit 23 transmits the performance information Y generated by the generation model G from the communication device 13 to the terminal device. The terminal device generates the music piece data Da (and further the music piece data Dy) from the performance information Y received from the information processing system 100.

[0100] (12) As described above, the functions of the information processing system 100 exemplified above are realized by the cooperation of one or more processors constituting the control device 11 and the program stored in the storage device 12. The program according to the present disclosure can be provided in a form stored in a computer-readable recording medium and installed in a computer. The recording medium is, for example, a non-transitory recording medium, and a good example is an optical recording medium (optical disk) such as a CD-ROM, but also includes any known type of recording medium such as a semiconductor recording medium or a magnetic recording medium. Note that the non-transitory recording medium includes any recording medium except for a transient propagating signal, and does not exclude volatile recording media. In addition, in a configuration in which a distribution device distributes a program via a communication network, the storage medium that stores the program in the distribution device corresponds to the non-transitory recording medium described above.

[0101] F: Notes From the above-described exemplary embodiments, the following configurations can be understood, for example.

[0102] An information processing method according to one aspect (aspect 1) of the present disclosure acquires first performance information representing a first part corresponding to one or more instruments in a musical piece, and processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second part corresponding to an instrument different from the one or more instruments in the musical piece. According to the above aspect, second performance information of a second part corresponding to an instrument different from the instrument of the first part is generated from first performance information representing an existing first part of a musical piece. Therefore, it is possible to reduce the burden on a user who creates a second part other than the existing first part for a musical piece.

[0103] "(First / second) performance information" is data in any format that represents parts that make up a piece of music. For example, performance information is data that represents a time series of notes that correspond to a specific part. For example, performance information is composed of a time series of multiple tokens that represent a specific part of a piece of music. Each token specifies, for example, the sounding conditions (e.g., pitch, position, and duration) of each note that makes up the piece of music.

[0104] A "generative model" is a statistical model of any configuration that has learned the relationship between the control data and the second performance information by machine learning. For example, the generative model is trained in advance by machine learning using training data including the control data for learning and the second performance information. In other words, the trained generative model generates second performance information that is statistically valid for unknown control data based on the latent relationship between the control data and the second performance information in multiple training data.

[0105] In a specific example (Aspect 2) of Aspect 1, the control data includes instrument information specifying the type of instrument corresponding to the second part. According to the above aspect, it is possible to generate second performance information for a part corresponding to an instrument of a type specified by the instrument information. Therefore, for example, in a form in which the instrument information specifies an instrument of a type selected by a user, it is possible to generate a second part according to the user's intention or preference. In addition, by changing the type of instrument specified by the instrument information, it is possible to generate second parts of various instruments using a single generative model.

[0106] In a specific example (Aspect 3) of Aspect 1 or Aspect 2, after the second performance information is generated, the control data including the first performance information and the second performance information is processed by the generative model to generate third performance information representing a third part corresponding to an instrument different from the instruments corresponding to the first part and the second part. According to the above aspect, after the second performance information is generated, the control data including the first performance information and the second performance information is processed by the generative model to generate third performance information representing the third part. Therefore, it is possible to generate third performance information for the third part that is musically consistent with both the first part and the second part.

[0107] In a specific example (Aspect 4) of any one of Aspects 1 to 3, the first performance information is a time series of tokens representing the first part, and the second performance information is a time series of tokens representing the second part. In the above aspects, the first performance information, which is a time series of tokens representing the first part, is processed by a generative model to generate second performance information, which is a time series of tokens representing the second part. Therefore, by utilizing a generative model suitable for processing a time series of multiple tokens, an effective second part that is musically consistent with the first part can be generated.

[0108] In a specific example (Aspect 5) of Aspect 4, the first performance information is further generated from first music piece data representing the time series of notes constituting the first part, and second music piece data representing the time series of notes constituting the second part is generated from the second performance information. In the above aspects, the first performance information is generated from the first music piece data representing the time series of notes constituting the first part. Therefore, existing music piece data can be used to generate the second performance information. Also, the second music piece data representing the time series of notes constituting the second part is generated from the second performance information. Therefore, the musical sounds of the second part can be generated by an existing sound source that generates musical sounds according to music piece data.

[0109] "(First / second) music piece data" is, for example, data in MIDI format that specifies the pitch and duration of each note that constitutes a piece of music. (First / second) performance information is also expressed as intermediate or alternative data that has a different format from music piece data.

[0110] In a specific example (aspect 6) of aspect 5, the second music piece data is data representing a time sequence of notes constituting the first part and the second part, and further, score data representing a score of a music piece including the first part and the second part is generated from the second music piece data. According to the above aspect, the score can be effectively used for an edited music piece including the first part and the second part. Specifically, the score data is used, for example, for displaying the score or automatically playing the music piece.

[0111] An information processing system according to one embodiment (embodiment 7) of the present disclosure includes an information acquisition unit that acquires first performance information representing a first part in a musical piece corresponding to one or more instruments, and an information generation unit that processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second part in the musical piece corresponding to an instrument different from the one or more instruments.

[0112] A program according to one embodiment (embodiment 8) of the present disclosure causes a computer system to function as an information acquisition unit that acquires first performance information representing a first part in a musical piece corresponding to one or more instruments, and an information generation unit that processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second part in the musical piece corresponding to an instrument different from the one or more instruments. [Explanation of symbols]

[0113] 100...information processing system, 200...machine learning system, 11,31...control device, 12,32...storage device, 13,33...communication device, 14...operation device, 15...sound source device, 16...sound emission device, 21...pre-processing unit, 22...information acquisition unit, 23...information generation unit, 24...post-processing unit, 25...musical score generation unit, 41...training data generation unit, 42...training processing unit.

Claims

[Claim 1] obtaining first performance information representing a first part corresponding to one or more instruments in a musical piece; Processing the control data including the first performance information with a machine learning trained generative model to generate second performance information representing a second part in the musical piece corresponding to an instrument different from the one or more instruments. An information processing method implemented by a computer system.

Citation Information

Patent Citations

  • Electronic apparatus, information processing method, and program

    JP2019159146A