Information processing method
A machine learning-based generative model automates the process of rearranging musical pieces into different parts, reducing the effort and expertise needed, enabling users to create diverse and instrument-specific arrangements.
Patent Information
- Application Number
- JP2023191477
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-21
AI Technical Summary
The process of arranging a piece of music into multiple parts corresponding to different instruments or rearranging it to fit a specific instrument requires specialized musical knowledge and significant effort, especially when the number of parts differs from the original piece.
An information processing method using a generative model trained by machine learning to rearrange a musical piece from N parts into M parts, where N and M differ, by processing control data to generate performance information representing the rearranged piece, utilizing a system with pre-processing, information acquisition, generation, and post-processing units.
Reduces the burden of musical arrangement by automating the process, allowing users to create musically diverse pieces with different numbers of parts and instruments according to their preferences, while maintaining musical consistency.
Smart Images

Figure 2025079062000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to techniques for generating information representing music. [Background technology]
[0002] Various techniques for editing existing music have been proposed. For example, Patent Document 1 discloses a configuration in which each note of an existing piece of music is judged to be acceptable or unacceptable according to a simplification level, and a simplified score is generated from the notes that are judged to be acceptable. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2007-241026 A Summary of the Invention [Problem to be solved by the invention]
[0004] There is a demand for arranging a piece of music consisting of multiple parts corresponding to different instruments, based on a specific part of a piece of music. There is also a demand for arranging a piece of music consisting of multiple parts, based on an existing piece of music, with parts corresponding to a specific instrument. In any of the arrangements exemplified above, specialized knowledge of music and a great deal of effort on the part of the producer are required. In consideration of the above circumstances, one aspect of the present disclosure aims to reduce the burden of arranging a piece of music whose total number of parts is different from that of an existing piece of music. [Means for solving the problem]
[0005] In order to solve the above problems, an information processing method according to one embodiment of the present disclosure acquires first performance information representing a first musical piece consisting of N parts (N is a natural number equal to or greater than 1), and processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second musical piece obtained by arranging the first musical piece into M parts (M is a natural number equal to or greater than 1 and different from N).
[0006] An information processing system according to one embodiment of the present disclosure includes an information acquisition unit that acquires first performance information representing a first musical piece composed of N parts (N is a natural number equal to or greater than 1), and an information generation unit that processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second musical piece obtained by arranging the first musical piece into M parts (M is a natural number equal to or greater than 1 and different from N).
[0007] A program according to one embodiment of the present disclosure causes a computer system to function as an information acquisition unit that acquires first performance information representing a first musical piece composed of N parts (N is a natural number equal to or greater than 1), and an information generation unit that processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second musical piece obtained by arranging the first musical piece into M parts (M is a natural number equal to or greater than 1 and different from N). [Brief description of the drawings]
[0008] [Figure 1] 1 is a block diagram illustrating a configuration of an information processing system according to a first embodiment. [Diagram 2] FIG. 1 is an explanatory diagram illustrating an overview of processing by an information processing system. [Diagram 3] FIG. 2 is a block diagram illustrating an example of a functional configuration of an information processing system. [Figure 4] FIG. [Diagram 5] FIG. 2 is an explanatory diagram of tokens constituting performance information. [Figure 6] FIG. 4 is a schematic diagram of control data. [Figure 7] 4 is a schematic diagram of performance information generated from control data. FIG. [Figure 8] 13 is a flowchart of an arrangement process. [Figure 9] FIG. 11 is an explanatory diagram relating to an overview of processing by an information processing system according to a second embodiment; [Figure 10] FIG. 11 is a block diagram illustrating a functional configuration of an information processing system according to a second embodiment. [Figure 11] FIG. 13 is a block diagram illustrating the configuration of a machine learning system according to a third embodiment. [Figure 12] FIG. 1 is a block diagram illustrating an example of the functional configuration of a machine learning system. [Figure 13] 13 is a flowchart of a training data generation process. [Figure 14] 13 is a flowchart of a training process. [Figure 15] FIG. 13 is a block diagram illustrating an example of a functional configuration of an information processing system according to a fourth embodiment. [Figure 16] FIG. 13 is a schematic diagram of control data in the fourth embodiment. [Figure 17] FIG. 13 is a block diagram illustrating an example of a functional configuration of an information processing system according to a fifth embodiment. [Figure 18] FIG. 13 is a schematic diagram of performance information in a modified example. [Figure 19] FIG. 11 is a block diagram illustrating a functional configuration of an information processing system according to a modified example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] A: First embodiment 1 is a block diagram illustrating an example of the configuration of an information processing system 100 in the first embodiment. The information processing system 100 is a computer system that generates a second piece of music from a first piece of music, and is realized by an information device such as a smartphone, a tablet terminal, or a personal computer.
[0010] 2 is an explanatory diagram regarding an overview of processing by the information processing system 100. The first musical piece is a solo piece composed of one performance part corresponding to a specific instrument. The second musical piece is an ensemble composed of M performance parts (M is a natural number equal to or greater than 2) corresponding to different instruments. That is, the information processing system 100 of the first embodiment generates an ensemble from a solo piece.
[0011] A "performance part" is a part (vocal part) that constitutes a piece of music, and means, for example, a group of one or more performers (or a group of one or more instruments) that play common notes. For example, performance parts are distinguished by instrument, range, or musical role (e.g., main melody / minor melody). A "solo piece" is a piece of music that is composed of a single performance part. Typically, a "solo piece" is a piece of music that is performed by a single performer, but in this disclosure, the concept of a "solo piece" also includes a piece of music that is composed of a single performance part including multiple performers. On the other hand, an "ensemble piece" is a piece of music that is composed of multiple performance parts.
[0012] As described above, the number of performance parts differs between the first and second pieces of music, but musical elements such as the theme (e.g., main melody) or impression are common or similar between the first and second pieces of music. In other words, the second piece of music is an arrangement of the first piece of music.
[0013] The types of instruments corresponding to the performance parts of the first piece of music are different from the types of instruments corresponding to each of the M performance parts of the second piece of music. The first piece of music is composed of, for example, an acoustic piano performance part. The second piece of music is composed of, for example, three performance parts (M=3) including an acoustic guitar, a trumpet, and a drum set.
[0014] The information processing system 100 generates music data Dy from music data Dx. The music data Dx is time-series data representing a first music piece. Specifically, the music data Dx represents a time series of notes that make up the first music piece. On the other hand, the music data Dy is time-series data representing a second music piece. Specifically, the music data Dy represents a time series of notes that make up each performance part of the second music piece. Note that the music data Dx is an example of "first music data" and the music data Dy is an example of "second music data."
[0015] The music data Dx and the music data Dy specify the pitch and the sounding period for each note that constitutes each performance part. The pitch is one of a plurality of discretely set scale notes (e.g., a note number). The sounding period is specified, for example, by the start point and duration (or end point) of the note. The music data Dx and the music data Dy are, for example, music files that comply with the MIDI (Musical Instrument Digital Interface) standard.
[0016] 1, the information processing system 100 includes a control device 11, a storage device 12, a communication device 13, an operation device 14, a sound source device 15, and a sound emission device 16. The information processing system 100 may be realized as a single device, or may be realized as a plurality of devices configured separately from each other.
[0017] The control device 11 is composed of one or more processors that control each element of the information processing system 100. For example, the control device 11 is composed of one or more types of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a sound processing unit (SPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC). The communication device 13 communicates with an external device via a communication network such as the Internet.
[0018] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. For example, the storage device 12 stores music piece data Dx of a first music piece. The storage device 12 is configured with a known recording medium such as a magnetic recording medium or a semiconductor recording medium. The storage device 12 may be configured with a combination of multiple types of recording media. In addition, a portable recording medium that is detachable from the information processing system 100, or a recording medium to which the control device 11 can write or read via a communication network (e.g., cloud storage) may be used as the storage device 12.
[0019] The operation device 14 is an input device that accepts instructions from a user. The operation device 14 is, for example, an operator operated by the user, or a touch panel that detects contact by the user. The user can specify M types of instruments corresponding to each performance part of the second musical piece by operating the operation device 14. For example, the user operates the operation device 14 to select M types of instruments from a plurality of candidates prepared in advance. Note that the operation device 14, which is separate from the information processing system 100, may be connected to the information processing system 100 by wire or wirelessly.
[0020] The sound source device 15 generates an audio signal representing a waveform of a musical tone designated by the music data Dx or Dy. The function of the sound source device 15 may be realized by the control device 11 executing a program.
[0021] The sound emitting device 16 reproduces sound waves under the control of the control device 11. The sound emitting device 16 is an output device such as a speaker or a headphone. Specifically, the sound emitting device 16 reproduces musical sounds represented by an acoustic signal generated by the sound source device 15. Note that the sound source device 15 or the sound emitting device 16, which are separate from the information processing system 100, may be connected to the information processing system 100 by wire or wirelessly.
[0022] 3 is a block diagram illustrating an example of a functional configuration of the information processing system 100. The control device 11 executes a program stored in the storage device 12 to realize a plurality of functions (a pre-processing unit 21, an information acquisition unit 22, an information generation unit 23, and a post-processing unit 24) for generating music data Dy of a second music piece from music data Dx of a first music piece.
[0023] The pre-processing unit 21 generates performance information X from the music data Dx. The performance information X is time-series data representing the first music piece. The performance information X is also expressed as intermediate or alternative data having a different format from the music data Dx. The performance information X is an example of the "first performance information."
[0024] FIG. 4 is a schematic diagram of performance information X. The performance information X is a time series of multiple tokens T that represent a first piece of music. The tokens T are units that make up the performance information X. The performance information X is described, for example, by a series of character strings. In FIG. 4, the performance information X is illustrated as a character string spanning multiple lines. Each token T is separated by a blank (for example, a half-width space).
[0025] 5 is an explanatory diagram of the tokens T. The multiple tokens T constituting the performance information X include a note token Ta, a beat token Tb, and an auxiliary token Tc.
[0026] The note token Ta is a token T that represents a note of the first music piece. Specifically, the note token Ta is divided into a performance token Ta1 and a note value token Ta2. The performance token Ta1 specifies the type of instrument used to play the note and the pitch of the note (e.g., note number). For example, a performance token Ta1 written as "acp_68" means a note with a pitch of 68 that should be played by an acoustic piano (acp: acoustic piano). The performance token Ta1 is also expressed as identification information that specifies the type of instrument. In other words, each performance token Ta1 in the performance information X indicates the type of instrument corresponding to the performance part of the first music piece.
[0027] The note value token Ta2 specifies the note value (duration) of the note. For example, the note value token Ta2 written as "len_L" (L is a natural number) means a time length equivalent to L unit times. The unit time is a time length corresponding to the beat interval of the first piece of music. Specifically, the unit time is a time length equivalent to, for example, 1 / 12 beat of the first piece of music. The combination of the performance token Ta1 and note value token Ta2 expresses one note with a specified instrument, pitch, and note value.
[0028] The beat token Tb is a token T that represents a beat of the first piece of music. Specifically, a specific point in time in the first piece of music is represented by the beat token Tb. Specifically, the beat token Tb is divided into a bar token Tb1 (bar), a beat token Tb2 (beat), and a position token Tb3 (pos). The beat token Tb that represents a specific point in time in the first piece of music is placed at a position corresponding to that point in the time series of tokens T in the performance information X.
[0029] The bar token Tb1 is a token T that signifies a bar line of the first piece of music (i.e., the first beat of each bar). The beat token Tb2 is a token T that signifies each beat of the first piece of music. The position token Tb3 is a token T that expresses a specific point in time in the first piece of music. For example, the position token Tb3 written as "pos_K" (K=1 to 11) signifies a point in time when a time equivalent to K unit times has elapsed from the bar line (bar token Tb1) or beat (beat token Tb2) immediately preceding the position token Tb3.
[0030] The position token Tb3 is used to express the position of a note, for example. Specifically, the position of one note is expressed by the position token Tb3 arranged immediately before the note token Ta (performance token Ta1 and note value token Ta2) that represents the note. For example, the notation "pos_6 acp_48 len_6" in the performance information X means that a note with a pitch of 48 and a note value of 6 played by an acoustic piano (acp) starts at a point (pos_6) six unit times after the previous beat. Note that if the start point of the note coincides with a bar line (bar token Tb1) or a beat (beat token Tb2), the position token Tb3 is omitted. As shown in the above example, the sounding conditions (pitch, note value, and position) of one note are specified by the performance token Ta1, note value token Ta2, and position token Tb3.
[0031] The auxiliary token Tc is a token T that represents various information related to the performance of the first musical piece. Specifically, the auxiliary token Tc is divided into a section token Tc1 (section), a key token Tc2 (key), and a tempo token Tc3 (tempo). The section token Tc1 is a token T that indicates the start point of each structural section of the first musical piece. The structural section is a section obtained by dividing the first musical piece on the time axis according to musical meaning. For example, each section such as an intro, an A melody, a B melody, a chorus, and an outro is exemplified as a structural section. The section token Tc1 is placed at a position in the performance information X that corresponds to the start point of a structural section in the first musical piece.
[0032] The key token Tc2 is a token T that represents the key of the first music piece. The numeric value specified by the key token Tc2 is a number for identifying the key. The key token Tc2 is placed at a position in the performance information X that corresponds to a time point at which the key changes in the first music piece. The speed token Tc3 is a token T that represents the tempo of the first music piece. The numeric value specified by the speed token Tc3 indicates the tempo (BPM: Beats Per Minute). The speed token Tc3 is placed at a position in the performance information X that corresponds to a time point at which the tempo changes in the first music piece.
[0033] The above are specific examples of the tokens T that make up the performance information X. The pre-processing unit 21 in Fig. 3 converts the music data Dx into the performance information X under the rules exemplified above. According to the first embodiment, the widely used existing music data Dx can be used to generate music data Dy (performance information Y described later).
[0034] The information acquisition unit 22 acquires performance information X. The information acquisition unit 22 in the first embodiment generates control data C including the performance information X and M pieces of instrument information Q1 to QM. Each piece of instrument information Qm (m=1 to M) is information specifying the type of instrument for each of the M performance parts of the second musical piece. Specifically, the M types of instruments specified by the user through an operation on the operation device 14 are specified by the M pieces of instrument information Q1 to QM.
[0035] 6 is a schematic diagram of control data C. The information acquiring unit 22 generates the control data C by adding M pieces of musical instrument information Q1 to QM to the beginning of performance information X. Each piece of musical instrument information Qm is added to the performance information X as, for example, one token T. Note that the position at which the musical instrument information Qm is added to the performance information X is arbitrary.
[0036] The information generating unit 23 in FIG. 3 generates performance information Y from the control data C. The performance information Y is time-series data representing the second music piece. FIG. 7 is a schematic diagram of the performance information Y. The performance information Y is a time-series of multiple tokens T representing M performance parts of the second music piece. Like the performance information X, the performance information Y is described, for example, as a series of character strings. Note that the performance information Y is an example of "second performance information."
[0037] The types and meanings of each token T constituting the performance information Y are the same as those of each token T in the performance information X described above with reference to Fig. 5. That is, the multiple tokens T constituting the performance information Y include note tokens Ta (Ta1, Ta2), beat tokens Tb (Tb1, Tb2, Tb3), and auxiliary tokens Tc (Tc1, Tc2, Tc3), as in the example of Fig. 5. Note that the notation "acg" of the performance token Ta1 shown in Fig. 5 means acoustic guitar, and the notation "trp" means trumpet.
[0038] As mentioned above, the second piece of music is composed of three parts: an acoustic guitar, a trumpet, and a drum set. The performance sounds of the drum set are not considered to have pitch or value. Therefore, in the performance information Y shown in FIG. 7, the performance token Ta1 does not include a pitch designation, and the value token Ta2 is not used.
[0039] Furthermore, the performance token Ta1 of the drum set specifies the type of each percussion instrument that constitutes the drum set. For example, a performance token Ta1 written as "bass" indicates a bass drum performance, a performance token Ta1 written as "hhcls" indicates a hi-hat performance, and a performance token Ta1 written as "snare" indicates a snare drum performance.
[0040] As shown in the above example, each performance token Ta1 of the performance information Y indicates the type of musical instrument corresponding to each of the M performance parts of the second musical piece. The second musical piece represented by the performance information Y is composed of M performance parts corresponding to the musical instruments specified by each musical instrument information Qm.
[0041] A trained generative model G is used for generating the performance information Y by the information generating unit 23. The generative model G is a statistical model that has learned the relationship between the control data C and the performance information Y by prior machine learning. The information generating unit 23 processes the control data C by the trained generative model G to generate the performance information Y.
[0042] The generative model G is realized by a combination of a program that causes the control device 11 to execute a calculation to generate performance information Y from control data C, and a number of variables (e.g., bias and weighting values) that are applied to the calculation. The numerical values of each of the multiple variables are set in advance by machine learning.
[0043] For example, a transformer, which is an encoder-decoder model including a self-attention mechanism (specifically, a multi-head attention mechanism), is used as the generative model G. The transformer is disclosed, for example, in Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin, "Attention Is All You Need", 31st Conference on Neural Information Processing Systems (NIPS 2017). In addition, MEGA (Moving Average Equipped Gated Attention) may be adopted as the generative model G. MEGA is disclosed in Ma, C. Zhou, X. Kong, J. He, L. Gui, G. Neubig, J. May, and L. Zettlemoyer. "Mega: moving average equipped gated attention", arXiv:2209.10655, 2022.
[0044] The format of the performance information X and the performance information Y, which are configured as a time series of multiple tokens T, is suitable for processing by the generation model G. As described above, according to the first embodiment, by utilizing the generation model G suitable for processing the time series of multiple tokens T, an effective second piece of music that is musically consistent with the first piece of music can be generated.
[0045] The post-processing unit 24 in Fig. 3 generates the music data Dy of the second music piece from the performance information Y. Specifically, the post-processing unit 24 generates the music data Dy of the second music piece from the performance information Y under the rules described above with reference to Fig. 5. Since the music data Dy is generated from the performance information Y as described above, musical tones of the second music piece can be generated by an existing sound source device 15 capable of processing the music data Dy.
[0046] 8 is a flowchart of the process (hereinafter referred to as "arrangement process") in which the control device 11 generates music data Dy from music data Dx. For example, the arrangement process is started in response to an operation of the operation device 14 by the user.
[0047] When the arrangement process is started, the control device 11 (preprocessing unit 21) generates performance information X from the music piece data Dx (Sa1). The control device 11 (information acquiring unit 22) generates M pieces of musical instrument information Q1 to QM that specify different types of musical instruments (Sa2). The control device 11 (information acquiring unit 22) acquires the performance information X and generates control data C including the M pieces of musical instrument information Q1 to QM and the performance information X (Sa3).
[0048] The control device 11 (information generating unit 23) processes the control data C using the generation model G to generate performance information Y (Sa4). The control device 11 (post-processing unit 24) generates music piece data Dy of the second music piece from the performance information Y (Sa5).
[0049] As described above, in the first embodiment, performance information Y of a second musical piece consisting of M performance parts is generated from performance information X of a first musical piece consisting of one performance part. This reduces the burden of arranging a second musical piece having a different total number of performance parts from an existing first musical piece. Specifically, for example, a musically diverse second musical piece consisting of M performance parts can be arranged from a first musical piece consisting of one performance part that is simply created by a user. In other words, according to the first embodiment, a unique customer experience can be provided to a user, in which the second musical piece is obtained by arranging an existing first musical piece.
[0050] In particular, in the first embodiment, the performance information Y of the second musical piece can be generated corresponding to the type of musical instrument specified by the musical instrument information Qm. Since the musical instrument information Qm specifies the musical instrument selected by the user, the second musical piece can be generated according to the user's intention or preference. In addition, by changing the type of musical instrument specified by the musical instrument information Qm, the second musical piece can be generated for various musical instruments using the single generation model G.
[0051] B: Second embodiment A second embodiment will be described. Note that, for elements having the same functions as those in the first embodiment in each of the following exemplary aspects, the same reference numerals as those in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.
[0052] FIG. 9 is an explanatory diagram of an overview of processing by the information processing system 100 of the second embodiment. The first musical piece is an ensemble composed of N performance parts (N is a natural number equal to or greater than 2) corresponding to different musical instruments. The second musical piece is a solo piece composed of one performance part corresponding to a specific musical instrument. That is, the information processing system 100 of the first embodiment generates an ensemble from a solo piece, whereas the information processing system 100 of the second embodiment generates a solo piece from an ensemble. As described above, the number of performance parts differs between the first musical piece and the second musical piece, but musical elements such as a theme (e.g., main melody) or an impression are common or similar between the first musical piece and the second musical piece. That is, similar to the first embodiment, the second musical piece is an arrangement of the first musical piece.
[0053] The types of instruments corresponding to the N performance parts of the first music piece are different from the types of instruments corresponding to the performance parts of the second music piece. The first music piece is composed of three performance parts (M=3), for example, an acoustic guitar, a trumpet, and a drum set, and the second music piece is composed of a performance part, for example, an acoustic piano. The information processing system 100 of the second embodiment generates music data Dy of the second music piece from music data Dx of the first music piece, as in the first embodiment.
[0054] The configuration of the information processing system 100 of the second embodiment is similar to the configuration of the information processing system 100 of the first embodiment illustrated in Fig. 1. As illustrated in Fig. 10, the control device 11 of the second embodiment executes a program stored in the storage device 12 to realize a plurality of functions (a pre-processing unit 21, an information acquiring unit 22, an information generating unit 23, and a post-processing unit 24) similar to those of the first embodiment.
[0055] The functions of the elements illustrated in Fig. 10 are the same as those in the first embodiment. The pre-processing unit 21 generates performance information X from the music data Dx of the first music piece. As described above, the performance information X is a time series of multiple tokens T representing the first music piece. Specifically, the performance information X (e.g., each performance token Ta1) represents the type of musical instrument corresponding to each of the N performance parts.
[0056] The information acquisition unit 22 acquires performance information X. The information acquisition unit 22 of the second embodiment generates control data C including the performance information X and instrument information Q. The instrument information Q is information specifying the type of instrument corresponding to one performance part of the second musical piece. Specifically, the instrument specified by the user by operating the operation device 14 is specified by the instrument information Q. The information acquisition unit 22 generates the control data C by adding the instrument information Q to the beginning of the performance information X, as in the example of FIG. 6.
[0057] The information generating unit 23 generates performance information Y from the control data C. As described above, the performance information Y is a time series of multiple tokens T representing the second music piece. Furthermore, the performance information Y (e.g., each performance token Ta1) represents the type of musical instrument corresponding to one performance part of the second music piece.
[0058] Specifically, the information generating unit 23 processes the control data C with a trained generation model G to generate performance information Y. The generation model G is a statistical model that has learned the relationship between the control data C and the performance information Y through prior machine learning. The post-processing unit 24 generates music piece data Dy of the second music piece from the performance information Y.
[0059] The procedure of the editing process is the same as that of the first embodiment. In the second embodiment, performance information Y of a second musical piece arranged into one performance part is generated from performance information X of a first musical piece consisting of N performance parts. Therefore, as in the first embodiment, the load of arranging a second musical piece having a different total number of performance parts from an existing first musical piece can be reduced. Specifically, for example, from an existing first musical piece consisting of N performance parts, a second musical piece with one performance part suitable for a user to practice playing a specific instrument can be arranged. That is, according to the second embodiment, as in the first embodiment, a unique customer experience can be provided to a user, in which the second musical piece is arranged from an existing first musical piece.
[0060] The first and second embodiments are collectively expressed as a system that generates a second musical piece consisting of M performance parts (M is a natural number greater than or equal to 1) from a first musical piece consisting of N performance parts (N is a natural number greater than or equal to 1). The total number N of performance parts in the first musical piece is different from the total number M of performance parts in the second musical piece. The first embodiment is a form in which the total number N of performance parts in the first musical piece is 1, and the second embodiment is a form in which the total number M of performance parts in the second musical piece is 1.
[0061] In the first and second embodiments, the types of instruments corresponding to the N performance parts of the first music piece are different from the types of instruments corresponding to the M performance parts of the second music piece. Therefore, it is possible to arrange a second music piece with a different type of instrument from that of the first music piece. Note that the types of instruments may be the same between some or all of the N performance parts of the first music piece and some or all of the M performance parts of the second music piece. In other words, the N types of instruments corresponding to the first music piece and the M types of instruments corresponding to the second music piece may be the same in part or all. Note that the total number N of performance parts of the first music piece and the total number M of performance parts of the second music piece may be the same.
[0062] C: Third embodiment 11 is a block diagram illustrating the configuration of a machine learning system 200 in the third embodiment. The machine learning system 200 is a computer system that establishes the above-mentioned generative model G used by the information processing system 100 through machine learning. The machine learning system 200 includes a control device 31, a storage device 32, and a communication device 33. Note that the machine learning system 200 may be realized as a single device, or may be realized as a plurality of devices configured separately from each other.
[0063] The control device 31 is composed of one or more processors that control each element of the machine learning system 200. For example, the control device 31 is composed of one or more types of processors such as a CPU, a GPU, an SPU, a DSP, an FPGA, or an ASIC.
[0064] The communication device 33 communicates with an external device via a communication network such as the Internet. For example, the communication device 33 communicates with the information processing system 100. The generative model G established by the machine learning system 200 is provided to the information processing system 100 by the communication device 33.
[0065] The storage device 32 is one or more memories that store the programs executed by the control device 31 and various data used by the control device 31. The storage device 32 is configured with a known recording medium such as a magnetic recording medium or a semiconductor recording medium. The storage device 32 may be configured with a combination of multiple types of recording media. In addition, a portable recording medium that is detachable from the information processing system 100, or a recording medium (e.g., cloud storage) to which the control device 31 can write or read via a communication network may be used as the storage device 32.
[0066] The storage device 32 stores a plurality of material data R corresponding to different musical pieces. Each of the plurality of material data R includes musical piece data Rx and musical piece data Ry. The musical piece data Rx is time-series data representing a musical piece for machine learning (hereinafter referred to as a "first reference piece"). The first reference piece is composed of N performance parts. The musical piece data Rx represents a time series of notes for each of the N performance parts. Furthermore, the musical piece data Rx specifies the type of instrument for each of the N performance parts.
[0067] The music data Ry is time-series data representing a music piece for machine learning (hereinafter referred to as the "second reference music piece"). The second reference music piece is composed of M performance parts. The music data Ry represents a time series of notes for each of the M performance parts. The music data Ry also specifies the type of instrument for each of the M performance parts.
[0068] The first and second reference pieces corresponding to one piece of material data R have different numbers of performance parts, but have common or similar musical elements such as a theme (e.g., main melody) or impression. In other words, one of the second and first reference pieces is an arrangement of the other.
[0069] The song data Rx and the song data Ry are, for example, music files that comply with the MIDI (Musical Instrument Digital Interface) standard. In generating the generative model G in the first embodiment (N=1), the song data Ry is, for example, data for karaoke of a song created in the past, and the song data Rx is data representing a score (for example, a practice score) created for solo performance of the song. On the other hand, in generating the generative model G in the second embodiment (M=1), the song data Rx is, for example, data for karaoke of a song created in the past, and the song data Ry is data representing a score created for solo performance of the song.
[0070] 12 is a block diagram illustrating an example of a functional configuration of the machine learning system 200. The control device 31 executes a program stored in the storage device 32 to realize a plurality of functions (a training data generation unit 41, a training processing unit 42) for establishing the generative model G.
[0071] The training data generation unit 41 generates multiple pieces of training data Z to be used for machine learning of the generative model G. Each of the multiple pieces of training data Z is composed of a combination of training control data Ct and training performance information Yt. The control data Ct is data in the same format as the above-mentioned control data C. The performance information Yt is data in the same format as the above-mentioned performance information Y. The training data generation unit 41 generates multiple pieces of training data Z from multiple pieces of material data R stored in the storage device 32. The training processing unit 42 establishes the generative model G by machine learning using the multiple pieces of training data Z.
[0072] 13 is a flowchart of a process (hereinafter referred to as "training data generation process") in which the control device 31 (training data generation unit 41) generates multiple training data Z. For example, the training data generation process is started in response to an instruction from an administrator of the machine learning system 200. When the training data generation process is started, the control device 31 selects one of multiple pieces of material data R stored in the storage device 32 (hereinafter referred to as "selected material data R") (Sb1).
[0073] The control device 31 causes the music data Rx and music data Ry contained in the selected material data R to correspond in time (Sb2). That is, alignment is performed between the music data Rx and music data Ry. Specifically, the control device 31 extracts sections that correspond musically to each other in each of the music data Rx and music data Ry, and adjusts the position of each section on the time axis so that the positions of corresponding notes on the time axis match.
[0074] The control device 31 generates performance information Yt from the adjusted music piece data Ry (Sb3). The process of generating the performance information Yt from the music piece data Ry is similar to the process of the pre-processing unit 21 generating the performance information X from the music piece data Dx.
[0075] The control device 31 generates performance information X from the adjusted music data Rx (Sb4). The process of generating performance information X from music data Rx is similar to the process of generating performance information X from music data Dx by the pre-processing unit 21. In addition, the control device 31 generates M pieces of instrument information Q1-QM that specify each instrument specified by the music data Ry (Sb5).
[0076] The control device 31 generates training control data Ct including the performance information X and M pieces of musical instrument information Q1 to QM (Sb6). Then, the control device 31 generates training data Z by associating the control data Ct with the performance information Yt, and stores the training data Z in the storage device 32 (Sb7).
[0077] The control device 31 determines whether or not the above processes (Sb2 to Sb7) have been performed for all of the material data R stored in the storage device 32 (Sb8). If there is unprocessed material data R (Sb8: NO), the control device 31 transitions the process to step Sb1. That is, the control device 31 selects the unprocessed material data R as new selected material data R (Sb1). As described above, the generation of training data Z (Sb2 to Sb7) is repeated for each of the multiple material data R. On the other hand, if all of the material data R has been processed (Sb8: YES), the control device 31 ends the training data generation process.
[0078] 14 is a flowchart of a process (hereinafter referred to as the "training process") in which the control device 31 (training processing unit 42) establishes a generative model G through machine learning using a plurality of training data Z. After the training data generation process is executed, the training process is started, for example, in response to an instruction from an administrator of the machine learning system 200. The training process is an example of a method for generating a generative model G.
[0079] When the training process is started, the control device 31 selects one of the multiple training data Z (hereinafter referred to as "selected training data Z") stored in the storage device 32 (Sc1). As illustrated in FIG. 12, the control device 31 generates performance information Y by processing control data Ct of the selected training data Z using an initial or provisional generation model G (hereinafter referred to as "provisional model G0") (Sc2). The control device 31 calculates a loss function that represents the error between the performance information Y generated by the provisional model G0 and the performance information Yt of the selected training data Z (Sc3). The control device 31 updates multiple variables of the provisional model G0 so that the loss function is reduced (ideally minimized) (Sc4).
[0080] The control device 31 determines whether a predetermined termination condition is satisfied (Sc5). The termination condition is, for example, that the loss function falls below a predetermined threshold, or that the amount of change in the loss function falls below a predetermined threshold. If the termination condition is not satisfied (Sc5: NO), the control device 31 selects the unselected training data Z stored in the storage device 32 as new selected training data Z (Sc1). That is, the process of updating multiple variables of the provisional model G0 (Sc2 to Sc4) is repeated until the termination condition is satisfied (Sc5: YES).
[0081] When the termination condition is met (Sc5: YES), the control device 31 ends the training process. The provisional model G0 at the time when the termination condition is met is determined as the trained generative model G.
[0082] As can be understood from the above explanation, the generative model G learns the latent relationship between the control data Ct and the performance information Yt in multiple training data Z. Therefore, the trained generative model G outputs the performance information Y that is statistically appropriate for the unknown control data C based on the above relationship.
[0083] The trained generative model G is transmitted from the communication device 33 to the information processing system 100. The control device 11 of the information processing system 100 receives the generative model G via the communication device 13, and stores the generative model G in the storage device 12. The generative model G provided by the above procedure is used in the musical arrangement process illustrated in FIG.
[0084] D: Fourth embodiment 15 is a block diagram illustrating a functional configuration of an information processing system 100 in the fourth embodiment. In the fourth embodiment, the total number N of performance parts of the first music piece and the total number M of performance parts of the second music piece are generalized. That is, the configuration of the fourth embodiment is applicable to both the first and second embodiments.
[0085] The information acquisition unit 22 of the fourth embodiment acquires a difficulty level F in addition to the performance information X and the instrument information Qm. The difficulty level F is an index of the difficulty of playing the second piece of music. The difficulty level F is also expressed as an index of the technical level required to play the second piece of music, or an index of the complexity of the second piece of music. For example, the user operates the operation device 14 to select the difficulty level F from a plurality of candidate values. The information acquisition unit 22 acquires the difficulty level F designated by the user.
[0086] 16 is a schematic diagram of the control data C in the fourth embodiment. The control data C includes a difficulty level F in addition to the performance information X and each piece of musical instrument information Qm. Specifically, the information acquiring unit 22 generates the control data C by adding the difficulty level F and the musical instrument information Qm to the beginning of the performance information X. The difficulty level F is added to the performance information X as, for example, one token T, similar to each piece of musical instrument information Qm. Note that the position at which the difficulty level F is added to the performance information X or the musical instrument information Qm is arbitrary.
[0087] As in the first embodiment, the information generating unit 23 processes the control data C using a trained generation model G to generate performance information Y. The second musical piece represented by the performance information Y is a musical piece corresponding to a difficulty level F. That is, even if the performance information X and the musical instrument information Qm are common, the difficulty of playing the second musical piece changes depending on the difficulty level F.
[0088] The musical piece data Ry of the second reference piece in each of the plurality of material data R includes a difficulty level F of the second reference piece. For example, the musical piece data Ry specifies one of a plurality of stages such as beginner / intermediate / advanced as the difficulty level F for the second reference piece. In the training data generation process (FIG. 13), the training data generation unit 41 generates control data C including the difficulty level F in addition to the performance information X and the information on each musical instrument Qm. The generation model G is trained by the training process (FIG. 14) using the training data Z including the difficulty level F. Therefore, the generation model G generates performance information Y of the second musical piece corresponding to the difficulty level F for the control data C including the difficulty level F.
[0089] In the fourth embodiment, the same effects as in the first or second embodiment are achieved. Moreover, in the fourth embodiment, a second piece of music can be generated that corresponds to the difficulty level F included in the control data C. Since the difficulty level F is set according to an instruction from the user, a second piece of music can be generated with a difficulty level F that matches the user's intention. Moreover, by changing the difficulty level F, second pieces of music with various difficulty levels F can be generated using a single generation model G.
[0090] E: Fifth embodiment 17 is a block diagram illustrating a functional configuration of an information processing system 100 in the fifth embodiment. The control device 11 in the fifth embodiment functions as a musical score generating unit 25 in addition to the same elements as those in the first or second embodiment (a pre-processing unit 21, an information acquiring unit 22, an information generating unit 23, and a post-processing unit 24).
[0091] The score generation unit 25 generates score data E from the music data Dy. The score data E is data that represents the score of the second music piece. For example, data in MusicXML format or PDF format is exemplified as the score data E. Any known technology may be used to generate the score data E. The score generation unit 25 displays the score represented by the score data E on a display device (not shown).
[0092] The fifth embodiment also achieves the same effects as the first or second embodiment. In the fifth embodiment, the score data E of the second music piece represented by the performance information Y is generated, so that the user can confirm the contents of the score of the second music piece. The configuration of the fourth embodiment is also applicable to the fifth embodiment.
[0093] F: Variation Specific modified embodiments added to each of the above-mentioned embodiments are exemplified below. Two or more embodiments selected from the following examples may be appropriately combined as long as they are not mutually contradictory.
[0094] (1) In the performance information X or Y, for an instrument that is played with both hands, such as a keyboard instrument, the right hand part and the left hand part may be expressed separately, as shown in the example of the performance information X in Fig. 18. The symbol "R" in the performance token Ta1 in Fig. 18 means the right hand part, and the symbol "L" means the left hand part.
[0095] 18 shows an example of performance information X, but the right hand part and the left hand part are also expressed separately for performance information Y. According to the above embodiment, a second musical piece for a keyboard instrument in which the right hand part and the left hand part are differentiated can be generated from a first musical piece for a keyboard instrument in which the right hand part and the left hand part are differentiated.
[0096] (2) In each of the above-mentioned embodiments, the section of the first musical piece that serves as the unit of processing by the information processing system 100 may be any section. For example, the above-mentioned arrangement process may be performed sequentially or in parallel for each of a plurality of sections (hereinafter referred to as "unit sections") obtained by dividing the first musical piece on the time axis. The musical piece data Dx represents a string of notes corresponding to one unit section of the first musical piece, and the musical piece data Dy represents a string of notes corresponding to one unit section of the second musical piece. Each unit section is, for example, a section of a time length equivalent to a predetermined number of bars in the musical piece (for example, 4 to 8 bars). The musical piece data Dy covering all sections of the second musical piece may be generated by a single arrangement process targeting all sections of the first musical piece.
[0097] (3) In the second embodiment in which a second musical piece with one performance part is generated from a first musical piece with N performance parts, the musical instrument information Q may be omitted. In other words, the musical instruments constituting the performance parts of the second musical piece may be fixed to a specific musical instrument. According to the embodiment in which the control data C includes the musical instrument information Qm as in each of the above-mentioned embodiments, as described above, the second musical piece with various instruments can be generated by the single generation model G by changing the musical instrument information Qm.
[0098] (4) In each of the above-described embodiments, a transformer is exemplified as the generative model G, but the configuration or type of the generative model G is arbitrary. For example, a deep neural network such as a recurrent neural network (RNN) or a long short-term memory (LSTM) may be used as the generative model G. The generative model G may be configured by combining multiple types of statistical models.
[0099] (5) In the second embodiment, the information processing system 100 and the machine learning system 200 are illustrated as separate systems, but the functions of the machine learning system 200 (the training data generation unit 41 and the training processing unit 42) may be incorporated in the information processing system 100.
[0100] (6) In the above-mentioned embodiments, the performance parts of the first and second music pieces correspond to musical instruments, but some or all of the performance parts constituting the first or second music piece may be singing parts corresponding to singing voices. That is, the "musical instrument" in the above-mentioned embodiments may be replaced with a "singer." Also, the "musical instrument" in the above-mentioned embodiments may be replaced with a "sound source" including both an instrument and a singer. That is, the performance parts of the first or second music piece correspond to different types of sound sources.
[0101] (7) The information processing system 100 may be realized by a server device that communicates with a terminal device such as a smartphone, a tablet terminal, or a personal computer. For example, the control device 11 receives the music data Dx and the musical instrument information Q (Q1 to QM) transmitted from the terminal device via the communication device 13, and generates the music data Dy using the music data Dx and the musical instrument information Q. The control device 11 transmits the music data Dy to the terminal device via the communication device 13.
[0102] In addition, in a configuration in which the pre-processing unit 21 that generates the performance information X from the music data Dx is installed in the terminal device, the pre-processing unit 21 is omitted from the information processing system 100 as shown in Fig. 19. That is, the information acquiring unit 22 receives the performance information X and the musical instrument information Q transmitted from the terminal device through the communication device 13, and generates control data C including the performance information X and the musical instrument information Q.
[0103] In addition, in a configuration in which the post-processing unit 24 that generates music data Dy from the performance information Y is installed in the terminal device, the post-processing unit 24 is omitted from the information processing system 100, as illustrated in Fig. 19. That is, the information generating unit 23 transmits the performance information Y generated by the generation model G from the communication device 13 to the terminal device. The terminal device generates the music data Dy from the performance information Y received from the information processing system 100.
[0104] (8) The functions of the information processing system 100 exemplified above are realized by the cooperation of one or more processors constituting the control device 11 and the program stored in the storage device 12, as described above. The program according to the present disclosure can be provided in a form stored in a computer-readable recording medium and installed in a computer. The recording medium is, for example, a non-transitory recording medium, and a good example is an optical recording medium (optical disk) such as a CD-ROM, but also includes any known type of recording medium such as a semiconductor recording medium or a magnetic recording medium. Note that the non-transitory recording medium includes any recording medium except for a transient propagating signal, and does not exclude volatile recording media. In addition, in a configuration in which a distribution device distributes a program via a communication network, the storage medium that stores the program in the distribution device corresponds to the non-transitory recording medium described above.
[0105] G: Notes From the above-described exemplary embodiments, the following configurations can be understood, for example.
[0106] An information processing method according to one aspect (aspect 1) of the present disclosure acquires first performance information representing a first musical piece consisting of N parts (N is a natural number equal to or greater than 1), and processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second musical piece obtained by arranging the first musical piece into M parts (M is a natural number equal to or greater than 1 and different from N). According to the above aspect, second performance information of the second musical piece arranged into M parts is generated from the first performance information of the first musical piece consisting of N parts. This reduces the load of arranging a musical piece whose total number of parts differs from that of an existing musical piece.
[0107] "(First / second) performance information" is data in any format that represents parts that make up a piece of music. For example, performance information is data that represents a time series of notes that correspond to a specific part. For example, performance information is composed of a time series of multiple tokens that represent a specific part of a piece of music. Each token specifies, for example, the sounding conditions (e.g., pitch, position, and duration) of each note that makes up the piece of music.
[0108] A "generative model" is a statistical model of any configuration that has learned the relationship between the control data and the second performance information by machine learning. For example, the generative model is trained in advance by machine learning using training data including the control data for learning and the second performance information. In other words, the trained generative model generates second performance information that is statistically valid for unknown control data based on the latent relationship between the control data and the second performance information in multiple training data.
[0109] Some or all of the types of instruments corresponding to the N parts of the first piece of music may be different from some or all of the types of instruments corresponding to the M parts of the second piece of music, or they may be the same. In other words, the M types of instruments corresponding to the M parts of the second piece of music may or may not include some or all of the N types of instruments corresponding to the N parts of the first piece of music.
[0110] In a specific example (Aspect 2) of Aspect 1, the first performance information indicates the type of instrument corresponding to each of the N parts, and the second performance information indicates the type of instrument corresponding to each of the M parts, and the type of instrument corresponding to each of the N parts is different from the type of instrument corresponding to each of the M parts. According to the above aspect, it is possible to arrange a second musical piece having a different type of instrument from the first musical piece.
[0111] In a specific example (Aspect 3) of Aspect 1 or Aspect 2, the N parts are one part, and the M parts are two or more parts. According to the above aspect, a second piece of music consisting of two or more parts can be generated from a first piece of music (solo piece of music) consisting of one part. Therefore, for example, a musically diverse second piece of music consisting of multiple parts can be arranged from a first piece of music consisting of one part that is simply created by a producer.
[0112] In a specific example (Aspect 4) of Aspect 1 or Aspect 2, the N parts are two or more parts, and the M parts are one part. According to the above aspect, a second piece of music (solo piece of music) consisting of one part can be generated from a first piece of music consisting of two or more parts. Therefore, a second piece of music for solo performance can be generated from an existing first piece of music consisting of multiple parts.
[0113] In a specific example (Aspect 5) of any one of Aspects 1 to 4, the control data includes instrument information specifying the type of instrument for each of the M parts, and the second musical piece is composed of the M parts corresponding to the type of instrument specified by the instrument information. According to the above aspects, it is possible to generate a second musical piece composed of M parts corresponding to the type of instrument specified by the instrument information. Therefore, in a form in which the instrument information specifies an instrument of a type selected by a user, for example, it is possible to generate a second musical piece according to the user's intention or preference. In addition, by changing the type of instrument specified by the instrument information, it is possible to generate a second musical piece with a variety of instruments using a single generative model.
[0114] In a specific example (Aspect 6) of any one of Aspects 1 to 5, the control data includes a difficulty level for performance, and the second musical piece is a musical piece corresponding to the difficulty level. According to the above aspects, it is possible to generate a second musical piece corresponding to the difficulty level included in the control data. Therefore, for example, in a form in which the difficulty level specified by the user is included in the control data, it is possible to generate a second musical piece with a difficulty level in line with the user's intention. Also, by changing the difficulty level, it is possible to generate second musical pieces with various levels of difficulty using a single generation model.
[0115] In a specific example (Aspect 7) of any one of Aspects 1 to 6, the first performance information or the second performance information represents a right hand part and a left hand part in a performance of a keyboard instrument, with distinction therebetween. According to the above aspects, a second musical piece for a keyboard instrument, in which a right hand part and a left hand part are distinguished, can be generated.
[0116] In a specific example (Aspect 8) of any one of Aspects 1 to 7, the first performance information is a time series of tokens representing the first musical piece, and the second performance information is a time series of tokens representing the second musical piece. In the above aspects, the first performance information, which is a time series of tokens representing the first musical piece, is processed by a generative model to generate second performance information, which is a time series of tokens representing the second musical piece. Therefore, by using a generative model suitable for processing a time series of multiple tokens, it is possible to generate an effective second musical piece that is musically consistent with the first musical piece.
[0117] In a specific example (aspect 9) of aspect 8, the first performance information is further generated from first music piece data representing a time series of notes constituting the first music piece, and second music piece data representing a time series of notes constituting the second music piece is generated from the second performance information. In the above aspect, the first performance information is generated from the first music piece data representing the time series of notes of the first music piece. Therefore, existing music piece data can be used to generate the second performance information. Also, the second music piece data representing the time series of notes of the second music piece is generated from the second performance information. Therefore, the musical tones of the second music piece can be generated by an existing sound source that generates musical tones according to music piece data.
[0118] "(First / second) music piece data" is, for example, data in MIDI format that specifies the pitch and duration of each note that constitutes a piece of music. (First / second) performance information is also expressed as intermediate or alternative data that has a different format from music piece data.
[0119] In a specific example (aspect 10) of aspect 9, furthermore, sheet music data representing the sheet music of the second musical piece is generated from the second musical piece data. In the above aspect, because sheet music data representing the sheet music of the second musical piece is generated, the user can confirm the contents of the sheet music of the second musical piece.
[0120] An information processing system according to one embodiment (embodiment 11) of the present disclosure includes an information acquisition unit that acquires first performance information representing a first musical piece consisting of N parts (N is a natural number equal to or greater than 1), and an information generation unit that processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second musical piece obtained by arranging the first musical piece into M parts (M is a natural number equal to or greater than 1 and different from N).
[0121] A program according to one embodiment (embodiment 12) of the present disclosure causes a computer system to function as an information acquisition unit that acquires first performance information representing a first musical piece consisting of N parts (N is a natural number equal to or greater than 1), and an information generation unit that processes control data including the first performance information using a generative model trained by machine learning to generate second performance information representing a second musical piece obtained by arranging the first musical piece into M parts (M is a natural number equal to or greater than 1 and different from N). [Explanation of symbols]
[0122] 100...information processing system, 200...machine learning system, 11,31...control device, 12,32...storage device, 13,33...communication device, 14...operation device, 15...sound source device, 16...sound emission device, 21...pre-processing unit, 22...information acquisition unit, 23...information generation unit, 24...post-processing unit, 25...musical score generation unit, 41...training data generation unit, 42...training processing unit.
Claims
[Claim 1] A first performance information representing a first piece of music composed of N parts (N is a natural number equal to or greater than 1) is obtained; The control data including the first performance information is processed by a generative model trained by machine learning, thereby generating second performance information representing a second musical piece obtained by arranging the first musical piece into M parts (M is a natural number equal to or greater than 1 and different from N). An information processing method implemented by a computer system.
Citation Information
Patent Citations
Simple musical score creating device and simple musical score creating program
JP2007241026A
Cited By
Melamine-formaldehyde foam with improved weather resistance
US12435175B2