Information processing method, information processing device, and program
The method uses a trained model to generate output tracks by modifying input tracks, ensuring high consistency and enhancing creativity in music or text generation.
Patent Information
- Application Number
- JP2022517612
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-01
- Filing Date
- 2021-04-13
- Publication Date
- 2025-10-22
- Estimated Expiration
- 2041-04-13
Smart Images

Figure 0007757958000001 
Figure 0007757958000002 
Figure 0007757958000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing method, an information processing device, and a program. [Background technology]
[0002] For example, Patent Document 1 discloses a technique for generating sequence information for automatically generating a music program or the like. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2002-207719 Summary of the Invention [Problem to be solved by the invention]
[0004] It is also possible to automatically generate music itself. For example, a track using a certain instrument can be used as an input track, and another track can be generated from it. In this case, it is desirable that the generated track has high consistency with the input track so that it harmonizes with the input track. The same can be said for generating various information other than music (for example, generating translated text).
[0005] An object of one aspect of the present disclosure is to provide an information processing method, an information processing device, and a program that are capable of generating a track that has improved consistency with an input track. [Means for solving the problem]
[0006] An information processing method according to one aspect of the present disclosure is an information processing method that generates an output track using an input track including a plurality of first information elements provided over a certain period or section, and a trained model, wherein the output track includes a first track that is the same as the input track or a track to which a modification has been made, and a second track that includes a plurality of second information elements provided over a certain period or section, and the trained model is a trained model generated using training data so as to output the output track when the first track is input.
[0007] An information processing device according to one aspect of the present disclosure includes a generation unit that generates an output track using an input track including a plurality of first information elements provided over a certain period or a certain section, and a trained model, wherein the output track includes a first track that is the same as the input track or a track in which a portion of the input track has been modified, and a second track that includes a plurality of second information elements provided over a certain period or a certain section, and the trained model is a trained model generated using training data so that when input data corresponding to the first track is input, output data corresponding to the output track is output.
[0008] A program according to one aspect of the present disclosure is a program for causing a computer to function, and causes the computer to generate an output track using an input track including a plurality of first information elements provided over a certain period or section, and a learned model, wherein the output track includes a first track that is the same as the input track or a track in which a portion of the input track has been modified, and a second track that includes a plurality of second information elements provided over a certain period or section, and the learned model is a learned model generated using training data so that when input data corresponding to the first track is input, output data corresponding to the output track is output. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a diagram illustrating an example of the appearance of an information processing apparatus according to an embodiment. [Figure 2] FIG. 10 is a diagram illustrating an example of an input screen of an information processing device. [Figure 3] FIG. 10 is a diagram illustrating an example of an output screen of an information processing device. [Figure 4] FIG. 2 is a diagram illustrating an example of functional blocks of an information processing device. [Figure 5] FIG. 10 is a diagram showing an example of a first track. [Figure 6] FIG. 10 is a diagram showing an example of a first track. [Figure 7] FIG. 10 is a diagram showing an example of a first track. [Figure 8] FIG. 10 is a diagram illustrating an example of a correspondence relationship between an input token and a token sequence. [Figure 9] FIG. 10 is a diagram illustrating an example of an additional token. [Figure 10] FIG. 1 is a diagram illustrating an example of functional blocks of a trained model. [Figure 11] FIG. 1 is a diagram illustrating an example of an overview of token sequence generation using a trained model. [Figure 12] FIG. 10 is a diagram illustrating an example of an output track. [Figure 13] FIG. 10 is a diagram illustrating an example of an output track. [Figure 14] FIG. 10 is a diagram illustrating an example of an output track. [Figure 15] 1 is a flowchart showing an example of a process (information processing method) executed in an information processing device. [Figure 16] 1 is a flowchart illustrating an example of generating a trained model. [Figure 17] FIG. 1 illustrates an example of a hardware configuration of an information processing device. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.
[0011] The present disclosure will be described in the following order: 1. Embodiment 1.1 Example of information processing device configuration 1.2 Examples of processing (information processing methods) performed by information processing devices 1.3 Example of generating a trained model 1.4 Hardware Configuration Examples 2. Variations 3. Effects
[0012] 1. Embodiment 1.1 Example of the outline configuration of an information processing device The following description will mainly take as an example an information processing device that can be used in the information processing method according to the embodiment. The information processing device according to the embodiment is used, for example, as an information generating device that generates various types of information. Examples of the generated information include music pieces, sentences, etc. The information to be handled is referred to as a "track." A track includes multiple information elements that are provided over a certain period or a certain section. When the track is a piece of music, an example of the information element is sound information of an instrument. Examples of sound information include the pitch value of a sound, the period during which a sound is generated, etc. In this case, the track may indicate the sound of an instrument at each time during the certain period. When the track is a sentence, an example of the information element is a word, a morpheme, etc. (hereinafter simply referred to as "words, etc."). In this case, the track may indicate words, etc. at each position within the certain section. Unless otherwise specified, the following description will be given assuming that the track is a piece of music and the information element is sound information.
[0013] FIG. 1 is a diagram showing an example of the appearance of an information processing device according to an embodiment. The information processing device 1 is realized by, for example, executing a predetermined program (software) on a general-purpose computer. In the example shown in FIG. 1, the information processing device 1 is a laptop used by a user U. The display screen of the information processing device 1 is illustrated as a display screen 1a. In addition to a laptop, the information processing device 1 can also be realized by various other devices such as a PC or a smartphone.
[0014] FIG. 2 is a diagram showing an example of an input screen of an information processing device. In the item "Input Track Selection," the user U selects an input track. The input track is sound information (pitch value of the sound, duration of sound occurrence, etc.) of an instrument (first instrument) at each time during a certain period. Any instrument including bass, drums, etc. can be the first instrument. The user U selects an input track, for example, by specifying data (MIDI file, etc.) corresponding to the input track. Visualized information of the selected input track is displayed below the item "Input Track Selection."
[0015] In the item "Make changes to the input track", the user U selects whether or not to make changes to the input track, and if so, also selects the degree of change (amount of change). An example of the amount of change is the percentage (%) of sound information to be changed. The amount of change may be selected from a plurality of numerical values prepared in advance, or may be directly input by the user U1. The specific content of the change will be described later.
[0016] In the "Instrument Selection" section, the user U selects an instrument (second instrument) to be used in the newly generated track. The second instrument may be selected automatically or may be specified by the user U. As with the first instrument described above, any instrument can be the first instrument. The type of the second instrument may be the same as the type of the first instrument.
[0017] FIG. 3 is a diagram showing an example of an output screen. In this example, information visualizing the output track is displayed. The output track is a track set (multi-track) including multiple tracks, and in the example shown in FIG. 3, two tracks are included. The first track is shown at the bottom of the figure and is the same as the input track (FIG. 2) or a track with a partially modified input track. The second track is shown at the top of the figure and is a newly generated track that shows sound information of a second instrument at each time during a certain period. In this example, the first track and the second track are displayed in a selectable and playable manner. The two tracks may be played simultaneously. A track that is more consistent with the first track is generated as the second track according to the principles described below. Such a first track and second track constitute a harmony with each other and are suitable for simultaneous playback.
[0018] 1 to 3 described above are merely examples of the external appearance and input / output screen configuration of the information processing device 1, and various other configurations may be adopted.
[0019] 4 is a diagram showing an example of functional blocks of the information processing device 1. The information processing device 1 includes an input unit 10, a storage unit 20, a generation unit 30, and an output unit 40.
[0020] An input track is input to the input unit 10. For example, the input unit 10 receives an input track selected by a user U as described above with reference to Fig. 2. The input unit 10 may also receive a selection of whether or not to make changes to the input track, and may also receive a selection of instruments to be used in the generated track.
[0021] The storage unit 20 stores various information used in the information processing device 1. Of these, FIG. 4 illustrates a trained model 21 and a program 22. The trained model 21 is a trained model generated using training data so that, when input data corresponding to a first track is input, output data corresponding to an output track is output. Details of the trained model 21 will be explained again later. The program 22 is a program (software) for realizing processing executed in the information processing device 1.
[0022] The generation unit 30 generates an output track using the input track input to the input unit 10 and the trained model 21. In Figure 4, functional blocks that perform representative processing by the generation unit 30 are exemplified as a track modification unit 31, a token generation unit 32, and a track generation unit 33.
[0023] The track modification unit 31 modifies the input track. The modified input track is a form of the first track. For example, the track modification unit 31 modifies part of the multiple pieces of sound information (such as the pitch value and duration of the sound of the first instrument) included in the input track. This will be described with reference to FIGS. 5 to 7.
[0024] 5 to 7 are diagrams showing examples of the first track. The horizontal axis indicates time (in this example, time (bars)), and the vertical axis indicates pitch value (in this example, MIDI pitch). Note that bars indicate bar numbers, which will be treated as unit time below.
[0025] The first track illustrated in Fig. 5 is the same as the input track. That is, this track is the input track input to the input unit 10. This track, which is not changed by the track change unit 31, is also one aspect of the first track. For comparison with Fig. 7 described later, the two sounds in Fig. 5 are labeled as sound P1 and sound P2.
[0026] The first track illustrated in FIG. 6 differs from the input track (FIG. 5) in that it includes sounds P11 to P13. Sounds P11 to P13 are sounds obtained by modifying the corresponding sounds in the input track. Sounds P11 and P13 have been modified to be higher in pitch (higher in pitch value). The degree of modification may vary between the sounds. Sound P12 has been modified to be lower in pitch (lower in pitch value). As another modification, the corresponding sounds in the input track may be deleted (masked so that information is lost) and sounds P11 to P13 added. The corresponding sounds in the input track may be replaced with sounds P11 to P13.
[0027] The first track illustrated in FIG. 7 differs from the input track illustrated in FIG. 5 in that it includes sounds P21 to P29. Sounds P21 to P28 are sounds obtained by modifying the corresponding sounds in the input track. Sound P29 is a newly added sound. Sounds P21, P25, and P26 have been modified to lower their pitches. The degree of modification may vary between sounds. Sound P22 has been modified to raise its pitch. Sounds P23 and P24 are sounds obtained by dividing sound P1 in the input track into sound P23, which has been modified to raise its pitch, and sound P24, which has been modified to lower its pitch. Sounds P27 and P28 have been modified to lower their pitches and lengthen their durations. As another modification, the corresponding sounds in the input track may be deleted (masked) and sounds P21 to P28 may be added. The proportion of the notes P21 to P29 in the whole tone in FIG. 7 is greater than the proportion of the notes P11 to P13 in the whole tone in FIG. 6 described above.
[0028] The track change unit 31 obtains a track that is partially different from the input track as the first track without being constrained by the input track input to the input unit 10. The degree of constraint (constraint strength) can be adjusted depending on the proportion of the sound to be changed. The amount of adjustment is determined randomly, for example.
[0029] Returning to FIG. 4, the token generator 32 generates a token string based on the first track. In one embodiment, the token generator 32 generates the token string by arranging a first token and a second token in chronological order. The first tokens are tokens that indicate the onset and stop of each sound included in the first track. The second tokens are tokens that indicate the period during which the state indicated by the corresponding first token is maintained. An example of generating a token string will be described with reference to FIG. 8.
[0030] 8 is a diagram showing an example of the correspondence between input tokens and token strings. The token string shown in the lower part of the figure is generated from the input token shown in the upper part of the figure. In the token string, the part represented by angle brackets <> corresponds to one token.
[0031] token<ON, M, 60> is the token (first token) that indicates that a note with a pitch value of 60 from instrument M begins to be generated at time 0.<SHIFT, 1> is a token (corresponding second token) indicating that the state (instrument M, pitch value 60) indicated by the corresponding first token will be maintained for one unit of time. In other words, SHIFT means that only time moves (only time passes) while the state indicated by the immediately preceding token remains the same.
[0032] token<ON, M, 64> is the token (first token) that indicates the start of a note with a pitch value of 64 on instrument M.<SHIFT, 1> is a token (corresponding second token) indicating that the state (instrument M, pitch value 60, instrument M, pitch value 64) indicated in the corresponding first token is maintained for one unit time period.
[0033] token<ON, M, 67> is the token (first token) that indicates the start of a note at pitch 67 on instrument M.<SHIFT, 2> is a token (corresponding second token) indicating that the state indicated in the corresponding first token (instrument M, pitch value 60, instrument M, pitch value 64, instrument M, pitch value 67) will be maintained for a period of two unit time.
[0034] token<OFF, M, 60> is a token (first token) that indicates the end of the sound produced by instrument M at pitch value 60.<OFF, M, 64> is the token (first token) that indicates the end of the sound produced by instrument M at pitch value 64.<OFF, M, 67> is the token (first token) that indicates the end of the sound produced by instrument M at pitch value 67.<SHIFT, 1> is a token (corresponding second token) indicating that the state indicated by the corresponding first token (no sound is produced by any instrument) is maintained for one unit time period.
[0035] token<ON, M, 65> is the token (first token) that indicates the start of a note with a pitch value of 65 on instrument M.<SHIFT, 1> is a token (corresponding second token) indicating that the state (instrument M, pitch value 65) indicated in the corresponding first token is maintained for one unit time period.
[0036] token<OFF, M, 65> indicates the end of the note occurrence at pitch value 65 on instrument M (first token).
[0037] In the above example, when multiple sounds are present at the same time, the tokens are arranged in order from the lowest to the highest sound. By determining the order in this way, it becomes easier to train the trained model 21.
[0038] The token generator 32 may add (embed) further tokens to the token string generated as described above (the token string shown at the bottom in FIG. 8) as a basic token string. As examples of the additional tokens, a first additional token and a second additional token will be described.
[0039] The first additional token indicates the period that has elapsed up to the time when each token appears in the token string. The token generator 32 may include (embed) in each token a token that indicates the total period indicated in the second token up to the time when each token appears in the token string. As explained above, the SHIFT in the second token means that only the time moves while maintaining the state indicated in the immediately preceding token, so the embedding of the first additional token can also be called TSE (Time Shift Summarization Embedding).
[0040] The second additional token is a token that indicates the position of each token in the token sequence. The token generation unit 32 may include (embed) a token that indicates the position of each token in the token sequence in each token. Embedding the second additional token can also be called PE (Position Embedding).
[0041] An example of embedding the above-mentioned additional tokens (first additional token and second additional token) will be described with reference to FIG.
[0042] 9 is a diagram showing an example of an additional token. In this example, the token<ON, b, 24> ,token<SHIFT, 6> and tokens<OFF, b, 24> These indicate that the generation of a sound from instrument b with a pitch value of 24 begins at time 0, and the generation of the sound is maintained for a period of 6 time units before ceasing.
[0043] As the first additional token corresponding to each of the above basic tokens,<SUM, 0> ,token<SUM, 6> and tokens<SUM, 6> For example, token<SUM, 0> is the token<ON, b, 24> indicates that the time elapsed until the token appears is 0.<SUM, 6> is the token<SHIFT, 6> and tokens<OFF, b, 24> indicates that 6 units of time have elapsed up to the time when appears.
[0044] As a second additional token for each of the above basic tokens,<POS, 0> ,token<POS, 1> and tokens<POS, 2> For example, token<POS, 0> is the token<ON, b, 24> indicates that the token is at the 0th position in the token sequence.<POS, 1> is the token<SHIFT, 6> indicates that is in the first position in the token sequence.<POS, 2> is the token<OFF, b, 24> is in the second position in the token sequence.
[0045] As described above, by including additional tokens in addition to basic tokens, a lot of information is added to the token sequence. In particular, by embedding the first additional token (TSE), actual time information corresponding to the basic token can be included in the token sequence. This allows time-related learning to be bypassed in generating the trained model 21, reducing the processing load associated with learning.
[0046] Returning to Fig. 4, the track generation unit 33 generates an output track. Specifically, the track generation unit 33 generates an output track using the input track and the trained model 21. An example of generating an output track using the trained model 21 will be described with reference to Fig. 10.
[0047] 10 is a diagram showing an example of functional blocks of a trained model. In this example, the trained model 21 includes an encoder 21a and a decoder 21b. An example of the trained model 21 having such a configuration is Seq2Seq (Sequence to Sequence), and an RNN (Recurrent Neural Network) or a Transformer can be used as the architecture.
[0048] The encoder 21a extracts features from an input token sequence. The decoder 21b generates (reconstructs) an output token sequence from the features extracted by the encoder 21a, for example, using the most probable token sequence. The encoder 21a may be trained by unsupervised learning such as a variational autoencoder (VAE) or a generative adversarial network (GAN). The input token sequence to the encoder 21a is compared with the output token sequence generated by the decoder 21b, and parameters of the encoder 21a and the decoder 21b are adjusted. A trained model 21 in which the parameters of the encoder 21a and the decoder 21b are optimized by repeating the adjustment is generated. An example of a flow for generating the trained model 21 will be described again later with reference to FIG. 16.
[0049] FIG. 11 is a diagram showing an example of an overview of token sequence generation by a trained model. In the diagram, the token sequence shown below the encoder 21a is the input token sequence input to the encoder 21a and corresponds to the first track (the input track or a track to which a change has been made). The token sequence shown above the decoder 21b is the output token sequence generated (reconstructed) by the trained model 21 and corresponds to the output track. As shown in the diagram, the output token sequence includes tokens related to instrument m in addition to tokens related to instrument b included in the input token sequence. That is, a token sequence corresponding to a track set including not only the first track using instrument b (the first instrument) but also a new track using instrument m (corresponding to the second instrument) is generated as the output token sequence.
[0050] By generating a token sequence corresponding to a track set of a first track and a second track as described above, a token sequence of the second track is generated that is more consistent with the first track set (i.e., with the input track) than, for example, a token sequence corresponding only to the second track is generated. Such music generation that takes consistency with the first track set into consideration is highly compatible with the human music generation process and is likely to produce a synergistic effect of creativity. Examples of a highly compatible human music generation process include creating tracks one by one or creating music inspired by a certain track.
[0051] In one embodiment, the decoder 21b of the trained model 21 may generate each token in chronological order. In this case, in the process of generating a token sequence, the decoder 21b may generate the next token by referring to previously generated tokens (attention function).
[0052] For example, as shown below the decoder 21b in the figure, the start token <start>, then as the base token, the token<ON, b, 24> ,token<ON, m, 60> ,token<SHIFT, 4> ,token<OFF, m, 60> and tokens<SHIFT, 2> are generated in this order. At that time, the decoder 21b also generates the additional token described above (it does not have to be output). In particular, by generating the first additional token, the decoder 21b can generate the next token while also referring to the token at the corresponding time in the input token sequence. As a result, the output track has a higher consistency between the new track using instrument m and the first track using instrument b.
[0053] For example, the track generation unit 33 generates an output track by using the token sequence generated as described above by the trained model 21. Some examples of the output track will be described with reference to FIGS. 12 to 14.
[0054] 12 to 14 are diagrams showing examples of output tracks. The track shown at the bottom of the diagrams is the first track (input track or modified track) shown in FIGS. 5 to 7 described above, and shows the sound of a first instrument. The track shown at the top of the diagrams is a second track newly generated based on the first track, and shows the sound of a second instrument. As can be seen from these diagrams, different output tracks are obtained when the input track is used as the first track as is (FIG. 12) and when modifications are made (FIGS. 13 and 14). In either case, a track set of the first track and the second track is generated as the output track as described above, thereby obtaining a second track that is more consistent with the first track.
[0055] 4, the output unit 40 outputs the track generated by the generation unit 30. For example, the output track is displayed as described above with reference to FIG.
[0056] 1.2 Examples of processing (information processing methods) performed by information processing devices FIG. 15 is a flowchart showing an example of a process (information processing method) executed in the information processing device.
[0057] In step S1, an input track is input. For example, a user U1 selects an input track as described above with reference to FIG. 2. The input unit 10 accepts the input track. A selection of whether or not to make changes to the input track may be input, and further, a selection of an instrument to be used in the track to be generated (second track) may also be input.
[0058] In step S2, it is determined whether or not to make changes. This determination is made based on, for example, the input result of the previous step S1 (such as the selection of whether or not to make changes to the input track). If changes are to be made (step S2: Yes), the process proceeds to step S3. If not (step S2: No), the process proceeds to step S4.
[0059] In step S3, the input track is modified. For example, the track modification unit 31 modifies the input track input in the previous step S1. The specific content of the modification has been described above with reference to Figures 6 and 7, and therefore will not be described again here.
[0060] In step S4, an input token sequence is generated. For example, the token generation unit 32 generates an input token sequence corresponding to the input track input in the previous step S1 and / or the input track (first track) to which a change was made in the previous step S3. The specific contents of the generation have been described above with reference to Figures 8 and 9, and therefore will not be described again here.
[0061] In step S5, an output token sequence is obtained using the trained model. For example, the track generation unit 33 inputs the input token sequence generated in the previous step S4 into the trained model 21, thereby obtaining an output token sequence corresponding to the output track. The specific details of the obtaining process have been described above with reference to FIG. 11 etc., and therefore will not be repeated here.
[0062] In step S6, an output track is generated. For example, the track generation unit 33 generates an output track corresponding to the output token sequence acquired in the previous step S4.
[0063] In step S7, the output track is output. For example, the output section 40 outputs the output track generated in the previous step S6, as described above with reference to FIG.
[0064] After the process of step S7 is completed, the process of the flowchart ends. For example, by such a process, an output track is generated from an input track and output.
[0065] 1.3 Example of generating a trained model 16 is a flowchart illustrating an example of generating a trained model. In this example, training is performed using a set of mini-batch samples.
[0066] In step S11, a set of mini-batch samples for a track set (corresponding to an output track) is prepared. Each mini-batch sample is composed, for example, of a combination of parts of mini-data for a plurality of songs prepared in advance. A set of mini-batch samples is obtained by collecting a plurality of such mini-batch samples (for example, 256 samples). One mini-batch sample from the set of mini-batch samples is used in one flow. Another mini-batch sample is used in another flow.
[0067] In step S12, the tracks are modified. The modifications have been described above and will not be described again. The number of track sets can be increased by the amount of the modifications.
[0068] In step S13, forward calculation is performed. Specifically, a token sequence corresponding to some tracks (corresponding to the first track) of the track set prepared in the previous step S12 is input to a neural network including an encoder and a decoder, and a token sequence corresponding to a new track set (corresponding to the first track set and the second track set) is output. An error function is calculated from the output track set and the previously prepared track set.
[0069] In step S14, backward calculation is performed. Specifically, a cross-entropy error is calculated from the error function obtained in step S13. From the calculated cross-entropy error, the parameter error of the neural network and the error gradient are obtained.
[0070] In step S15, the parameters are updated. Specifically, the parameters of the neural network are updated according to the error obtained in the previous step S14.
[0071] After the process of step S15 is completed, the process returns to step S11, where a different mini-batch sample from the previously used mini-batch sample is used.
[0072] For example, a trained model can be generated as described above. The above is merely an example, and various known training methods may be used in addition to the above-described method using a mini-batch sample set.
[0073] 1.4 Hardware Configuration Examples 17 is a diagram showing an example of the hardware configuration of an information processing device. In this example, the information processing device 1 is realized by a computer 1000. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, an HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.
[0074] The CPU 1100 operates and controls each unit based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.
[0075] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .
[0076] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records an information processing program according to the present disclosure, which is an example of program data 1450.
[0077] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.
[0078] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disk), magneto-optical recording media such as an MO (Magneto-Optical disk), tape media, magnetic recording media, and semiconductor memories.
[0079] For example, when the computer 1000 functions as the information processing device 1, the CPU 1100 of the computer 1000 executes an information processing program loaded onto the RAM 1200 to realize functions of the generation unit 30, etc. Also, the HDD 1400 stores a program according to the present disclosure (program 22 in the storage unit 20) and data in the storage unit 20. Note that the CPU 1100 reads and executes program data 1450 from the HDD 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.
[0080] 2. Variations Although one embodiment of the present disclosure has been described above, the present disclosure is not limited to the above embodiment.
[0081] In the above embodiment, an example has been described in which the input track includes one track and the output track includes two tracks. However, the input track may include two or more tracks. The output track may include three or more tracks. As the number of tracks increases, the input / output mode of the information processing device 1 (such as the input screen in FIG. 2) is also changed appropriately.
[0082] In the above embodiment, an example has been described in which the trained model is a model including an encoder and a decoder, such as an RNN or Seq2Seq. However, the trained model is not limited to these models, and various trained models capable of reconstructing a token sequence from an input token sequence may be used.
[0083] In the above embodiment, the track is a piece of music and the information elements are sound information. However, various tracks including information elements other than sound information may be used. For example, the track may be a sentence and the information elements may be words, etc. In this case, the multiple first information elements are words, etc. in a first language given over a certain interval, and the input track indicates the words, etc. in the first language at each position within the certain interval. The multiple second information elements are words, etc. in a second language given over a certain interval, and the second track indicates the words, etc. in the second language at each position within the certain interval. Regarding tokens, the first token indicates the occurrence and stop of each word, etc. The second token indicates the interval (e.g., the length of the word, etc.) during which the state indicated by the corresponding first token is maintained. The token generation unit 32 generates a token sequence by arranging the first token and the second token in order of their positions within the certain interval. The first additional token indicates the interval that has elapsed up to the time when each token appears in the token sequence. The second token indicates the position of each token in the token sequence.
[0084] Some of the functions of the information processing device 1 may be realized outside the information processing device 1 (for example, an external server). In this case, the information processing device 1 may include some or all of the functions of the storage unit 20 and the generation unit 30 in the external server. The information processing device 1 communicates with the external server, thereby similarly realizing the processing of the information processing device 1 described above.
[0085] 3. Effects The information processing method described above can be specified, for example, as follows. As described with reference to FIG. 5 and FIGS. 10 to 15, the information processing method generates an output track using an input track and a trained model 21 (step S6). As described with reference to FIG. 5, the input track includes a plurality of first information elements provided over a certain period or a certain section. The output track includes a first track (the same as the input track or a track with modifications made thereto) and a plurality of second information elements provided over a certain period or a certain section. For example, the plurality of first information elements are sound information of a first instrument provided over a certain period, and the input track indicates sound information of the first instrument at each time during the certain period. The plurality of second information elements are sound information of a second instrument provided over a certain period, and the second track indicates sound information of the second instrument at each time during the certain period. The input track indicates words in a first language at each position during the certain section. The second track included in the output track indicates words in a second language at each position during the certain section. The trained model 21 is a trained model generated using training data so that when input data corresponding to the first track is input, output data corresponding to the output track is output.
[0086] According to the information processing method, a track set of a first track and a second track is generated as output tracks. This makes it possible to generate second tracks that are more consistent with the first track set (i.e., with the input tracks) than when only the second track is generated as an output set. Such music generation that takes into account consistency with the first track set is highly compatible with the human music generation process and is likely to produce a synergistic effect of creativity.
[0087] 6 and 7, if the first track is a track in which a part of the input track has been changed, the information processing method may generate the first track by changing a part of a plurality of first information elements (for example, the sound of a first instrument) included in the input track (step S3). This makes it possible to obtain an output track that is different from the case in which the input track is used as the first track without being restricted by the input track.
[0088] As described with reference to FIGS. 8 to 11, the input data may be an input token sequence corresponding to a first track, and the output data may be an output token sequence corresponding to an output track. The information processing method may obtain an output token sequence by inputting the input token sequence into the trained model 21 (step S5). The information processing method may generate an input token sequence by arranging first tokens and second tokens in time order within a certain period or in position order within a certain section (step S4). The first tokens indicate the occurrence and stop of each of a plurality of first information elements (e.g., the sound of a first instrument). The second tokens indicate a period or section during which the state indicated by the corresponding first token is maintained. For example, such a token sequence can be generated and used in a trained model.
[0089] As described with reference to Figures 9 to 11, the information processing method may generate an input token sequence by including, in the first token and the second token, an additional token indicating the time or position at which each of the first token and the second token appeared in the input token sequence. The additional token may be a token indicating the total duration or section indicated in the second token up to the time at which each of the first token and the second token appeared in the input token sequence. This allows time or position information to be included in the token sequence, thereby bypassing learning related to time or position in generating the trained model 21, for example, thereby reducing the processing load associated with learning.
[0090] 5 to 7, the sound information of the first instrument may include the pitch value of the sound of the first instrument and / or the duration of the sound. For example, by changing such sound information of the first instrument (step S3), the first track can be obtained.
[0091] The information processing device 1 described with reference to Figures 1 to 4 etc. is also one aspect of the present disclosure. That is, the information processing device 1 includes a generation unit 30 that generates an output track using the above-mentioned input track and a trained model 21. As described above, the information processing device 1 can also generate a second track that has improved consistency with the input track.
[0092] The program 22 described with reference to Figures 4, 17, etc. is also one aspect of the present disclosure. That is, the program 22 is a program for causing a computer to function, and causes the computer to generate an output track using the above-mentioned input track and the trained model 21. As described above, the program 22 can also generate a second track that has improved consistency with the input track.
[0093] The effects described in this disclosure are merely examples and are not limited to the disclosed contents. Other effects may also be obtained.
[0094] Although the embodiments of the present disclosure have been described above, the technical scope of the present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure. Furthermore, components of different embodiments and modifications may be combined as appropriate.
[0095] Furthermore, the effects of each embodiment described in this specification are merely examples and are not intended to be limiting, and other effects may also be obtained.
[0096] The present technology can also be configured as follows. (1) An information processing method for generating an output track using an input track including a plurality of first information elements given over a certain period or a certain section and a trained model, comprising: the output tracks include a first track that is the same as the input track or a track to which a change has been made, and a second track that includes a plurality of second information elements provided over the fixed period or fixed section; The trained model is a trained model generated using training data so as to output output data corresponding to the output track when input data corresponding to the first track is input. Information processing methods. (2) the first track is a track obtained by changing a part of the input track, the information processing method includes generating the first track by changing some of the plurality of first information elements included in the input track. The information processing method described in (1). (3) the input data is an input token sequence corresponding to the first track; the output data is an output token sequence corresponding to the output track; the information processing method includes inputting the input token sequence to the trained model to obtain the output token sequence. The information processing method according to (1) or (2). (4) generating the input token sequence by arranging first tokens indicating occurrence and stop of each of the plurality of first information elements and second tokens indicating a period or interval during which a state indicated by the corresponding first token is maintained in order of time within the certain period or in order of position within the certain interval; (3) The information processing method described in (3). (5) generating the input token sequence by including additional tokens in the first token and the second token that indicate a time or position at which each of the first token and the second token appeared in the input token sequence; (4) The information processing method described in (4). (6) the additional token is a token indicating the sum of the period or interval indicated by the second token up to the time when each of the first token and the second token appeared in the input token string; (5) The information processing method described in (5). (7) the plurality of first information elements are sound information of a first instrument given over the certain period of time, and the input track indicates the sound information of the first instrument at each time during the certain period of time; the plurality of second information elements are sound information of a second instrument given over the certain period of time, and the second track indicates the sound information of the second instrument at each time during the certain period of time; The information processing method according to any one of (1) to (6). (8) the sound information of the first musical instrument includes at least one of a pitch value of the sound of the first musical instrument and a duration of the sound; (7) The information processing method described in (7). (9) a generation unit that generates an output track using an input track including a plurality of first information elements provided over a certain period or a certain section and a trained model; the output track includes a first track that is the same as the input track or a track obtained by changing a part of the input track, and a second track that includes a plurality of second information elements provided over the fixed period or fixed section; The trained model is a trained model generated using training data so as to output output data corresponding to the output track when input data corresponding to the first track is input. Information processing device. (10) A program for causing a computer to function, generating an output track using an input track including a plurality of first information elements provided over a certain period or a certain section and a trained model; causing the computer to execute the output track includes a first track that is the same as the input track or a track obtained by changing a part of the input track, and a second track that includes a plurality of second information elements provided over the fixed period or fixed section; The trained model is a trained model generated using training data so as to output output data corresponding to the output track when input data corresponding to the first track is input. program. [Explanation of symbols]
[0097] 1. Information processing equipment 1a Display screen 10 Input section 20 Memory section 21 Pre-trained models 21a Encoder 21b decoder 22 Programs 30 Generation part 31 Track Change Section 32 Token generation unit 33 Track Generation Section 40 Output section< / start>
Claims
1. An information processing method for generating an output track using an input track including a plurality of first information elements provided over a certain period or a certain section, and a trained model, comprising: the output tracks include a first track that is the same as the input track or a track to which a change has been made, and a second track that is neither the same as the input track nor a track to which a change has been made, and that is newly generated so as to include a plurality of second information elements provided over the fixed period or fixed section; each of the first information element and the second information element includes at least one of phonetic information, words, and morphemes; the trained model is a trained model generated using training data so as to generate and output an output token sequence corresponding to the output track when an input token sequence corresponding to the first track is input; the output token sequence is a token sequence corresponding to a track set of the first track and the second track; the output track is generated by inputting the input token sequence into the trained model to obtain the output token sequence. Information processing methods.
2. the first track is a track obtained by changing a part of the input track, the information processing method includes generating the first track by changing a part of the plurality of first information elements included in the input track. The information processing method according to claim 1 .
3. generating the input token sequence by arranging first tokens indicating occurrence and stop of each of the plurality of first information elements and second tokens indicating a period or interval during which a state indicated by the corresponding first token is maintained in order of time within the certain period or in order of position within the certain interval; The information processing method according to claim 1 .
4. generating the input token sequence by including additional tokens in the first token and the second token that indicate a time or position at which each of the first token and the second token appeared in the input token sequence; The information processing method according to claim 3 .
5. the additional token is a token indicating the total of the period or interval indicated by the second token up to the time when each of the first token and the second token appeared in the input token string; The information processing method according to claim 4.
6. the plurality of first information elements are sound information of a first instrument given over the certain period of time, and the input track indicates the sound information of the first instrument at each time during the certain period of time; the plurality of second information elements are sound information of a second instrument given over the certain period of time, and the second track indicates the sound information of the second instrument at each time during the certain period of time; The information processing method according to claim 1 .
7. the sound information of the first musical instrument includes at least one of a pitch value of the sound of the first musical instrument and a duration of the sound; The information processing method according to claim 6.
8. a generation unit that generates an output track using an input track including a plurality of first information elements provided over a certain period or a certain section and a trained model; the output tracks include a first track that is the same as the input track or a track to which a change has been made, and a second track that is neither the same as the input track nor a track to which a change has been made, and that is newly generated so as to include a plurality of second information elements provided over the fixed period or fixed section; the trained model is a trained model generated using training data so as to generate and output an output token sequence corresponding to the output track when an input token sequence corresponding to the first track is input; the output token sequence is a token sequence corresponding to a track set of the first track and the second track; the output track is generated by inputting the input token sequence into the trained model to obtain the output token sequence. Information processing device.
9. A program for causing a computer to function, generating an output track using an input track including a plurality of first information elements provided over a certain period or a certain section and a trained model; causing the computer to execute the output tracks include a first track that is the same as the input track or a track to which a change has been made, and a second track that is neither the same as the input track nor a track to which a change has been made, and that is newly generated so as to include a plurality of second information elements provided over the fixed period or fixed section; each of the first information element and the second information element includes at least one of phonetic information, words, and morphemes; the trained model is a trained model generated using training data so as to generate and output an output token sequence corresponding to the output track when an input token sequence corresponding to the first track is input; the output token sequence is a token sequence corresponding to a track set of the first track and the second track; the output track is generated by inputting the input token sequence into the trained model to obtain the output token sequence. program.
Citation Information
Patent Citations
Method and device for generating sequence information
JP2002207719A
Sound source separation device and program
JP2012053205A
Electronic apparatus, information processing method, and program
JP2019159146A
Generating using a bidirectional RNN variations to music
US20160314392A1