Tone output program, tone output device, and tone output method

The tone output program and device analyze music data to determine timbre types using a machine-learned model, enhancing user understanding of musical affinity by outputting association degrees and identification information.

WO2025203738A1PCT designated stage Publication Date: 2025-10-02ROLAND CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/030650
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2024-08-28
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing technologies fail to determine the type of timbre that corresponds to inputted musical melody information, limiting user understanding of musical affinity.

Method used

A tone output program and device that utilize a machine-learned tone output model to analyze the degree of association between music data and timbre types by superimposing spaces based on music data and tone waveform data, outputting the correspondence with identification information.

Benefits of technology

Enables users to intuitively understand the degree of association between music data and timbre types, facilitating the selection of suitable timbre types for input music.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024030650_02102025_PF_FP_ABST
    Figure JP2024030650_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a tone output program, a tone output device, and a tone output method by which a user can grasp the correspondence relation between music data and the kind of tone. SMF data is input to a PC (1), and the SMF data is transmitted to a server (50). In the server (50), the received SMF data is input to a tone output model (M), a tone distance indicating the correspondence relation between the SMF data and the kind of tone is calculated, and a tone distance score based on the reciprocal of the tone distance is transmitted to the PC (1) together with tone type information. A tone score obtained by converting the tone distance score received from the server (50) in the PC (1) is calculated and displayed on a display device (4). A user (H) can intuitively grasp the degree (correspondence relationship) of association of the kind of tone with the input SMF data by grasping the magnitude of the tone score.
Need to check novelty before this filing date? Find Prior Art

Description

Tone output program, tone output device, and tone output method

[0001] The present invention relates to a tone output program, a tone output device, and a tone output method.

[0002] Patent Document 1 describes a method in which the similarity between music melody information, which is composed of melodies and chords of music input by a user, and music information is calculated, and music information according to the similarity is displayed. Since the music information includes song titles and song data, the user can find out the song titles that are highly similar to the music melody information they input, and enjoy those songs.

[0003] Japanese Unexamined Patent Publication No. 8-123818

[0004] Recently, there has been an increasing need for users to be able to ascertain the type of timbre that has a high musical affinity and corresponds to inputted musical melody information. However, while the technology in Patent Document 1 allows users to ascertain the song title and tune based on inputted musical melody information, it has the problem that it is not possible to ascertain the type of timbre that corresponds to that musical melody information.

[0005] The present invention has been made to solve the above-mentioned problems, and aims to provide a tone output program, a tone output device, and a tone output method that allow a user to understand the correspondence between music data and tone types.

[0006] To achieve this object, the timbre output program of the present invention causes a computer to execute an output step of outputting the correspondence between input music data and timbre types together with identification information of the timbre types.

[0007] Another tone output program of the present invention causes a computer to execute an input step of inputting music data, an analysis step of analyzing the degree of association between the music data input in the input step and the type of tone using a machine-learned tone output model, and an output step of outputting the degree of association analyzed in the analysis step together with identification information for the type of tone, wherein the tone output model analyzes the degree of association by superimposing a space based on the music data and a space based on tone waveform data.

[0008] The tone color output device of the present invention comprises output means for outputting the correspondence between input music data and tone color types together with identification information for the tone color types.

[0009] Another tone output device of the present invention comprises an input means for inputting music data, an analysis means for analyzing the degree of association between the music data input by the input means and a tone type using a machine-learned tone output model, and an output means for outputting the degree of association analyzed by the analysis means together with identification information for the tone type, wherein the tone output model analyzes the degree of association by superimposing a space based on the music data and a space based on tone waveform data.

[0010] The tone color output method of the present invention includes an output step of outputting the correspondence between input music data and tone color types together with identification information of the tone color types.

[0011] Another tone output method of the present invention includes an input step of inputting music data, an analysis step of analyzing the degree of association between the music data input in the input step and a tone type using a machine-learned tone output model, and an output step of outputting the degree of association analyzed in the analysis step together with identification information for the tone type, wherein the tone output model analyzes the degree of association by superimposing a space based on the music data and a space based on tone waveform data.

[0012] 1 is a diagram showing the appearance of a PC; (a) is a diagram showing a timbre list screen, and (b) is a diagram showing the timbre list screen when the slide bar is moved to the right end; (a) is a block diagram showing the electrical configuration of a PC and a server, and (b) is a diagram showing a schematic diagram of a timbre score memory; (b) is a flowchart of a PC main process; (c) is a flowchart of a server main process; (d) is a flowchart of a display sound generation process; (e) is a diagram explaining a method of learning a timbre output model; (e) is a diagram explaining an error learning unit, and (b) is a diagram explaining distance learning using triplet loss; and (f) is a flowchart of a model creation process.

[0013] A preferred embodiment will now be described with reference to the accompanying drawings. First, an overview of the PC 1 of this embodiment will be described with reference to Figure 1. Figure 1 is a diagram showing the appearance of the PC 1. The PC 1 is an information processing device (computer, tone color output device) that displays the correspondence between input music data and tone colors of acoustic piano, electric guitar, etc.

[0014] In this embodiment, SMF (Standard MIDI File) data in MIDI (Musical Instrument Digital Interface) format is used as the music data input to the PC 1. Hereinafter, "MIDI format data" will be abbreviated as MIDI data. Note that the format used for music data is not limited to SMF or MIDI, and other formats may also be used.

[0015] The PC 1 is provided with a mouse 2 and keyboard 3 for inputting instructions from the user H, a display device 4 for displaying various information such as timbre type information, a MIDI keyboard 5 for inputting MIDI data based on performance operations by the user H, and a speaker 6 for outputting musical tones based on music data. Here, timbre type information is information for identifying the type of timbre, such as the timbre number, timbre category or name of the timbre type.

[0016] In the PC 1, MIDI data input from the MIDI keyboard 5 is converted into SMF (Standard MIDI File) data and used to display the correspondence with the types of timbre. In addition, pre-created SMF data input from another PC or the like is also used to display the correspondence with the types of timbre.

[0017] A server 50 is connected to the PC 1 via the Internet Nt. The server 50 is an information processing device that calculates the correspondence between the SMF data transmitted from the PC 1 and the timbre types, and transmits the calculated correspondence to the PC 1 together with the corresponding timbre type information. The server 50 is equipped with a timbre output model M (artificial intelligence), which will be described later. The SMF data transmitted from the PC 1 is input into the timbre output model M, which calculates information regarding the correspondence between the SMF data and the timbre types. The information regarding the calculated correspondence is transmitted to the PC 1 together with the corresponding timbre type information.

[0018] In this embodiment, the "correspondence between SMF data and timbre types" output from the timbre output model M is a timbre distance that represents the degree of association between the SMF data and the timbre type. The timbre distance is a Euclidean distance calculated according to the degree of association between the timbre type and the SMF data. The timbre distance is a numerical value greater than or equal to 0, and the more closely related the timbre type is to the SMF data, the smaller the timbre distance calculated from the timbre output model M, and the less closely related the timbre type is to the SMF data, the larger the timbre distance calculated from the timbre output model M.

[0019] Such timbre distances are calculated for all timbre types managed (recognized) by the timbre output model M. The reciprocals of the timbre distances for all or some of the timbre types output from the timbre output model M are calculated as timbre distance scores, and the calculated timbre distance scores are transmitted together with the corresponding timbre type information to the PC 1 to which the SMF data was transmitted.

[0020] In the PC 1, a display based on the timbre distance score received from the server 50 is performed on the display device 4. Specifically, a timbre score for display is calculated, which is a numerical value obtained by converting the timbre distance score received from the server 50 so that the maximum value is 100 and the minimum value is 0. In other words, the more closely related the timbre type is to the input SMF data and the shorter the timbre distance, the closer the timbre score will be to 100. On the other hand, the less closely related the timbre type is to the input SMF data and the longer the timbre distance, the closer the timbre score will be to 0.

[0021] A display based on the calculated timbre score is then displayed on the display device 4. By understanding the magnitude of the timbre score, the user H can intuitively understand the degree of association between the timbre type and the input SMF data.

[0022] Next, a description will be given of the display screens displayed on the display device 4. In this embodiment, a chat screen 4a and a timbre list screen 4b are provided as display screens for inputting SMF data and obtaining timbre scores related to the SMF data.

[0023] 1 displays a chat screen 4a. The chat screen 4a is used to input SMF files and display timbre type information for the timbre with the highest timbre score. The chat screen 4a displays a record button 40, a play button 41, a metronome button 42, a note display area 43, a panic button 44, a save button 45, a phrase history button Fb, and an optimal timbre button Tb.

[0024] The record button 40 is a button that instructs the start and end of real-time input of MIDI data from the MIDI keyboard 5. When the record button 40 is pressed by the user H, the MIDI data subsequently input from the MIDI keyboard 5 is acquired as the MIDI data used to acquire the timbre score. When the record button 40 is pressed again, the acquisition of the MIDI data is terminated.

[0025] The play button 41 is a button for instructing the output of musical tones of SMF data based on MIDI data from the MIDI keyboard 5 or SMF data input from another PC, etc. (hereinafter, SMF data specified by inputting in these ways will be collectively referred to as "specified SMF data.") When the play button 41 is pressed, the type of timbre with the highest acquired timbre score (i.e., the type of timbre displayed on the optimum timbre button Tb, described below), or the type of timbre specified on the timbre list screen 4b, described below, is applied to the specified SMF data and output as musical tones.

[0026] The metronome button 42 is a button that instructs the start or end of output of a specified rhythm sound. The rhythm sound that is output by pressing the metronome button 42 allows the user H to easily grasp the desired rhythm when inputting MIDI data from the MIDI keyboard 5.

[0027] The note display area 43 is a display area for illustrating the note lengths and scales included in the designated SMF data. By dragging and dropping SMF data input from another PC or the like into the note display area 43, the SMF data can be designated as the designated SMF data.

[0028] The panic button 44 is a button for turning off all notes of the musical tones of the specified SMF data being played back when the play button 41 is pressed. The save button 45 is a button for saving the specified SMF data in association with the timbre type information of the timbre with the maximum timbre score or the timbre type information of the timbre specified on the timbre list screen 4b (described later).

[0029] The phrase history button Fb displays the year, month, day, hour, and minute when the specified SMF data was specified. The optimum timbre button Tb displays the timbre type that achieved the highest timbre score among the timbre scores based on the specified SMF data displayed in the phrase history button Fb, so that it can be compared with other timbre types. When the optimum timbre button Tb is pressed, the display screen of the display device 4 switches to the timbre list screen 4b, which will be described later in FIG. 2.

[0030] The optimum timbre button Tb displays the name of the timbre type "Mute Trumpet," which is the timbre type information of the timbre with the highest timbre score that corresponds to the specified SMF data specified for "2024.2.29 11:59" displayed on the phrase history button Fb. This allows user H to easily grasp the timbre type that is most highly related to the specified SMF data that he or she has specified and has the highest musical affinity. Note that while the chat screen 4a and the timbre list screen 4b display timbre type names as timbre type information, this is not limiting, and other timbre type information, such as timbre numbers, may also be displayed.

[0031] Next, the tone color list screen 4b will be described with reference to Fig. 2. Fig. 2(a) is a diagram showing the tone color list screen 4b, and Fig. 2(b) is a diagram showing the tone color list screen 4b when the slide bar 47 is moved all the way to the right.

[0032] 2(a) and 2(b), the tone list screen 4b displays the play button 41, panic button 44, and save button 45 described above in FIG. 1, as well as tone buttons 46tp1, 46tp2, ..., 46a1, 46a2, ..., 46b1, 46b2, ..., 46y1, 46y2, ..., 46ws1, 46ws2, ... (hereinafter, when the tone button 46tp1 etc. is not to be distinguished, they will be referred to as "tone buttons 46"), a slide bar 47, a tone score display area 48, and a back button 49 for switching the display screen of the display device 4 to the chat screen 4a described above in FIG. 1.

[0033] The tone color buttons 46 are buttons that display the name and tone color score of each tone color type. When a tone color button 46 is pressed, the tone color type corresponding to that tone color button 46 is applied to the musical sound output by the play button 41. The tone color buttons 46 are displayed in ascending or descending order of tone color score, or by tone color category.

[0034] Specifically, the timbre buttons 46tp1, 46tp2, ... display the names of timbre types (timbre type information) and timbre scores in descending order of timbre score. The timbre buttons 46tp1, 46tp2, ... are displayed on the far left in the display area, with "Matched" displayed above them, indicating that the timbre scores (comprising all timbre categories) are arranged in descending order. In FIG. 2( a), the timbre type with the highest timbre score is "Mute Trumpet," so the timbre button 46tp1 displays the name of that timbre type, "Mute Trumpet," and a timbre score of "28.27." The timbre buttons 46tp2, 46tp3, ... then display the names and timbre scores of the timbre types with the second, third, and so on highest timbre scores (comprising all timbre categories), respectively.

[0035] 2A, three tone color buttons 46tp1, 46tp2, and 46tp3 are displayed, and the fourth tone color button 46tp4 and subsequent buttons are displayed by scrolling the area displaying these tone color buttons 46tp1, 46tp2, and 46tp3 up and down with the mouse 2. The same operation as scrolling up and down on the tone color button 46 with the mouse 2 applies to the other tone color buttons 46.

[0036] To the right of the tone color buttons 46tp1, ..., tone color buttons 46a1, ..., 46b1, ... for each tone color category are displayed. Here, tone color categories refer to divisions that classify tone types according to the properties of the sound. Specifically, tone color categories are set as broad conceptual names that encompass musical instruments of the same type, such as "piano" and "guitar." The tone color category "piano" includes various piano tone types such as "acoustic piano" and "mezzo piano," while the tone color category "guitar" includes various guitar tone types such as "acoustic guitar" and "electric guitar."

[0037] The timbre buttons 46a1, 46a2, ... display the names and timbre scores of the timbre category to which the timbre type with the highest timbre score belongs, in descending order of timbre score. Specifically, since the timbre type with the highest timbre score is "Mute Trumpet," the timbre buttons 46a1, 46a2, ... display the names and timbre scores of the timbre types that belong to the "Solo Brass" timbre category to which "Mute Trumpet" belongs, in descending order of timbre score. The name of that timbre category, "Solo Brass," is displayed above the timbre buttons 46a1, 46a2, ....

[0038] The timbre buttons 46b1, 46b2, ... display the names and timbre scores of the timbre types in the timbre category "Ensemble String" to which the timbre type "Tape String" with the second highest timbre score belongs, in descending order of timbre score.

[0039] Furthermore, the tone buttons 46c1, 46c2, ... display the names and scores of tone types in the "Mallet" tone category, to which the tone type with the fourth highest tone score, "Ax Steel Drum," belongs, in descending order of tone score. This is because the tone type with the third highest tone score, "AX Romantic," belongs to the "Solo Brass" tone category and is already displayed on tone button 46a2. This prevents the names and scores of tone types displayed in different tone categories from being displayed in duplicate.

[0040] Similarly, the names of the timbre types belonging to each timbre category and their timbre scores are displayed for each timbre category on the timbre buttons 46d1, ..., 46y1, ... below. The slide bar 47 will now be described. The slide bar 47 is a display part for switching between the timbre buttons 46 displayed on the timbre list screen 4b. The slide bar 47 is configured to be slidable left and right, and the timbre buttons 46 displayed correspond to the position of the slide bar 47.

[0041] By moving the slide bar 47 to the left, the timbre buttons 46 (e.g., timbre button 46a1) of the timbre categories to which timbre types with high timbre scores belong are displayed, and by moving the slide bar 47 to the right, the timbre buttons 46 (e.g., timbre button 46y1) of the timbre categories to which timbre types with low timbre scores belong are displayed. Figure 2(a) shows the timbre list screen 4b when the slide bar 47 is moved to the left end, and Figure 2(b) shows the timbre list screen 4b when the slide bar 47 is moved to the right end.

[0042] 2B, when the slide bar 47 is moved to the right end, the timbre buttons 46ws1, ws2, ... are displayed to the right of the timbre button 46y1, ... of the timbre category "Perc," to which the timbre type with the smallest timbre score belongs. Contrary to the timbre buttons 46tp1, ... described above, the timbre buttons 46ws1, ws2, ... display the names of the timbre types and their timbre scores in ascending order (comprehensively covering all timbre categories).

[0043] The timbre buttons 46a1, ..., 46b1, ... on the timbre list screen 4b display, for each timbre category, the timbre buttons 46 for the timbre types belonging to that timbre category. This allows user H to understand the degree of association between the input SMF data and the timbre type based on the timbre score, and also allows user H to easily understand the names and timbre scores of other timbre types belonging to the same timbre category as the input timbre type. This allows user H to easily understand the tendency for timbre types that are suitable for the input SMF data, thereby encouraging user H to select a suitable timbre type.

[0044] Furthermore, the timbre buttons 46a1, ..., 46b1, ... display timbre buttons 46 of timbre types belonging to the same timbre category in descending order of timbre score. This allows user H to easily grasp the timbre types that are highly (weakly) related to the SMF data entered in one timbre category, further encouraging user H to select a suitable timbre type in one timbre category.

[0045] Furthermore, the timbre buttons 46tp1, 46tp2, ... display the names and timbre scores of timbre types in descending order of timbre score, regardless of timbre category. This allows user H to grasp the timbre type that is most closely related to the input SMF data. In addition, user H can grasp the timbre type that is most closely related to the input SMF data, regardless of timbre category, which encourages user H to flexibly select timbre types.

[0046] While the name of the timbre type and the timbre score are displayed on the timbre button 46, this is not limiting and the display of the timbre score may be omitted. Also, while the timbre types are assigned to the timbre buttons 46a1, 46a2, etc. in descending order of timbre score, this is not limiting and the timbre types may be assigned to the timbre buttons 46a1, 46a2, etc. in alphabetical order of the timbre type names or in the order of the Japanese syllabary pronunciation, for example.

[0047] The timbre score display area 48 is a display area that illustrates the distribution of timbre scores for each timbre type. The timbre score display area 48 illustrates the position of each timbre type when arranged in a space of a predetermined dimension with a circle, and the circle is colored according to the magnitude of the timbre score of that timbre type.

[0048] 2, the circle with the highest timbre score is colored white, and the circle with the lowest timbre score is colored black. The other circles are colored gray, with lighter gray (closer to white) indicating a higher timbre score and darker gray (closer to black) indicating a lower timbre score. By checking the distribution of circles displayed in the timbre score display area 48 and the colors of the circles, user H can intuitively grasp the relevance of the timbre type to the input SMF data. Note that instead of coloring the circles according to the magnitude of the timbre score of that timbre type, colors according to the timbre category may be used.

[0049] Furthermore, when any of the tone color buttons 46tp1, 46a1, etc. is selected, the circle of the corresponding tone color score in the tone color score display area 48 may be made to flash, etc. Conversely, when any of the tone color score circles in the tone color score display area 48 is selected, the corresponding tone color button 46tp1, 46a1 may be made to flash, etc.

[0050] Next, the electrical configuration of the PC 1 and the server 50 will be described with reference to Figure 3. Figure 3(a) is a block diagram showing the electrical configuration of the PC 1 and the server 50. The PC 1 has a CPU 20, a hard disk drive (HDD) 21, and a RAM 22, which are each connected to an input / output port 24 via a bus line 23. The input / output port 24 is further connected to the above-mentioned mouse 2, keyboard 3, display device 4, MIDI keyboard 5, and speaker 6, as well as a communication device 7 that communicates with other information processing devices such as the server 50 via the Internet Nt.

[0051] The CPU 20 is a computing device that controls each component connected via a bus line 23. The HDD 21 is a rewritable, non-volatile storage device that stores programs executed by the CPU 20, fixed value data, and the like, and stores a tone color output program 21a and SMF selected tone color data 21b. When the tone color output program 21a is executed by the CPU 20, the PC main processing of Figure 4 is executed. The SMF selected tone color data 21b stores the specified SMF data described above in Figures 1 and 2, and tone color type information of the tone with the highest tone color score or the tone color type information of the tone selected by the tone button 46 in Figure 2, in association with each other.

[0052] The RAM 22 is a memory in which the CPU 20 rewrites various work data, flags, etc. when the program is executed, and is provided with an input SMF memory 22a that stores designated SMF data and the year, month, day, hour, and minute when the designated SMF data was specified, a timbre score memory 22b, and a selected timbre memory 22c that stores timbre type information for the timbre with the highest timbre score or timbre type information for the timbre selected with the timbre button 46 in Figure 2. The timbre score memory 22b will be described with reference to Figure 3(b).

[0053] 3B is a diagram schematically illustrating the timbre score memory 22b. As shown in FIG. 3B, the timbre score memory 22b stores, for each timbre type, a timbre number, a timbre category, a timbre type name ("Tone Name" in the figure), and a timbre score, all of which are associated with each other. Note that the timbre score memory 22b may store timbre type information other than the timbre category, timbre number, or timbre type name.

[0054] The timbre number is a unique number assigned to each type of timbre, and the PC 1 and the server 50 use this timbre number to identify which type of timbre is being processed. The timbre score memory 22b stores in advance the timbre numbers, timbre categories, and timbre type names (timbre type information) corresponding to the timbre types, and the timbre scores in the timbre score memory 22b are updated each time a timbre distance score is received from the server 50.

[0055] 3A, the server 50 has a CPU 51, an HDD 52, and a RAM 53, which are all connected to an input / output port 55 via a bus line 54. The input / output port 55 is further connected to a communication device 57 that communicates with other information processing devices such as the PC 1 via the Internet Nt.

[0056] The CPU 51 is a computing device that controls each unit connected via a bus line 54. The HDD 52 is a rewritable non-volatile storage device that stores programs executed by the CPU 51, fixed value data, etc., and stores a server program 52a and a tone color output model 52b. When the server program 52a is executed by the CPU 51, the server main processing of FIG. 5 is executed. The tone color output model 52b stores the above-mentioned tone color output model M. The creation of the tone color output model M will be described later with reference to FIGS. 7 to 9.

[0057] The RAM 53 is a memory that stores various work data, flags, etc. in a rewritable manner when the CPU 51 executes the program. A plurality of PCs 1 are connected to the Internet Nt, and SMF data is transmitted from each PC 1 to the server 50. The server 50 calculates a timbre distance score according to the received SMF data, and transmits the calculated timbre distance score together with the corresponding timbre type information to each PC 1.

[0058] Next, the processing executed by the CPU 20 of the PC 1 and the CPU 51 of the server 50 will be described with reference to Figures 4 to 6. Figure 4 is a flowchart of the PC main processing. The PC main processing is processing executed by the CPU 20 when an instruction to execute the tone color output program 21a is given in the PC 1.

[0059] 1 on the display device 4 (S1). After the process of S1, it is confirmed whether MIDI data has been input from the MIDI keyboard 5 (S2). If it is confirmed in the process of S2 that MIDI data has been input from the MIDI keyboard 5 (S2: Yes), the input MIDI data is converted into SMF data, and the converted SMF data is stored in the input SMF memory 22a (S3).

[0060] Specifically, when an instruction to start inputting MIDI data is given using the recording button 40, acquisition of MIDI data from the MIDI keyboard 5 begins, and the MIDI data acquired up until an instruction to end inputting MIDI data is given using the recording button 40 is converted into SMF data and stored in the input SMF memory 22a.

[0061] On the other hand, if it is not confirmed in the process of S2 that MIDI data has been input from the MIDI keyboard 5 (S2: No), it is confirmed whether SMF data has been input by dragging and dropping into the note display area 43 (S4). If it is confirmed in the process of S4 that SMF data has been input (S4: Yes), the input SMF data is saved in the input SMF memory 22a (S5). After the processes of S3 and S5, the SMF data in the input SMF memory 22a is transmitted to the server 50 via the communication device 7 (S6). The process of the server 50 will now be described with reference to FIG. 5.

[0062] 5 is a flowchart of the server main processing. The server main processing is executed by the CPU 51 after the server 50 is started. The server main processing first checks whether SMF data has been received from the PC 1 via the communication device 57 (S50). If it is confirmed in the processing of S50 that SMF data has been received (S50: Yes), the received SMF data is input into the timbre output model M stored in the timbre output model 52b, and the timbre distances for all timbre types for the SMF data are obtained (S51).

[0063] After the process of S51, the timbre distance score based on the reciprocal of the timbre distance obtained in the process of S51 is transmitted together with the corresponding timbre type information to the PC 1 that transmitted the SMF data received in the process of S50 via the communication device 57 (S52). If it is not confirmed in the process of S50 that the SMF data has been received (S50: No), or after the process of S52, the processes from S50 onwards are repeated.

[0064] Returning to Fig. 4, after the process of S6, it is confirmed whether or not the timbre distance scores and timbre type information have been received from the server 50 via the communication device 7 (S7). If it is not confirmed in the process of S7 that the timbre distance scores and timbre type information have been received from the server 50 (S7: No), the process of S7 is repeated. On the other hand, if it is confirmed in the process of S7 that the timbre distance scores and timbre type information for all timbre types have been received from the server 50 (S7: Yes), the timbre distance scores for all received timbre types are converted into timbre scores using the method described above with reference to Figs. 1 and 2, and these timbre scores are stored in the timbre score area for each corresponding timbre type in the timbre score memory 22b (S8).

[0065] After the processing of S8, the timbre type information of the timbre that obtained the highest timbre score among the timbre scores converted in the processing of S8 is saved in the selected timbre memory 22c (S9). As a result, the timbre type information of the timbre that is most closely related to the SMF data in the input SMF memory 22a is saved in the selected timbre memory 22c and used for outputting musical tones, as will be described later in Fig. 6. At this time, it is also possible to implement filtering by user H based on the timbre category or user H's attribute information or preference information, and saving the timbre type information of the timbre that is most closely related as a result to the selected timbre memory 22c.

[0066] If it is not confirmed in the process of S4 that SMF data has been input (S4: No), or after the process of S9, a display and sound process (S10) is executed, and after the process of S10, the processes from S2 onwards are repeated. The display and sound process of S10 will be described with reference to FIG.

[0067] FIG. 6 is a flowchart of the display / sound generation process. The display / sound generation process begins by checking the display screen displayed on the display device 4 (S20). If it is confirmed in S20 that the chat screen 4a is displayed on the display device 4 (S20: "chat screen"), a phrase history button Fb is displayed showing the year, month, day, hour, and minute for which the SMF data in the input SMF memory 22a was specified (S21). After S21, the name of the timbre type that achieved the highest timbre score among the timbre scores calculated in S8 of FIG. 4 is displayed on the optimal timbre button Tb (S22). At this time, filtering may be performed by user H based on the timbre category or user H's attribute information or preference information, and the resulting timbre with the highest relevance may be displayed on the optimal timbre button Tb.

[0068] After the process of S22, it is confirmed whether the optimum timbre button Tb has been pressed (S23). If it is confirmed in the process of S23 that the optimum timbre button Tb has been pressed (S23: Yes), the display screen of the display device 4 is switched to the timbre list screen 4b (S24). On the other hand, if it is not confirmed in the process of S23 that the optimum timbre button Tb has been pressed (S23: No), the process of S24 is skipped.

[0069] If it is confirmed in the processing of S20 that the timbre list screen 4b is displayed on the display device 4 (S20: "Tone list screen"), the timbre buttons 46 are displayed as described above in Fig. 2 based on the names and timbre scores of each timbre type stored in the timbre score memory 22b (S25). During the processing of S25, filtering may be performed by user H based on the specification of a timbre category or the user H's attribute information or preference information, and only the specified timbre buttons may be displayed.

[0070] After the process of S25, it is confirmed whether the tone color button 46 has been pressed (S26). If it is confirmed in the process of S26 that the tone color button 46 has been pressed (S26: Yes), the tone color type information corresponding to the pressed tone color button 46 is saved in the selected tone color memory 22c (S27). On the other hand, if it is not confirmed in the process of S26 that the tone color button 46 has been pressed (S26: No), the process of S27 is skipped.

[0071] After the processes of S26 and S27, it is confirmed whether the back button 49 has been pressed (S28). If it is confirmed in the process of S28 that the back button 49 has been pressed (S28: Yes), the display screen of the display device 4 is switched to the chat screen 4a (S29). On the other hand, if it is not confirmed in the process of S28 that the back button 49 has been pressed (S28: No), the process of S29 is skipped.

[0072] After the processes of S23, S24, S28, and S29, it is checked whether the play button 41 has been pressed (S30). If it is confirmed in the process of S30 that the play button 41 has been pressed (S30: Yes), a musical tone is output from the speaker 6 by applying the timbre type of the selected timbre memory 22c to the SMF data of the input SMF memory 22a (S31).

[0073] A sound source (not shown) is used to output the musical tones in the process of S31. The sound source stores the timbre numbers or names of timbre types in association with timbre data (waveform data). When the timbre number or name of a timbre type is designated, the corresponding timbre data is applied to the SMF data in the input SMF memory 22a, and the resulting musical tone is output from the speaker 6.

[0074] On the chat screen 4a, immediately after SMF data is saved in the input SMF memory 22a, the tone color type information of the tone color most closely related to the SMF data (i.e., the specified SMF data) is saved in the selected tone color memory 22c and set as the tone to be applied to the output of musical tones. This allows the user H to listen to the musical tones to which the tone color most closely related to the SMF data is applied simply by pressing the play button 41 without having to specify the tone color type on the tone color list screen 4b or the like, thereby eliminating the need for the user H to go to the trouble of selecting a tone color type.

[0075] On the other hand, if it is not confirmed in the process of S30 that the play button 41 has been pressed (S30: No), the process of S31 is skipped.

[0076] After the processes of S30 and S31, it is confirmed whether the panic button 44 has been pressed (S32). If it is confirmed in the process of S32 that the panic button 44 has been pressed (S32: Yes), all musical tones being output are set to note off (S33). On the other hand, if it is not confirmed in the process of S32 that the panic button 44 has been pressed (S32: No), the process of S33 is skipped.

[0077] After the processes of S32 and S33, it is checked whether the save button 45 has been pressed (S34). If it is confirmed in the process of S34 that the save button 45 has been pressed (S34: Yes), the SMF data in the input SMF memory 22a and the timbre type information in the selected timbre memory 22c are stored in the SMF selected timbre data 21b in association with each other (S35).

[0078] User H can output musical tones using the SMF data and timbre type information stored in the SMF selection timbre data 21b, or can select the type of timbre to apply to other SMF data by referring to the combination of SMF data and timbre type information stored in the SMF selection timbre data 21b. After the processing of S34 and S35, the display and sound generation processing is terminated.

[0079] 7 and 8, a method for creating the tone color output model M will be described. Fig. 7 is a diagram for explaining a method for creating the tone color output model M. The tone color output model M is trained using SMF data Md and tone color waveform data Td.

[0080] The SMF data Md is SMF data prepared in advance for learning. The SMF data Md is associated with a timbre number and a timbre category (i.e., timbre type information) of a timbre type that is optimal (corresponding) to the performance information (phrase). A plurality of SMF data Md (e.g., 100,000 pieces) are provided, and each is used to create a timbre output model M.

[0081] The timbre waveform data Td is pre-prepared timbre waveform (wave) data for learning, and is created by recording musical tones to which a predetermined timbre type is applied. The timbre waveform data Td also corresponds to the timbre number and timbre category (i.e., timbre type information) of the applied timbre type. Multiple pieces of timbre waveform data Td (e.g., 4,200 pieces) are provided, and each is used to create the timbre output model M.

[0082] The MIDI feature Mf and the acoustic feature Tf used for learning are created from the SMF data Md and the tone color waveform data Td, respectively. First, a method for creating the MIDI feature Mf will be described.

[0083] Each piece of training SMF data Md is input to the tokenizer 80. The tokenizer 80 is a module that converts the SMF data Md, which is composed of hexadecimal numbers, into a human-readable token format composed of character strings. In this embodiment, the tokenizer 80 is configured using known technology to convert the SMF data Md into tokens.

[0084] Each piece of SMF data Md converted into tokens by the tokenizer 80 is input to a natural language processor 81, which extracts MIDI features Mf. Specifically, the natural language processor 81 is a module that processes the input tokens using a model similar to that used in natural language processing, thereby extracting MIDI features Mf, which are the features of the tokens. The MIDI features Mf are considered to be high-dimensional (e.g., 800-dimensional) information that represents features such as the context of the tokens converted from the SMF data Md.

[0085] Next, a method for creating the acoustic feature Tf will be described. Each piece of timbre waveform data Td is input to the acoustic analysis unit 82, which extracts the acoustic feature Tf. Specifically, the acoustic analysis unit 82 is a module that acoustically analyzes the input timbre waveform data Td to extract the acoustic feature Tf, which is a feature of the timbre waveform data Td. The acoustic feature Tf is considered to represent the features of the waveform information of the sound in the timbre waveform data Td as high-dimensional (e.g., 200-dimensional) information.

[0086] The MIDI feature Mf extracted by the natural language processing unit 81 and the acoustic feature Tf extracted by the acoustic analysis unit 82 are input to an error learning unit 83, and a tone color output model M is created based on the results of learning in the error learning unit 83. The error learning unit 83 will now be described with reference to FIG.

[0087] 8A is a diagram illustrating the error learning unit 83. The error learning unit 83 includes a learning control unit 84 that controls the learning process, a MIDI neural network 85, an acoustic neural network 86, and an error function calculation unit 87.

[0088] The MIDI neural network 85 is a neural network that converts the input MIDI feature Mf, which is a high-dimensional MIDI feature Mf, into a MIDI feature Me, which is relatively low-dimensional (e.g., 64-dimensional) information. The number of dimensions of the MIDI feature Me is set to be the same as the number of dimensions of the acoustic feature Te converted by the acoustic neural network 86, which will be described later. In this manner, the MIDI feature Me is created from the training SMF data Md. Note that the MIDI neural network 85 may be configured using other techniques for converting high-dimensional MIDI feature Mf into low-dimensional MIDI feature Me.

[0089] The acoustic neural network 86 is a neural network that converts high-dimensional acoustic features Tf into acoustic features Te, which are relatively low-dimensional (e.g., 64-dimensional) information. The number of dimensions of the acoustic features Te is set to be the same as the number of dimensions of the MIDI features Me. In this way, the acoustic features Te are created from the training timbre waveform data Td.

[0090] The MIDI feature Me and acoustic feature Te created in this manner are input to an error function calculation unit 87. The error function calculation unit 87 is a module that calculates an error function called "triplet loss" using the input MIDI feature Me and acoustic feature Te. The learning control unit 84 performs learning of the MIDI neural network 85 and the acoustic neural network 86 by an error backpropagation method using the triplet loss calculated by the error function calculation unit 87. Here, distance learning using triplet loss will be described with reference to FIG. 8(b).

[0091] Fig. 8(b) is a diagram illustrating distance learning using triplet loss. Note that Fig. 8(b) shows a schematic two-dimensional representation for explaining the space in which the MIDI feature Me and the acoustic feature Te are defined. In distance learning using triplet loss, first, one of the created MIDI features Me is selected and set as "A data" (A: Anchor).

[0092] Then, from among the created acoustic features Te, acoustic features Te in the same timbre category as the A data are randomly selected and set as "P data" (P: Positive). Furthermore, from among the created acoustic features Te, acoustic features Te in a timbre category different from that of the A data are randomly selected and set as "N data" (N: Negative).

[0093] Next, the distance AP between the A data and the P data in the 64-dimensional space and the distance AN between the A data and the N data in the 64-dimensional space are calculated (calculation of the error function). In this embodiment, the distances AP and AN are Euclidean distances, but other distances such as Mahalanobis distance may also be used. The learning control unit 84 updates the parameters of the MIDI neural network 85 and the acoustic neural network 86 using the error backpropagation method so that the distance AP becomes closer than the distance AN. This parameter update (i.e., learning) is repeated while switching between the A data, P data, and N data. Note that the calculation of the error function and the error backpropagation method use known techniques, so detailed explanations will be omitted.

[0094] Furthermore, one of the created MIDI features Me is selected and set as the A data, and among the acoustic features Te, an acoustic feature Te with the same timbre number as the A data is randomly selected and set as the P data, and among the acoustic features Te, an acoustic feature Te with a timbre number different from that of the A data is randomly selected and set as the N data. Next, the distance AP between the A data and the P data in the above-mentioned 64-dimensional space and the distance AN between the A data and the N data in the 64-dimensional space are calculated (calculation of the error function). The learning control unit 84 updates the parameters of the MIDI neural network 85 and the acoustic neural network 86 using the error backpropagation method so that the distance AP is closer than the distance AN. This parameter update (i.e., learning) is repeated while switching between the A data, P data, and N data.

[0095] A timbre output model M is created by using a known model creation method for the trained MIDI neural network 85 and acoustic neural network 86. When using the timbre output model M, SMF data can be input to the timbre output model M, and the timbre distances to all timbres based on the correspondence between the input SMF data and each timbre type are output together with timbre type information for each timbre, thereby configuring the timbre output model M to calculate the timbre distances to all timbres for the input SMF data. Through distance learning using such triplet loss, the timbre output model M can define the correspondence between the input SMF data and all timbre types.

[0096] Here, we will explain the structure and operation of the created timbre output model M. The timbre output model M includes a portion of the functions of the trained MIDI neural network 85 and acoustic neural network 86, and by using a portion of the functions of these neural networks 85 and 86, the timbre distance of each timbre for the input SMF data is output.

[0097] First, the MIDI neural network 85 included in the tone output model M acquires coordinates in the above-mentioned 64-dimensional space corresponding to the SMF data input to the tone output model M. Specifically, the input SMF data is converted into 800-dimensional MIDI features Mf by functions equivalent to the tokenizer 80 and natural language processor 81 included in the tone output model M, and then converted into 64-dimensional MIDI features Me by the function of the MIDI neural network 85 included in the tone output model M. In this way, coordinates in the 64-dimensional space corresponding to the input SMF data are acquired. Hereinafter, the 64-dimensional space defined by the function of the MIDI neural network 85 will be referred to as the "MIDI feature 64-dimensional space."

[0098] Separately, in the tone output model M, coordinates in a 64-dimensional space of tones based on the acoustic features Te are allocated in advance by the acoustic neural network 86 included in the tone output model M. Such tone coordinates are allocated for all tones. Hereinafter, the 64-dimensional space defined by the function of the acoustic neural network 86 and in which the coordinates of each tone are allocated will be referred to as the "acoustic feature 64-dimensional space." This acoustic feature 64-dimensional space is optimized (for example, by adjusting the position of the coordinates of each tone) in the above-mentioned learning so that it can be superimposed on the above-mentioned MIDI feature 64-dimensional space.

[0099] Next, the 64-dimensional MIDI feature space in which the coordinates of the input SMF data are arranged is superimposed on the 64-dimensional acoustic feature space in which the coordinates of each timbre are arranged. Then, the timbre distance between the coordinates of the input SMF data and the coordinates of each timbre is calculated (analyzed) and output. In this way, the timbre distance of each timbre relative to the input SMF data is output from the timbre output model M.

[0100] As described above, the timbre output model M is created by performing distance learning (machine learning) on ​​the MIDI neural network 85 and the acoustic neural network 86 so that the relationship between the MIDI feature Me based on the training SMF data Md and the acoustic feature Te based on the training timbre waveform data Td is appropriate. In the distance learning, an error function is calculated in which an arbitrary MIDI feature Me is defined as A data, an arbitrary acoustic feature Te having the same timbre number or timbre category as the A data is defined as P data, and an arbitrary acoustic feature Te having a timbre number or timbre category different from the A data is defined as N data.

[0101] In this way, by performing learning based on the MIDI features Me and acoustic features Te, which have common information such as the timbre number or timbre category, it is possible to create a timbre output model M that can output the timbre distance, which is the correspondence between the input SMF data and the timbre type, and the timbre type information of each timbre for all timbre types.

[0102] The MIDI features Me used to create the tone output model M are created by converting the SMF data Md into tokens, processing the tokens with a model similar to that used in natural language processing, and then converting the extracted MIDI features Mf. Processing with a language model makes it possible to extract in detail the "context" of the performance information contained in the SMF data Md, such as the relationship between note-on and note-offs. This allows the MIDI features Me to include in detail the characteristics of the performance information of the SMF data Md.

[0103] The acoustic feature Te is created by converting the acoustic feature Tf, which is obtained by acoustically analyzing the timbre waveform data Td. The acoustic analysis allows for detailed extraction of the characteristics of the waveform data of the sound contained in the timbre waveform data Td. This allows the acoustic feature Te to include detailed sound characteristics of the timbre waveform data Td.

[0104] In this way, the MIDI feature Me, which contains detailed characteristics of the performance information of the SMF data Md, and the acoustic feature Te, which contains detailed characteristics of the sound of the timbre waveform data Td, are used to create the timbre output model M based on the results of distance learning using an error function. This allows the timbre output model M to accurately output the timbre distance for each timbre type corresponding to the input SMF data.

[0105] Here, the MIDI feature Me and the acoustic feature Te are created as the same 64-dimensional data, which allows for easy and accurate distance learning based on different types of data, namely, the SMF data Md and the timbre waveform data Td.

[0106] Then, the timbre output model M is created based on the results of distance learning using an error function between the MIDI feature amount Me and the acoustic feature amount Te. By using distance learning, the MIDI neural network 85 and the acoustic neural network 86 are trained using an error function between the MIDI feature amount Me and the acoustic feature amount Te, so that even if there is a limit to the number of samples of the SMF data Md and the timbre waveform data Td, it is possible to create a timbre output model M that can accurately output the timbre distances between the input SMF data and all timbre types.

[0107] Furthermore, error functions are calculated both based on timbre category and based on timbre number. This allows for the calculation of more error functions that represent the close / distant relationship between the MIDI feature Me and the acoustic feature Te, making it possible to learn both the rough correspondence based on timbre category and the more detailed correspondence based on timbre number. This also makes it possible to create a timbre output model M that can accurately output the timbre distance between the input SMF data and all timbre types.

[0108] Next, referring to Fig. 9, a model creation process for creating a tone color output model M based on the creation method described above with reference to Figs. 7 and 8 will be described. Fig. 9 is a flowchart of the model creation process. The model creation process is executed by the PC 1, and it is preferable that the information processing device be equipped with a GPU (Graphics Processing Unit). Note that the model creation process may also be executed by another information processing device such as the server 50.

[0109] The model creation process begins by setting training SMF data Md and the timbre number, timbre category, and timbre type name corresponding to the SMF data Md (S60). After S60, training timbre waveform data Td and the timbre number, timbre category, and timbre type name corresponding to the timbre waveform data Td are set (S61). After S61, the MIDI feature creation process (S62) is executed.

[0110] Specifically, the MIDI feature creation process acquires one piece of SMF data Md for learning set in the process of S60, converts the acquired SMF data Md into tokens, and processes the SMF data Md converted into tokens using a model similar to that used in natural language processing to extract MIDI features Mf.

[0111] The extracted MIDI feature Mf is associated with the timbre number, timbre category, and timbre type name set in the SMF data Md on which the MIDI feature Mf is based, and stored in the HDD 21, etc. The process of obtaining the SMF data Md, converting it to tokens, extracting the MIDI feature Mf, and storing the extracted MIDI feature Mf is performed for all SMF data Md set in the process of S60.

[0112] After the MIDI feature creation process of S62, an acoustic feature creation process (S63) is executed. Specifically, the acoustic feature creation process acquires one piece of timbre waveform data Td for learning set in the process of S61, acoustically analyzes the acquired timbre waveform data Td, and extracts an acoustic feature Tf.

[0113] The extracted acoustic features Tf are associated with the timbre number, timbre category, and timbre type name set in the timbre waveform data Td on which the acoustic features Tf are based, and are saved in a HDD, etc. The process of obtaining timbre waveform data Td, converting to tokens, extracting acoustic features Tf, and saving the extracted acoustic features Tf is performed for all of the timbre waveform data Td set in the process of S61.

[0114] After the acoustic feature creation process of S63, a timbre category learning process (S64) is executed. Specifically, the MIDI feature Mf and acoustic feature Tf extracted in the processes of S62 and S63 are converted into MIDI feature Me through the MIDI neural network 85 and acoustic feature Tf through the acoustic neural network 86. Using these MIDI feature Me and acoustic feature Te, a triplet loss (error function) based on the timbre category described above in FIG. 8(b) is calculated. Based on the calculated error function, the MIDI neural network 85 and acoustic neural network 86 are trained using the error backpropagation method.

[0115] In this embodiment, the process proceeds to calculation of the error function without waiting for all MIDI features Mf to be converted into MIDI features Me and all acoustic features Tf to be converted into acoustic features Te, but the present invention is not limited to this. The process may proceed to calculation of the error function after converting all MIDI features Mf into MIDI features Me and all acoustic features Tf into acoustic features Te.

[0116] After the timbre category learning process of S64, a timbre number learning process (S65) is executed. Specifically, the timbre number learning process calculates a triplet loss (error function) based on the timbre number shown in FIG. 8B using the MIDI feature Me and acoustic feature Te created in the process of S64. Based on the calculated error function, the MIDI neural network 85 and the acoustic neural network 86 are trained using the error backpropagation method.

[0117] After the tone color number learning process of S65, the MIDI neural network 85 and the acoustic neural network 86 learned in the processes of S64 and S65 are created as a tone color output model M by a known model creation method (S66). After the process of S66, the created tone color output model M is transmitted to the server 50 via the Internet Nt or the like (S67). After the process of S67, the model creation process ends.

[0118] The above has been explained based on the above embodiment, but it can be easily imagined that various improvements and modifications are possible.

[0119] In the above embodiment, the timbre output model M outputs a timbre distance according to the level of association as the correspondence between the input SMF data and the timbre type. However, this is not limited to this, and the timbre output model M may output a timbre distance score or a timbre score as the correspondence between the input SMF data and the timbre type, or may output another value according to the correspondence between the input SMF data and the timbre type.

[0120] In the above embodiment, the timbre score based on the SMF data input to the PC 1 is displayed on the display device 4, but the output of the timbre score is not limited to this. For example, the timbre score and the name of the timbre type may be output aloud, or the timbre score and the name of the timbre type may be printed. Furthermore, the display device 4 is not limited to displaying the timbre score, and may instead display, for example, the timbre distance or timbre distance score received from the server 50, or other values ​​based on the timbre distance or timbre distance score.

[0121] In the above embodiment, the timbre output model M is created by distance learning using triplet loss, but this is not limiting and the timbre output model M may be created using other distance learning techniques.Furthermore, the timbre output model M is not limited to being created by distance learning and may be created by other learning techniques.

[0122] In the above embodiment, the MIDI feature Me and the acoustic feature Te are used in the distance learning, but this is not limiting. For example, instead of the MIDI feature Me, the MIDI feature Mf may be used, or SMF data Md converted into tokens may be used, or the SMF data Md may be used as is, or SMF data Md converted into data other than the above may be used. Furthermore, instead of the acoustic feature Te in the distance learning, the acoustic feature Tf may be used, or the timbre waveform data Td may be used as is, or the timbre waveform data Td converted into data other than the above may be used.

[0123] In the above embodiment, in the distance learning, both error functions based on timbre category and timbre number are calculated, but this is not limited to this. For example, only the error function based on timbre category may be calculated, or only the error function based on timbre number may be calculated. Furthermore, in the distance learning, the error function is calculated using the A data as the MIDI feature Me and the P data and N data as the acoustic feature Te, but this is not limited to this. For example, the A data may be the acoustic feature Te, and the P data and N data may be the MIDI feature Me.

[0124] In the above embodiment, the tone output model M is configured to output the correspondence between the input SMF data and the tone type (specifically, the tone distance), but this is not limited to this. For example, the tone output model M may be configured to output the correspondence between the input SMF data and the sound effect, including the effect. In this case, the sound effect type may be used instead of the tone type, and further, in creating the tone output model M, a recording of the sound to which the sound effect has been applied may be used instead of the tone waveform data Td.

[0125] In the above embodiment, the types of tones are classified by tone category and used to display the tone score and for distance learning of the tone output model M, but this is not limited to this, and the types of tones may also be classified by music genre such as pop, rock, jazz, etc., or by information related to other tones.

[0126] 5, the input of SMF data to the timbre output model M and the acquisition of the timbre distance and timbre distance score for the SMF data are performed by the server 50. However, this is not limiting. For example, the input of SMF data to the timbre output model M and the acquisition of the timbre distance and timbre distance score for the SMF data may be performed by the PC 1.

[0127] In the above embodiment, the tone data for each tone type is stored in a sound source (not shown) used to output the musical tones in the process of S31 in Fig. 6, but this is not limiting. For example, in the process of S7 in Fig. 4, the corresponding tone data may be received from the server 50 along with the tone distance score and tone type information, and used in the process of S31 in Fig. 6.

[0128] In this case, in the process of S8 in Fig. 4, the timbre score and timbre data based on the timbre distance score received from the server 50 are stored in the timbre score memory 22b in association with a preset timbre number, timbre category, and timbre type name (timbre type information). Then, in the process of S9 in Fig. 4 or the process of S27 in Fig. 6, the timbre type information corresponding to the timbre with the highest timbre score or the timbre selected with the timbre button 46 in Fig. 2, and the timbre data acquired from the timbre score memory 22b corresponding to that timbre type information, are stored in the selected timbre memory 22c. Then, in the process of S31 in Fig. 6, the timbre data stored in the selected timbre memory 22c is used to output musical tones. This reduces the memory usage within the sound source.

[0129] In the above embodiment, the PC 1 is used as an example of a computer that executes the tone color output program 21a, but the present invention is not limited to this and the tone color output program 21a may be executed by an information processing device such as a smartphone or tablet terminal, or an electronic musical instrument such as a synthesizer. Furthermore, the tone color output program 21a may be stored in a ROM or the like, and the present invention may be applied to a dedicated device (tone color output device) that executes only the tone color output program 21a.

[0130] 1 PC (computer, tone color output device) M tone color output model Md SMF data (music learning data) Td tone color waveform data (tone color learning data) 21a tone color output program S2 to S5 input steps S6, S51 analysis steps S22, S25 output steps, output means S31 musical tone output step

Claims

1. A timbre output program that causes a computer to execute an output step of outputting the correspondence between input music data and timbre types together with identification information for said timbre types.

2. The tone output program according to claim 1, characterized in that the correspondence is the degree of association between each tone type and the music data, and the output step outputs identification information and the degree of association of the tone type according to the degree of association.

3. The tone color output program according to claim 2, wherein said output step outputs identification information and a degree of association of the tone color type having the highest degree of association among said tone color types.

4. The tone output program according to claim 2, characterized in that the tone types are classified into tone categories that correspond to the properties of the sound, and the output step outputs the identification information and relevance of the tone types together for each tone category to which the tone types belong.

5. A tone output program as described in claim 4, characterized in that the output step outputs the identification information and relevance of the tone type for each tone category in the order of the tone category to which the tone type with the highest relevance belongs.

6. A tone output program as described in claim 4 or 5, characterized in that the output step, in outputting the identification information and relevance of the tone types for each tone category, outputs the identification information and relevance of the tone types belonging to the tone category in descending order of the relevance.

7. The tone output program according to claim 2, further comprising causing the computer to execute a musical tone output step of outputting a musical tone obtained by applying the tone type having the highest degree of association among the tone types to the music data.

8. The tone color output program according to claim 1, wherein the correspondence is analyzed using a tone color output model that has been machine-learned.

9. The tone output program of claim 8, wherein the tone output model is constructed based on the results of machine learning using music learning data consisting of a combination of music data for training and identification information for the type of tone corresponding to that music data, and tone learning data consisting of a combination of tone waveform data for training and identification information for the type of tone corresponding to that tone waveform data.

10. A tone output program that causes a computer to execute an input step of inputting music data, an analysis step of analyzing the degree of association between the music data input in the input step and a tone type using a machine-learned tone output model, and an output step of outputting the degree of association analyzed in the analysis step together with identification information for the tone type, wherein the tone output model analyzes the degree of association by superimposing a space based on the music data and a space based on tone waveform data.

11. A tone color output device comprising output means for outputting the correspondence between input music data and tone color types together with identification information for said tone color types.

12. A tone output device comprising: an input means for inputting music data; an analysis means for analyzing the degree of association between the music data input by the input means and a tone type using a machine-learned tone output model; and an output means for outputting the degree of association analyzed by the analysis means together with identification information for the tone type, wherein the tone output model analyzes the degree of association by superimposing a space based on the music data and a space based on tone waveform data.

13. A tone color output method comprising an output step of outputting the correspondence between input music data and tone color types together with identification information for said tone color types.

14. A tone output method comprising: an input step for inputting music data; an analysis step for analyzing the degree of association between the music data input in the input step and a tone type using a machine-learned tone output model; and an output step for outputting the degree of association analyzed in the analysis step together with identification information for the tone type, wherein the tone output model analyzes the degree of association by superimposing a space based on the music data and a space based on tone waveform data.

Citation Information

Patent Citations

  • Musical information retrieving system and method thereof

    JP1996123818A

  • Information processing device and method, and program

    WO2020218075A1

  • Information processing system, electronic musical instrument, information processing method, and program

    WO2022153875A1