Sentence providing method, program and sentence providing device

A neural network-based method generates explanatory text for chord progressions, addressing the challenge of understanding chord progressions in music by providing clear textual explanations, improving user comprehension.

JP2025106476APending Publication Date: 2025-07-15YAMAHA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025064504
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-03-23
Filing Date
2025-04-09
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing methods for understanding chord progressions in music rely on image-based representations that require a certain level of music theory knowledge, making it difficult for users without such knowledge to comprehend the meaning behind chord progressions.

Method used

A method using a trained neural network model, such as RNN or CNN, to generate explanatory text from code input data arranged in time series, providing explanations on chord progressions, functions, and connections between chords based on learned relationships from a teacher dataset.

Benefits of technology

Enables users to understand chord progressions through textual explanations, enhancing comprehension without requiring music theory knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106476000001_ABST
    Figure 2025106476000001_ABST
Patent Text Reader

Abstract

To provide an explanation sentence related to a plurality of codes arrayed in time series from the codes.SOLUTION: A sentence providing method includes: providing code input data having codes arranged in time series for a learned model utilizing RNN (Recurrent Neural Network), CNN (Convolutional Neural Network), etc., having learned the relation between code string data having codes arranged in time series and an explanation sentence related to codes included in the code string data; and obtaining a sentence corresponding to the code input data from the learned model.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method for providing text.

Background Art

[0002] A plurality of chords that make up a piece of music change the impression given to the listener by their combination (for example, a chord progression arranged in time series). An ordinary listener receives the impression from the music sensuously. The listener can verify the impression by analyzing the music based on music theory such as chord progressions. A technique of detecting a cadence in a musical score showing a chord progression of a piece of music, displaying an arrow symbol at the cadence part, and varying the color according to the type of cadence is disclosed in, for example, Patent Document 1. The user can recognize the part corresponding to the cadence and the type of cadence among the chords included in the music by the arrow symbol and the color.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The types of chord progressions included in a piece of music are various. Understanding the types of chord progressions is an important factor in verifying the impression of the music. According to the technique described in Patent Document 1, the user can recognize the part and type of the cadence included in the musical score of the music from image information such as an arrow symbol and a color. However, without a certain level of knowledge of music theory, the user cannot understand the meaning from the image information and cannot utilize the obtained information.

[0005] One of the objects of the present disclosure is to provide an explanatory text about chords from a plurality of chords arranged in time series.

Means for Solving the Problem

[0006] According to an embodiment of the present disclosure, there is provided a sentence providing method including obtaining a sentence corresponding to code input data in which codes are arranged in time series based on the relationship between code sequence data in which codes are arranged in time series and explanatory sentences regarding the codes included in the code sequence data.

Effect of the Invention

[0007] According to the present disclosure, it is possible to provide explanatory sentences regarding codes from a plurality of codes arranged in time series.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Embodiments for Carrying Out the Invention

[0009] Hereinafter, one embodiment of the present disclosure will be described in detail with reference to the drawings. The following embodiments are examples, and the present disclosure is not construed as being limited to these embodiments. In the drawings referred to in this embodiment, the same parts or parts having the same functions are denoted by the same reference numerals or similar reference numerals (reference numerals with A, B, etc. attached after the numbers), and the repeated description thereof may be omitted.

[0010] [1-1. Article Provision System] FIG. 1 is a diagram showing an article provision system in one embodiment. The article provision system 1000 includes an article provision server 1 (article provision device) and a model generation server 3 connected to a network NW such as the Internet. The communication terminal 9 is a smartphone, a tablet personal computer, a laptop personal computer, a desktop personal computer, etc., and is connected to the network NW to perform data communication with other devices.

[0011] The article providing server 1 receives data related to music from the communication terminal 9 via the network NW, and transmits an explanatory article corresponding to the chord progression included in the music to the communication terminal 9. In the communication terminal 9, the explanatory article can be displayed on the display. The article providing server 1 generates an explanatory article using a learned model obtained by machine learning. When the learned model 155 receives code input data in which the codes constituting the music are arranged in time series, it outputs an explanatory article regarding the chord progression by arithmetic processing using a neural network. The model generation server 3 executes machine learning processing using a teacher data set, and generates a learned model used in the article providing server 1. Hereinafter, the article providing server 1 and the model generation server 3 will be described.

[0012] [1-2. Article providing server] The article providing server 1 includes a control unit 11, a communication unit 13, and a storage unit 15. The control unit 11 includes a CPU (processor), a RAM, and a ROM. The control unit 11 performs processing according to the instructions described in the program by executing the program stored in the storage unit 15 by the CPU. This program includes a program 151 for performing the article providing process described later.

[0013] The communication unit 13 includes a communication module, connects to the network NW, and performs transmission and reception of various data with other devices.

[0014] The storage unit 15 includes a storage device such as a non-volatile memory, and stores the program 151 and the learned model 155. In addition, various data used in the text providing server 1 are stored. The storage unit 15 may store a music database 159. The music database 159 will be described in another embodiment. The program 151 only needs to be executable by a computer, and may be provided to the text providing server 1 in a state stored in a computer-readable recording medium such as a magnetic recording medium, an optical recording medium, a magneto-optical recording medium, or a semiconductor memory. In this case, the text providing server 1 only needs to be equipped with a device for reading the recording medium. The program 151 may also be provided by being downloaded via the communication unit 13.

[0015] The learned model 155 is generated by machine learning in the model generation server 3 and provided to the text providing server 1. When the code input data is provided, the learned model 155 outputs an explanatory text about the code by arithmetic processing using a neural network. In this example, the learned model 155 is a model using an RNN (Recurrent Neural Network). The learned model 155 uses Seq2Seq (Sequence To Sequence), that is, it includes an encoder and a decoder described later. The code input data and the explanatory text are an example of data described in time series, and the details will be described later. Therefore, it is preferable that the learned model 155 adopts a model advantageous for handling time series data.

[0016] The learned model 155 may be a model that utilizes LSTM (Long Short Term Memory) or GRU (Gated Recurrent Unit). The learned model 155 may also be a model that utilizes CNN (Convolutional Neural Network), Attention (Self-Attention, Source Target Attention), etc. The learned model 155 may be a model that combines multiple models. The learned model 155 may be stored in another device connected via the network NW. In this case, the article providing server 1 may be connected to the learned model 155 via the network NW.

[0017] [1-3. Model Generation Server] The model generation server 3 includes a control unit 31, a communication unit 33, and a storage unit 35. The control unit 31 includes a CPU (processor), a RAM, and a ROM. The control unit 31 performs processing according to the instructions described in the program by executing the program stored in the storage unit 35 using the CPU. This program includes a program 351 for performing model generation processing described later. The model generation processing is processing for generating the learned model 155 using the teacher dataset.

[0018] The communication unit 33 includes a communication module, is connected to the network NW, and transmits and receives various data to and from other devices.

[0019] The storage unit 35 includes a storage device such as a non-volatile memory, and stores the program 351 and the teacher dataset 355. Various other data used in the model generation server 3 is also stored. The program 351 only needs to be executable by a computer, and may be provided to the model generation server 3 in a state stored in a computer-readable recording medium such as a magnetic recording medium, an optical recording medium, a magneto-optical recording medium, or a semiconductor memory. In this case, the article providing server 1 only needs to be equipped with a device for reading the recording medium. The program 351 may be provided by being downloaded via the communication unit 33.

[0020] A plurality of teacher data sets 355 may be stored in the memory unit 35. The teacher data set 355 is data that associates code sequence data 357 with explanatory text data 359, and is used when generating the learned model 155. Details of the teacher data set 355 will be described later.

[0021] [2. Article providing process] Next, the article providing process (article providing method) executed by the control unit 11 in the article providing server 1 will be described. The article providing process is started, for example, in response to a request from the communication terminal 9.

[0022] FIG. 2 is a flowchart showing the article providing process in one embodiment. The control unit 11 waits until receiving music code data from the communication terminal 9 (step S101; No). The music code data is data in which a plurality of codes constituting a music piece are arranged in time series. For example, the music code data is described as, for example, "CM7 - Dm7 - Em7 - ···". When arranging in time series, each code may be arranged in units of a predetermined unit period (for example, 1 measure, 1 beat, etc.), or may be arranged in order without considering the unit period. For example, assuming that in the above example each code is arranged in units of 1 measure, when the first code in the above example continues for 2 measures, the music code data is described as "CM7 - CM7 - Dm7 ···". On the other hand, assuming that the number of measures is not considered, the music code data is described as "CM7 - Dm7 - ···" as in the above example.

[0023] When the user operates the communication terminal 9 to instruct the transmission of music code data, the communication terminal 9 transmits the music code data to the text providing server 1. When the text providing server 1 receives the music code data, the control unit 11 generates code input data from the music code data (step S103). The code input data is described by converting each code included in the music code data into a predetermined format. Specifically, the code input data is data in which each code is described by a chroma vector.

[0024] FIGS. 3 and 4 are diagrams for explaining chroma vectors representing codes in one embodiment. As shown in FIGS. 3 and 4, the chroma vector is described by the presence "1" or absence "0" of a sound corresponding to each note name (C, C#, D,...). In this example, each code is converted into data (hereinafter referred to as conversion data) that combines a chroma vector corresponding to the constituent sound, a chroma vector corresponding to the bass sound, and a chroma vector corresponding to the tension sound. In this example, the conversion data is data in which three chroma vectors are described as matrix data (3×12). The conversion data may be described as vector data in which three chroma vectors are connected in series.

[0025] FIG. 3 is an example showing the code "CM7" as conversion data. FIG. 4 is an example showing the code "C / B" as conversion data. "CM7" and "C / B" have the same constituent sounds, but different bass sounds and tension sounds. Therefore, according to the conversion data, "CM7" and "C / B" can be distinguished. That is, the conversion data can clearly represent the function of the code. The conversion data only needs to include at least the chroma vector of the constituent sound, and may not include at least one or both of the bass sound and the tension sound. The structure of the conversion data may be appropriately set according to the required result.

[0026] The code input data is data obtained by arranging the conversion data in time series. As in the example described above, when the music code data is "CM7-Dm7-···", the code input data is described as data arranged in the order of the conversion data corresponding to "CM7", the conversion data corresponding to "Dm7", ···.

[0027] Returning to FIG. 2 to continue the explanation. The control unit 11 provides the code input data to the learned model 155 (step S105). The control unit 11 executes arithmetic processing by the learned model 155 (step S107) and acquires sentence output data from the learned model 155 (step S109). The control unit 11 transmits the acquired sentence output data to the communication terminal 9 (step S111). The sentence output data corresponds to the above-described explanatory text and includes a character group indicating an explanation regarding the code defined by the code input data. The explanatory text includes at least one of a first character group explaining the chord progression, a second character group explaining the function of the chord, and a third character group explaining the connection technique between chords. In this example, the explanatory text includes the first character group, the second character group, and the third character group.

[0028] FIG. 5 is a diagram for explaining an example of an explanatory text obtained from the code input data. The learned model 155 includes an encoder (also referred to as an input layer) that generates intermediate state data by performing arithmetic operations on the provided code input data using an RNN, and a decoder (also referred to as an output layer) that outputs sentence output data by performing arithmetic operations on the intermediate state data using an RNN. More specifically, a plurality of conversion data included in the code input data are provided to the encoder in time series order. The decoder outputs a plurality of characters (character groups) arranged in time series as an explanatory text. Here, the character may mean a single word (morpheme) classified by morphological analysis. The intermediate state may also be referred to as a hidden state or a hidden layer.

[0029] The code input data shown in Fig. 5 is presented as the music code data "CM7 - Dm7 - Em7 - ···", but as described above, it is data in which each code is described as conversion data. In the code input data, an end marker (EOS: End Of Sequence) is attached to the part where the code ends. When this code input data is provided to the learned model 155, the learned model 155 outputs text output data including the explanatory text exemplified in Fig. 5.

[0030] According to the code input data shown in Fig. 5, the text output data, that is, the explanatory text, is composed of combinations of the following character groups. "In the first half, the diatonic codes are sequentially ascended to form a two-five between Fm7 and Bb7. While Bb7 functions as a substitute code for the subdominant minor code Fm7, it also functions as the back code of the dominant 7th code E7 for the subsequent Am7. In the second half, through the repetition of the substitute code Am7 for the tonic code CM7 and the substitute code AbM7 for the subdominant minor code Fm7, the root note moves up and down by a semitone while the 3rd and 7th notes are held the same."

[0031] Among the explanatory text shown in Fig. 5, the first character group (explaining the chord progression) corresponds to "forming a two-five between Fm7 and Bb7".

[0032] Among the explanatory text shown in Fig. 5, the second character group (explaining the function of the chords) corresponds to "Bb7 functions as a substitute code for the subdominant minor code Fm7", "Bb7 functions as the back code of the dominant 7th code E7 for the subsequent Am7", "the substitute code Am7 for the tonic code CM7", "the substitute code AbM7 for the subdominant minor code Fm7". Actually, in the explanatory text, the explanations of the two functions regarding Bb7 are summarized and expressed as "Bb7 functions as ··· while it also functions as ··· for the subsequent Am7."

[0033] Among the explanatory texts shown in Fig. 5, the third character group (explanation of the code connection technique) corresponds to "sequentially ascending the diatonic codes" and "a progression in which the root note moves up and down by a semitone while the 3rd and 7th notes are held the same by repeating Am7 and AbM7. Actually, in the explanatory text, regarding sequentially ascending the diatonic codes, it is expressed as "sequentially ascending the diatonic codes and... " so as to be connected to the following text.

[0034] The text output data obtained in this way is transmitted to the communication terminal 9 that transmitted the music code data. As a result, an explanatory text corresponding to the music code data is provided to the user of the communication terminal 9. The above is the explanation of the text providing process.

[0035] [3. Model Generation Process] Next, the model generation process (model generation method) executed by the control unit 31 in the model generation server 3 will be described. The model generation process is started in response to a request from a terminal or the like used by the administrator of the model generation server 3. The model generation process may also be started in response to a user's request, that is, a request from the communication terminal 9.

[0036] Fig. 6 is a flowchart showing the model generation process in an embodiment. The control unit 31 acquires the teacher data set 355 from the storage unit 35 (step S301). As described above, the teacher data set 355 includes the code sequence data 357 and the explanatory text data 359 that are associated with each other. The code sequence data 357 is described in the same format as the code input data. That is, the code sequence data 357 is described as data in which the codes represented by the conversion data are arranged in time series.

[0037] The explanatory text data 359 is data including an explanatory text as shown in FIG. 5. This explanatory text is a text that explains the code defined by the code sequence data 357. As described above, the explanatory text includes at least one of a first character group that explains the code progression, a second character group that explains the function of the code, and a third character group that explains the concatenation technique between the codes. In this example, the explanatory text data 359 is attached with an identifier for identifying words obtained by morphological analysis of the explanatory text. Each word is described by a "One Hot Vector". The analysis text may be described in word representations such as "word2vec" and "GloVe".

[0038] The code sequence data 357 included in the teacher dataset 355 includes, in this example, the sequence of codes corresponding to one piece of music, and is attached with at least one end marker EOS. The teacher dataset 355 can take various forms. Using FIGS. 7 to 9, a plurality of examples that the teacher dataset 355 can take will be described.

[0039] FIGS. 7 to 9 are diagrams for explaining an example of the teacher dataset. In the teacher dataset 355 described here, the code sequence data 357 corresponding to the codes of the music is shown in a plurality of divisions (music intervals CL(A) to CL(E)). Here, the music intervals CL(A) to CL(E) each correspond to a divided range such as a phrase constituting the music, for example, a range in units of 8 measures, and each includes a plurality of codes arranged in time series. Each music interval does not have to be the same length as other music intervals.

[0040] The code sequence data 357 shown in FIG. 7 has a format in which the codes corresponding to the music intervals CL(A) to CL(E) are described in a series, and includes the end marker EOS only at the end of the data.

[0041] The code sequence data 357 in FIG. 8 has a format in which the codes corresponding to the music sections CL(A) to CL(E) are divided and described for each music section. An end marker EOS is described at the division position. The section divided by the end marker EOS is called a division area. A plurality of music sections may be included in one division area. On the other hand, in this example, one music section is not included in a plurality of division areas.

[0042] As shown in FIG. 8, the code sequence data 357 in FIG. 9 has a format in which the codes corresponding to the music sections CL(A) to CL(E) are divided for each music section, and then the codes of the music sections before and after the music section in each division area are added and described. That is, in the code sequence data 357 in FIG. 9, a plurality of consecutive music sections are arranged in one division area, and at least one music section is included in a plurality of division areas. In this example, three consecutive music sections are arranged in each division area, and only two consecutive music sections are arranged in the first and last division areas. The number of consecutive music sections is not limited to this example.

[0043] The explanatory text data 359 includes explanatory texts ED(A) to ED(E) corresponding to the music sections CL(A) to CL(E) respectively. For example, the explanatory text ED(A) includes a character group that explains the code corresponding to the music section CL(A). The explanatory text data 359 shown in FIGS. 8 and 9 is divided by the end marker EOS in the same manner as the code sequence data 357.

[0044] Returning to FIG. 6, the description will be continued. The control unit 31 inputs the code sequence data 357 into a model for machine learning (here, called a training model) (step S303). The training model is a model that performs arithmetic processing using the same neural network (in this example, an RNN) as the learned model 155. The training model may be the learned model 155 stored in the text providing server 1.

[0045] The control unit 31 performs machine learning by error backpropagation using the value output from the training model in response to the input of the code sequence data and the explanatory text data 359 (step S305). Specifically, the weight coefficients in the neural network of the training model are updated by machine learning. If there is another teacher dataset 355 to be learned (step S307; Yes), machine learning is performed using the remaining teacher dataset 355 (steps S301, S303, S305). If there is no other teacher dataset 355 to be learned (step S307; No), the control unit 31 ends the machine learning.

[0046] The control unit 31 generates the trained model that has completed machine learning as a learned model (step S309) and ends the model generation process. The generated learned model is provided to the text providing server 1 and used as the learned model 155. In this way, the learned model 155 is a model that has learned the correlation between the codes defined in the code sequence data 357 and the explanatory text related to those codes.

[0047] When the code sequence data 357 input in the machine learning includes the end marker EOS in the middle of the data as shown in FIGS. 8 and 9, the control unit 31 resets the intermediate state at the time point of the end marker EOS. That is, in machine learning, the codes in a specific divided region and the codes in other divided regions divided from that region are not treated as continuous time-series data. In the example shown in FIG. 8, the codes in a specific music section and the codes in a music section different from that music section are treated as independent time-series data. On the other hand, the codes included in one music section are treated as time-series data.

[0048] In the example shown in FIG. 9, separated music sections, for example, music section CL(B) and music section CL(E), are not included in one divided area and are treated as independent time-series data. On the other hand, music section CL(B) and music section CL(C) may be included in one divided area or may be included in different divided areas. Therefore, depending on the divided area, the codes of music section CL(B) and the codes of music section CL(C) may be treated as a series of time-series data or may be treated as independent time-series data.

[0049] Summarizing the three exemplified teacher data sets, it is as follows. The first example is a teacher data set in which no divided area is set as shown in FIG. 7. The second example is a teacher data set in which a plurality of divided areas are set as shown in FIG. 8 and no music section is included in a plurality of divided areas. The third example is a data set in which a plurality of divided areas are set as shown in FIG. 9 and at least one music section is included in a plurality of divided areas.

[0050] In particular, by increasing the number of codes treated as time-series data as in the first example and the third example, it is possible to realize highly accurate machine learning that widely considers the context of the code sequence. By narrowing the range of context as in the third example, it is possible to exclude parts that are too far apart and have a weak relationship from the target of machine learning and realize more accurate machine learning. In machine learning, only one of these examples may be used, or a plurality of examples may be used in combination.

[0051] [4. Example of Code Interpretation] Next, the correlation between code interpretation and explanatory text will be described in more detail. Here, an example in which two-five-one (II-V-I) is detected as a typical code progression example will be described.

[0052] FIG. 10 is a diagram for explaining a chord progression detected as two-five-one. In FIG. 10, as the chord progression of two-five-one, examples (basic form, derivative form) within the scales of Cmaj or Amin, and other examples outside the scales (hidden chord, delayed resolution) are shown. The "hidden chord" means that the hidden chord is used in a part of the chord progression. The derivative form and the hidden chord corresponding to the basic form are surrounded by a single range with a dashed line. In the derivative form and the hidden chord, the parts different from the basic form are indicated by underlines. Regarding "delayed resolution", it is an example in which a change is made by inserting the chord indicated by () while taking the form of two-five-one.

[0053] The trained model 155 generated by the above-described machine learning can output, as an explanatory text, that there is a two-five-one even in a chord progression expressed other than the basic form. In the sequence of chords constituting a piece of music, there may be a chord progression that accidentally corresponds to two-five-one without intending two-five-one. Even in such a case, the trained model 155 generated by machine learning including the context can output an explanatory text considering whether the chord progression corresponds to two-five-one or not.

[0054] FIGS. 11 and 12 are diagrams for explaining examples of explanatory texts obtained from chord input data. Both FIG. 11 and FIG. 12 have a sequence of chords "Em7 - A7 - GbM7 - Ab7", but in FIG. 12, DbM7 is further added to the last part. That is, the chord located at the end of the time series by the end marker EOS is Ab7 in FIG. 11, while it is DbM7 in FIG. 12, which is different.

[0055] The trained model 155 infers that the elements of "Em7 - A7 - Ab7" in the chord input data shown in FIG. 11 are related to two-five-one, and outputs the following explanatory text as text output data. "Em7 - A7 - Ab7 is a derivative form of II - V - I (Em7 - A7 - Dm7) when regarded as the diatonic chords of the Cmaj scale in the Dmaj scale, with Dm7 changed to the back chord. GbM7 is the two - five for Ab7 in Dbmaj and is inserted to temporarily delay the resolution (cadence) to Ab7."

[0056] On the other hand, the learned model 155 infers that among the chord input data shown in FIG. 12, the element "GbM7 - Ab7 - DbM7" is related to the two - five - one, and further outputs the following explanatory text that also mentions the element "Em7 - A7 - Ab7" as the text output data. "There is a temporary modulation from Cmaj to Dmaj. GbM7 - Ab7 - DbM7 is obtained by changing II in II - V - I (Ebm7 - Ab7 - DbM7) in the Dbmaj scale to IV which is the same sub - dominant (Ebm7 → GbM7). To make the modulation smooth, a kind of two - five - one using the back chord Em7 - A7 - Ab7 is incorporated."

[0057] In this way, by performing machine learning as described above using a large number of teacher datasets 355, even if the chord sequences included in the chord input data are similar, the learned model 155 can output text output data including appropriate explanatory text considering the context of the similar parts.

[0058] FIG. 13 and FIG. 14 are diagrams for explaining modified examples of chord progressions detected as two - five - one. In the example shown in FIG. 13, when the connection technique of baseline descent is applied to the basic form of the two - five - one chord progression "Bm7(-5) - E7 - Am7", the chord sequence becomes, for example, "Bm7(-5) - Bm7(-5) / F - E7 - E7 / G# - Am7". Even in this case, it can be recognized that it is a two - five - one chord progression without being affected by the change in the baseline.

[0059] In the example shown in FIG. 14, when the concatenation technique called passing diminished is applied to the basic form of the chord progression of two-five-one, which is "Dm7 - Db7 - CM7", the chord sequence becomes, for example, "Dm7 - Ddim7 - Db7 - CM7". Even in this case, it is possible to recognize that it is the chord progression of two-five-one without being affected by the addition of Ddim7.

[0060] [5. Extraction of Specific Section] In the above-described embodiment, the chord input data may identify the sequence of all chords included in the music chord data, or may identify the sequence of some chords extracted therefrom. In the following description, the section of the music corresponding to the chords included in the chord input data is referred to as a specific section. The specific section may be set by the user or may be set by a predetermined method exemplified below.

[0061] An example of the predetermined method will be described. The chord input data provided to the learned model 155 does not have to be all of the music chord data. If characteristic parts of the music can be used, an explanatory text characteristic of the music can be obtained. Therefore, it is preferable to set such characteristic parts of the music as the specific section. The characteristic parts of the music can be set in various ways, and an example thereof will be described.

[0062] In the example described here, the control unit 11 divides the music into a plurality of predetermined determination sections (for example, the music sections described above), and sets the determination section that satisfies a predetermined condition as the specific section. In this example, by calculating the chord progression importance in each determination section, the determination section having a chord progression importance exceeding a predetermined threshold is set as the specific section.

[0063] The chord progression importance is calculated based on various data registered in the music database 159 and the chord progression in the determination section. An example of this calculation method will be described.

[0064] FIG. 15 is a diagram for explaining a music database in one embodiment. The music database 159 is stored, for example, in the storage unit 15 of the text providing server 1. In the music database 159, information on a plurality of pieces of music is registered. For example, genre information, scale information, chord appearance rate data, and chord progression appearance rate data that are associated with each other are registered.

[0065] The genre information is information indicating the genre of a piece of music such as "rock", "pop", "jazz",... The scale information is information indicating scales such as "C major scale", "C minor scale", "C# major scale",... (including keys in this example). Each scale has tones that constitute it (hereinafter referred to as scale constituent tones) set.

[0066] The chord appearance rate data indicates the ratio of each type of chord to the total number of chords in all the pieces of music registered in the music database. For example, if the total number of chords is "10000" and the number of the chord "Cm" is "100" among them, the appearance rate of that chord is "0.01".

[0067] In calculating the appearance rate of chords, for the identity of chords that are similar to each other, any of the following determination criteria exemplified below may be used. Even if the chord names are different from each other, they may be treated as different chords (e.g., "CM7" and "C / B" are different). If the constituent tones are the same for each other, they may be treated as the same chord (e.g., "CM7" and "C / B" are the same). If the constituent tones and the bass tone are the same for each other, they may be treated as the same chord (e.g., "CM7" and "G / C" are the same). Even if the constituent tones are different from each other, if they are the same except for the tension tones, they may be treated as the same chord (e.g., "CM7" and "C" are the same).

[0068] The chord progression appearance rate data indicates the ratio of each type of chord progression to the total number of chord progressions of all the music pieces registered in the music database. The chord progressions referred to here are set in advance by the user or the like. For example, if the total number of chord progressions is "20000" and among them, the number of the chord progression "Dm - G7 - CM7" is "400", the appearance rate of the chord progression is "0.02".

[0069] The criterion for determining the identity of chords may be the same as the method for determining the chord appearance rate described above. The criterion for determining the identity of chord progressions may use any of the criteria exemplified below. Chord progressions that are similar to each other may be treated as the same chord progression. For example, the forms using derivative forms and back chords with respect to the basic form shown in FIG. 10 may be treated as the same chord progression.

[0070] Chord progressions in which at least two of the chord progressions match may be treated as the same chord progression. For example, for the chord progression "Dm - G7 - CM7", "* - G7 - CM7", "Dm - * - CM7", and "Dm - G7 - *" may be treated as the same chord progression. Here, "*" indicates a chord that is not specified (any of all chords).

[0071] The chord appearance rate data and the chord progression appearance rate data include data for all music pieces. In this example, the chord appearance rate data and the chord progression appearance rate data further include data defined corresponding to each genre specified in the genre information. For example, the chord appearance rate data and the chord progression appearance rate data corresponding to the genre "rock" may include the appearance rate of chords and the appearance rate of chord progressions obtained only from the music pieces corresponding to the genre "rock". Regarding the population of the appearance rate (the total number of chords and the total number of chord progressions), data for the entire music pieces may be targeted.

[0072] In terms of chords and chord progressions, the appearance rates in the genre "rock" and the appearance rates in the genre "jazz" are different. Therefore, due to the existence of the appearance rates of chords and chord progressions for each genre, the characteristic parts of the music can be determined more accurately. Genre information does not necessarily have to be used. In this case, the chord appearance rate data and chord progression appearance rate data for each genre do not have to exist.

[0073] FIG. 16 is a diagram for explaining a method of calculating chord progression importance. In the example shown in FIG. 16, the respective index values and importance when the chord progression in the determination section is "C - Cm - CM7 - Cm7" are shown. The index values include the chord progression rarity (CP) determined for the chord progression, the scale element (S) determined for each chord constituting the chord progression, and the chord rarity (C). Based on these indexes, the chord importance (CS) for each chord and the chord progression importance (CPS) for the chord progression are calculated. Both the index values and the importance have values in the range from "0" to "1". The higher the numerical value, the more characteristic the element is indicated.

[0074] In this example, for the music, the key is C, the scale is the major scale, and the genre is pop. These pieces of information may be set in advance by the user, or may be set by analyzing the music chord data. When analyzing the music chord data, for example, it may be set in relation to chords similar by comparison with the music registered in the music database 159, or may be estimated from the order of chords using a learned model obtained by machine learning or the like.

[0075] The scale element (S) is set to "0" when all the constituent tones of the chord are included in the scale constituent tones, and is set to "1" when any of the constituent tones of the chord is not included in the scale constituent tones. This is because a chord containing a tone not included in the scale constituent tones can be said to be a characteristic part of the music.

[0076] The chord rarity (C) is obtained by a predetermined calculation formula. The calculation formula is determined such that the higher the chord appearance rate, the smaller the chord rarity (C). In the case of the C major scale, since C and CM7 have relatively high chord appearance rates, the chord rarity (C) is set to a relatively small value.

[0077] The chord progression rarity (CP) is obtained by a predetermined calculation formula. The calculation formula is determined such that the higher the chord progression appearance rate, the smaller the chord progression rarity (CP). In this example, since the appearance rate of the chord progression "C - Cm - CM7 - Cm7" is extremely low, the chord progression rarity (CP) is set to a high value of "1".

[0078] The chord importance (CS) is calculated using the scale factor (S), the chord rarity (C), and the chord progression rarity (CP). In this example, the calculation formula is CS = a×S + b×C + c×CP, where a = 1 / 4, b = 1 / 4, and c = 1 / 2. The chord progression importance (CPS) is the average value of the chord importance (CS).

[0079] The chord progression importance (CPS) obtained in this way indicates that the larger the value (the closer to "1"), the more unusual the chord progression is compared to other pieces. That is, it can be said that the determination interval with a large chord progression importance (CPS) is a characteristic part of the piece.

[0080] The above method for calculating the index values and importance is just an example, and various calculation methods can be adopted as long as the importance of the chord progression (the characteristic part of the piece) is obtained as a whole. Subsequently, a method for generating chord input data using a specific interval will be described. This method for generating chord input data can be replaced with, for example, the process in step S103 shown in FIG. 2.

[0081] FIG. 17 is a flowchart showing a process of generating code input data in one embodiment. The control unit 11 sets a key, a scale, and a genre (step S1031). As described above, the key, the scale, and the genre may be obtained by being received from the communication terminal 9 according to user settings, or may be obtained by analyzing music code data. The control unit 11 divides a music piece into a plurality of determination sections (step S1033), and calculates a chord progression importance (CPS) in each determination section (step S1035).

[0082] Based on the chord progression importance (CPS) calculated for each determination section, the control unit 11 sets at least one determination section as a specific section (step S1037). Here, a determination section in which the chord progression importance (CPS) is greater than a predetermined threshold is set as the specific section. A predetermined number of determination sections may be set as the specific section in order from the determination section with the highest chord progression importance (CPS).

[0083] The control unit 11 generates code input data corresponding to the specific section (step S1039). In the code input data, an end marker EOS may be arranged for each specific section so that one specific section is arranged in one divided area, or when a plurality of consecutive determination sections are set as a plurality of specific sections, the plurality of specific sections may be arranged so as to be included in one divided area.

[0084] In this way, by providing the generated code input data to the learned model 155, the learned model 155 can generate an explanatory text for the chord progression representing the characteristic part of the music piece and output the text output data.

[0085] [6. Modification Example] The present disclosure is not limited to the above-described embodiments, but includes various other modifications. For example, the above-described embodiments have been described in detail for the purpose of clearly explaining the present disclosure, and are not necessarily limited to those having all the configurations described. For a part of the configuration of the embodiment, other configurations may be added, deleted, or replaced. Hereinafter, some modifications will be described.

[0086] (1) In the above-described embodiment, the text providing server 1 used the learned model 155 to generate the explanatory text from the code input data, but a model that does not use a neural network (for example, a rule-based model) may be used. According to the learned model 155, the accuracy of the explanatory text can be improved by using a large amount of training data sets 355 for machine learning.

[0087] According to the rule-based model, it is necessary to set a rule for generating the explanatory text from the code input data, that is, the correspondence between the information corresponding to the above-described code sequence data 357 and the information corresponding to the explanatory text data 359. This rule requires a large amount of information. For example, as the code sequence determined as the two-five-one code progression as described above, various types are assumed. Therefore, in order to improve the accuracy of the explanatory text, it is necessary to set the explanatory text for each of the many types that can be assumed. In order to reduce the amount of information, it may be necessary to simplify the explanatory text more than when using the learned model 155. Although it is assumed that the efficiency may be lower than when using the learned model 155, it is possible to generate the explanatory text from the code input data using the rule-based model.

[0088] (2) The code appearance rate data and the code progression appearance rate data may be defined to be equivalent regardless of the key of the music. For example, the code appearance rate data may be interpreted such that the code "CM7" when the key of the music is "C" and the code "EM7" when the key is "E" are the same code. The code progression appearance rate data may be interpreted such that the code progression "Dm - G7 - CM7" when the key of the music is "C" and the code progression "Fm - B7 - EM7" when the key is "E" are the same code progression.

[0089] That is, both the code appearance rate data and the code progression appearance rate data may be defined by the codes of the relative expressions with respect to the key of the music. The relative expression may be, for example, the one converted when the key is "C", or the one converted into descriptions such as "I", "II". For example, the code "Em7" in the key of "C" is expressed as "IIIm7".

[0090] In this case, the control unit 11 converts the code appearance rate data and the code progression appearance rate data defined by the codes of the relative expressions into the codes of the absolute expressions based on the set key of the music. The control unit 11 calculates the code importance (CS) and the code progression importance (CPS) based on the appearance rate by the converted codes.

[0091] (3) Instead of using the learned model 155, the article providing server 1 may use an arithmetic model such as SVM (Support Vector Machine) or HMM (Hidden Markov Model). In this case, the control unit 11 uses this arithmetic model to obtain a specific code progression from the code input data, for example, "two five one". The control unit 11 combines the obtained code progression with a predetermined template to control the explanatory text. The predetermined template is, for example, "XXXX is used for this code progression." By using the obtained code progression (in the above example, "two five one") in the "XXXX" part, an explanatory text such as "Two five one is used for this code progression." is generated. In the case of HMM, the codes included in the code input data may be sequentially input. In the case of SVM, a predetermined number of codes included in the code input data may be input together.

[0092] (4) A server storing a plurality of learned models 155 may be connected to the network NW. This server may be the model generation server 3. The article providing server 1 may select any one of the plurality of learned models 155 stored in this server and execute the above-described article providing process. The article providing server 1 may download the learned model 155 used for the article providing process from that server and store it in the storage unit 15, or may transmit the code input data by communicating with the server storing this without downloading the learned model 155 and receive the article output data.

[0093] Among the plurality of trained models 155, at least a part of the teacher datasets 355 used in machine learning are different from each other. For example, when performing machine learning using a plurality of teacher datasets 355 classified by genre (such as jazz, classical, etc.), a plurality of trained models 155 corresponding to each of the plurality of genres are generated. The teacher dataset 355 may be classified by the type of genre or the type of musical instrument. According to this classification, the code sequence data and the explanatory text data become specialized for that classification. The teacher dataset 355 may be classified by the creator of the explanatory text included in the explanatory text data 359.

[0094] For example, by providing the code input data corresponding to a piece of music classified as jazz to the trained model 155 corresponding to jazz, an explanatory text with high accuracy can be obtained. The object to which the music corresponding to the code input data is classified may be set by the user or may be set by analyzing the music.

[0095] By providing one piece of code input data to a plurality of trained models 155, a plurality of types of explanatory texts may be obtained. For example, if a plurality of trained models 155 corresponding to a plurality of creators are used, the plurality of types of obtained explanatory texts can be compared to select the one suitable for the user. Among the explanatory texts obtained from the plurality of trained models 155, a new explanatory text may be generated based on the common points.

[0096] (5) The code input data and the code sequence data 357 are not limited to being described in chroma vectors. For example, if the constituent sounds of the code are represented by data including vectors, they may be represented by other methods. Regarding the code, it may also be described in an expression to which "word2vec", "GloVe", etc. are applied.

[0097] The above is the description regarding the modification example.

[0098] As described above, according to one embodiment of the present disclosure, there is provided a sentence providing method including obtaining a sentence corresponding to code input data in which codes are arranged in time series based on the relationship between code sequence data in which codes are arranged in time series and explanatory sentences regarding the codes included in the code sequence data.

[0099] Obtaining the sentence may include providing the code data to a trained model that has learned the relationship and obtaining the sentence from the trained model.

[0100] The codes included in the code sequence data may include at least the constituent sound and the base sound of the code.

[0101] The codes included in the code sequence data may include at least the constituent sound and the tension sound of the code.

[0102] The code may be represented by data including a vector.

[0103] The code may be represented by data including a first chroma vector corresponding to the constituent sound of the code.

[0104] The code may be represented by data including a second chroma vector corresponding to the base sound of the code.

[0105] The code may be represented by data including a third chroma vector corresponding to the tension sound of the code.

[0106] The explanatory sentence may include a first character group that explains the chord progression.

[0107] The explanatory sentence may include a second character group that explains the function of the chord.

[0108] The explanatory sentence may include a third character group that explains the connection technique between chords.

[0109] Obtaining music code data in which the codes of the music are arranged in time series, and extracting the arrangement of the codes in a specific section that satisfies a predetermined condition from the music code data as the code input data may be included.

[0110] The predetermined condition may include a condition using the codes included in the music code data and the importance regarding the codes determined according to the key of the music.

[0111] The predetermined condition may include a condition using the codes included in the music code data and the importance regarding the codes determined according to the genre of the music.

[0112] A program for causing a computer to execute the article providing method may be provided. An article providing apparatus including a storage unit that stores the instructions of this program and a processor that executes the instructions may be provided.

Description of Reference Numerals

[0113] 1…Article providing server, 11…Control unit, 13…Communication unit, 15…Storage unit, 151…Program, 155…Trained model, 159…Music database, 3…Model generation server, 31…Control unit, 33…Communication unit, 35…Storage unit, 351…Program, 355…Teacher dataset, 357…Code sequence data, 359…Explanation article data, 9…Communication terminal, 1000…Article providing system

Claims

【Claim 1】 For a learned model that has learned the relationship between code sequence data in which codes are arranged in time series and explanatory text regarding the codes included in the code sequence data, code input data in which codes are arranged in time series is provided, and a sentence corresponding to the code input data is obtained from the learned model. A sentence providing method comprising the above.

Citation Information

Patent Citations

  • Device for explaining electronic musical instrument

    JP1993204301A

  • Musical score display apparatus and program for realizing musical score display method

    JP2011215181A

  • Information processing device and method, and recording medium

    WO2007077993A1

  • Musical performance information display device and musical performance information display method, musical performance information display program, and electronic musical instrument

    JP2020056938A