Information processing device, information processing method, and program
The information processing device analyzes song segments and parts to address the limitations of existing music analysis technologies, enhancing music production and marketing strategies through detailed music analysis.
Patent Information
- Application Number
- PCT/JP2025/020068
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-18
- Filing Date
- 2025-06-03
- Publication Date
- 2025-12-26
AI Technical Summary
Existing technologies for analyzing user emotional values regarding music content do not adequately analyze the music itself, limiting their effectiveness in music production and marketing strategies.
An information processing device and method that calculates the similarity between song segments and estimates parts based on song data, including chord progression, song structure, and beat analysis, to provide detailed analysis of music pieces.
Enables appropriate analysis of music pieces, facilitating effective music production and marketing strategies by providing detailed insights into song structure and fan/customer characteristics.
Smart Images

Figure JP2025020068_26122025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present disclosure relates to an information processing device, an information processing method, and a program, and more particularly to an information processing device, an information processing method, and a program that are capable of appropriately analyzing music.
[0002] There is a need to analyze the characteristics of songs and their fans and customers, and use this information to create music and develop and implement marketing strategies.
[0003] Meanwhile, a technology has been proposed for estimating a user's emotional values regarding a specific object, which can also be applied to music content (see Patent Document 1).
[0004] Therefore, it is conceivable that the technology of Patent Document 1 could be applied to statistically estimate the user's emotional values regarding a specific target, and analyze the characteristics of each song and fan or customer, thereby utilizing the results for song production and the formulation and implementation of marketing strategies.
[0005] International Publication No. 2020 / 162486
[0006] However, the analysis results using the technology of Patent Document 1 are estimates of a user's emotional values regarding a specific target, and do not analyze the music itself, such as its composition, so they may not be sufficient for use in music production or in formulating and implementing marketing strategies.
[0007] The present disclosure has been made in light of such circumstances, and in particular, is intended to enable appropriate analysis of music pieces.
[0008] An information processing device and program according to one aspect of the present disclosure include an information processing device and program that include a similarity calculation unit that calculates the similarity between segments that constitute a song based on song data, which is data about the song, and a part estimation unit that estimates parts based on the similarity between the segments that constitute the song.
[0009] An information processing method according to one aspect of the present disclosure is an information processing method that includes performing a similarity calculation process to calculate the similarity between segments that constitute a song based on song data, which is data of the song, and performing a part estimation process to estimate parts using the segments as units based on the similarity between the segments that constitute the song.
[0010] In one aspect of the present disclosure, the similarity between segments that make up a song is calculated based on song data, which is data on the song, and parts are estimated based on the similarity between the segments that make up the song.
[0011] 1 is a diagram illustrating an overview of the present disclosure. FIG. 1 is a diagram illustrating keys and scales. FIG. 2 is a diagram illustrating the circle of fifths. FIG. 3 is a diagram illustrating diatonic chords. FIG. 4 is a diagram illustrating an example configuration of a music analysis device of the present disclosure. FIG. 5 is a diagram illustrating estimation of parallel keys and modulations. FIG. 6 is a diagram illustrating estimation of parallel keys and modulations. FIG. 7 is a diagram illustrating estimation of parallel keys and modulations. FIG. 8 is a diagram illustrating estimation of parallel keys and modulations. FIG. 9 is a diagram illustrating estimation of parallel keys and modulations. FIG. 10 is a diagram illustrating estimation of parallel keys and modulations. FIG. 11 is a diagram illustrating an example configuration of a chord analysis unit. FIG. 12 is a diagram illustrating estimation of chord progressions. FIG. 13 is a diagram illustrating estimation of chord groups / functions. FIG. 14 is a diagram illustrating estimation of chord groups / functions. FIG. 15 is a diagram illustrating an example configuration of a music structure analysis unit. FIG. 16 is a diagram illustrating segment information assigned to lyrics data. FIG. 17 is a diagram illustrating the number of intros, outros, interludes, and music titles set for each segment. FIG. 18 is a diagram illustrating music structure features. FIG. 19 is a diagram illustrating a standardized edit distance matrix. FIG. 20 is a diagram illustrating estimation of parts. FIG. 21 is a diagram illustrating estimation of parts. FIG. 22 is a diagram illustrating a method of setting outliers. FIG. 23 is a diagram illustrating estimation of identical parts. FIG. 24 is a diagram illustrating a method of estimating parts from the relationship between identical parts and segments. FIG. 25 is a diagram illustrating an example configuration of a beat analysis unit. FIG. 26 is a diagram illustrating generation of beat belonging probabilities. FIG. 1 is a diagram illustrating beat types. FIG. 2 is a diagram illustrating beat types with overlapping features. FIG. 3 is a diagram illustrating a method for training a recognition model of a beat belonging probability generation unit. FIG. 4 is a diagram illustrating beat estimation. FIG. 5 is a diagram illustrating a method for training a recognition model of a beat estimation unit. FIG. 6 is a flowchart illustrating music analysis processing. FIG. 7 is a flowchart illustrating chord analysis processing. FIG. 8 is a flowchart illustrating music composition analysis processing. FIG. 9 is a flowchart illustrating beat analysis processing. FIG. 10 is a diagram illustrating a UI image example (1). FIG. 11 is a diagram illustrating a UI image example (2). FIG. 12 is a diagram illustrating a UI image example (3). FIG. 13 is a diagram illustrating a UI image example (4). FIG. 14 is a diagram illustrating a UI image example (5). FIG. 15 is a diagram illustrating an example configuration of a general-purpose computer.
[0012] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0013] Hereinafter, embodiments of the present technology will be described in the following order.
[0014] 1. Overview of the Disclosure 2. Glossary 3. Preferred Embodiments 4. Examples of Implementation by Software
[0015] <<1. Overview of the Present Disclosure>> The present disclosure is directed to enabling appropriate analysis of content, particularly music, etc. Therefore, first, an overview of the present disclosure will be described with reference to FIG.
[0016] As shown in FIG. 1, in this disclosure, a song is analyzed based on song data (including sound source data, lyric data, and chord data), which is data on content consisting of a song, and the analysis results, such as the song's intro, verse, bridge, chorus, interlude, and outro, as well as an analysis of the characteristics of each fan or customer, are presented as a UI (User Interface) image.
[0017] More specifically, in the first process St1, chord progression estimation, song structure analysis, and beat estimation are performed on the song data, and the structure divisions for each song structure, the change points in the structure, the chord progression for each structure, and the beat divisions for each structure are obtained as the estimation results and analysis results.
[0018] Then, in the second process St2, a UI image is generated and presented based on the estimation results and analysis results, which includes an analysis of the characteristics of the songs and the fans and customers.
[0019] As a result, according to the present disclosure, it is possible to appropriately analyze the elements that make up content consisting of music, which can be used for music production, the formulation and implementation of marketing strategies, and the like.
[0020] <<2. Explanation of Terms>> Below, a brief explanation will be given of terms related to music analysis that are required when describing an example configuration of a music analysis device according to the present disclosure.
[0021] <Key and Scale> A key refers to the tonic note and the series of notes (scale) starting from the tonic note. Scales are divided into major and minor scales.
[0022] As shown in the left part of Figure 2, when the tonic note is C (do), the major scale is considered to be a scale that starts from the tonic note C and continues in the order CDEFGABC. In other words, the notes C-D (do-re) are a whole step, DE (re-mi) are a whole step, EF (mi-fa) are a semitone, FG (fa-sol) are a whole step, GA (sol-la) are a whole step, AB (la-si) are a whole step, and BC (si-do) are a semitone.
[0023] Here, a whole step means that there is a note between them, and a semitone means that there is no note between them. For example, between C, D, and C, there is a C♯, which is a semitone between C, so both are whole steps. On the other hand, between E, E, and F, there is no semitone between them, so F is a semitone between E and E.
[0024] On the other hand, as shown in the right part of Figure 2, when the tonic note is A (La), the minor scale is considered to be a scale that continues in the order ABCDEFGA. That is, the notes AB (La Si) are a whole step, BC (Si Do) are a semitone, CD (Do Re) are a whole step, DE (Re Mi) are a whole step, EF (Mi Fa) are a semitone, FG (Fa So) are a whole step, and GA (So La) are a whole step.
[0025] <Relative keys> Relative keys are keys that share the same scale, consisting of the same key signature (sharp or flat), and examples of parallel keys include C major and A minor, as explained with reference to Figure 2. Figure 3 shows a list of parallel keys, called the circle of fifths, arranged on a circle in perfect fifth increments, with the outer edge of the dotted circle representing the major keys and the inner edge representing the minor keys. In other words, C major (C) on the outer edge of the top center of Figure 3 and A minor (Am) on the inner edge are parallel keys.
[0026] Each key has a set of chords that can be used, and these are the same for relative keys. These chords that can be used in each key are called available chords. However, there are also cases where chords other than the available chords are used as techniques.
[0027] <Diatonic chords> Diatonic chords are chords in a scale that fits the key. For example, in the case of the C major scale CDEFGABC (Do Re Mi Fa So La Si Do) shown in the top row of Figure 4, as shown in the second row from the top of Figure 4, the diatonic chord corresponding to C is C (Do Mi So), the diatonic chord corresponding to D is Dm (Re Fa La), the diatonic chord corresponding to E is Em (Mi So Si), the diatonic chord corresponding to F is F (Fa La Do), the diatonic chord corresponding to G is G (So Si Re), the diatonic chord corresponding to A is Am (La Do Mi), and the diatonic chord corresponding to B is Bm (♭5) (Si Re Fa).
[0028] That is, for each key, the diatonic chords are basically chords that are included in adjacent regions of the circle of fifths described with reference to Figure 3. For example, if the key is C, the seven chords shown in the left part of the third row in Figure 4 are considered diatonic chords, which correspond to range Z1 in Figure 3. Also, if the key is C♯m, the seven chords shown in the right part of the third row in Figure 4 are considered diatonic chords, which correspond to range Z2 in Figure 3.
[0029] Diatonic chords also include tonic (T), subdominant (SD), and dominant (D), each of which has a different function.
[0030] The tonic (T) is stable and functions as a tonic. The subdominant (SD) is more stable than the dominant (D) but unstable and functions as a variation on the tonic (T). The dominant (D) is unstable and has the ability to change into a tonic chord (T).
[0031] As shown in the fourth row of Figure 4, when the key is C, the tonics (T) are C, Am, Em, CM7, Am7, and Em7, the subdominants (SD) are F, Dm, FM7, and Dm7, and the dominants (D) are G, G7, Bm(♭5), and Bm7(♭5).
[0032] Other keys are shown in the table at the bottom of Figure 4. In the bottom row of Figure 4, the keys are arranged in the order of CDE♭EFGAB♭ from top to bottom in the left column, and the horizontal direction shows the degree notation (degree name notation) and the corresponding diatonic chords.
[0033] (Available Chords) Available chords are the set of commonly used chords in a particular key, and consist of diatonic chords, secondary dominant chords, related IIm (two minor) chords, and subdominant minor chords.
[0034] It is entirely possible to compose music using only diatonic chords, which are the basis of chords, but in order to broaden the range of expression without destroying the worldview of the key, secondary dominant chords, related IIm (two minor) chords, and subdominant minor chords are also used as available chords.
[0035] A secondary dominant chord is the seventh chord that is immediately to the right of the diatonic chord when using the circle of fifths.
[0036] A related IIm (two minor) chord is, for example, "IIm" / "IIm7 (♭5)" that can progress to the secondary dominant chord "V7".
[0037] A subdominant minor chord is, for example, a minorized version of a subdominant chord.
[0038] Hereinafter, available chords will be described as a group consisting of diatonic chords, secondary dominant chords, related IIm (two minor) chords, and subdominant minor chords.
[0039] (Types of parts that make up a piece of music) A piece of music is made up of multiple parts, and six types of parts will be explained below.
[0040] (1) Intro The intro is the opening part of a song that attracts the listener's attention and draws them into the world of the song.
[0041] (2) Verse: A song generally has multiple verse parts, and the verse part may take up the majority of the song. The lyrics of each verse tend to be straightforward and develop the story the song is trying to tell. The lyrics often have the same melody and rhymes each time in each verse, and tend to be a set length, such as 16 or 32 bars.
[0042] (3) Chorus The chorus is also called "Chorus," and as the word "Chorus" suggests, it is a part that is grander, fuller, and has more voices (sounds). The chorus is also the climax of the song, and is the part where the most grandiose statement is made in the story that the song is trying to express, both lyrically and musically.
[0043] (4) B-melody (Pre-chorus / Post-chorus) The B-melody is a part that has the functions of creating a smooth flow between parts, increasing tension for a more explosive chorus, and functioning as a harmonic axis for a new chord progression.
[0044] (5) C-melody (Bridge / Break) The C-melody is usually used just before the final chorus, and is the part that adds variety by using new chords and melodies to make the second half of the song more exciting.
[0045] (6) Outro The outro is the ending of a song, concluding the song. It can be similar to the intro, or it can be a unique part that “closes things out as a conclusion.”
[0046] In addition, while the B and C melodies have a relatively strong relationship with the chorus, there are also melodies D to F that have a weaker relationship with the chorus.
[0047] <<3. Preferred Embodiment>> Next, an example configuration of a music analysis device according to the present disclosure will be described with reference to the block diagram of FIG.
[0048] The music analysis device 101 in Figure 5 is composed of an input unit 111, a data acquisition unit 112, a parallel key modulation estimation unit 113, a chord analysis unit 114, a lyrics extraction unit 115, a music structure analysis unit 116, an acoustic feature extraction unit 117, a beat analysis unit 118, and a UI generation unit 119.
[0049] The input unit 111 may be configured, for example, as a keyboard or touch panel operated by the user, and accepts input of the title and artist name of the song to be analyzed and supplies it to the data acquisition unit 112.
[0050] The input unit 111 may also be configured to accept input of a sound source file of a song to be analyzed, specified by the user. The input unit 111 may include, for example, an identification unit 111a, and when accepting input as sound source data, may analyze the sound source data to identify the title name and artist name of the song, and supply the identified title name and artist name of the song to be analyzed to the data acquisition unit 112.
[0051] Furthermore, the input unit 111 may be configured to accept input of songs to be analyzed, such as trend charts, and when acquiring the title name and artist name of the songs to be analyzed, the acquired song title name and artist name may be supplied to the data acquisition unit 112. Also, when acquiring the song to be analyzed as sound source data rather than the title name and artist name, the input unit 111 may recognize the song title name and artist name from the acquired sound source data, and supply the song title name and artist name as the recognition result to the data acquisition unit 112.
[0052] The data acquisition unit 112 reads and acquires chord-related master information consisting of the relative positions from the key of C to each key, the diatonic chords of each key and their degree notations, secondary dominant chords and their degree notations, related two-minor chords and their degree notations, subdominant minor chords and their degree notations, and chord progression patterns from the external chord-related DB 103 via a network (not shown) such as the Internet, and supplies this information to the parallel key modulation estimation unit 113 and the music structure analysis unit 116.
[0053] The data acquisition unit 112 accesses the music DB 102 via the network based on the song title and artist name supplied by the input unit 111, acquires chord data and lyric data from the music data related to the song to be analyzed, and supplies the chord data to the parallel key modulation estimation unit 113 and the lyric data to the lyric extraction unit 115.
[0054] The data acquisition unit 112 accesses the music DB 102 via the network based on the song title and artist name supplied from the input unit 111, acquires the sound source data from the music data of the song to be analyzed, and supplies the sound source data to the acoustic feature extraction unit 117.
[0055] The parallel key modulation estimator 113 estimates the parallel key and modulation based on the chord data and supplies the result to the chord analyzer 114. The processing of the parallel key modulation estimator 113 will be described in detail later with reference to FIGS.
[0056] The chord analysis unit 114 estimates a chord progression pattern (chord progression for each structure), key, chord group, and chord function based on the information on the parallel key and modulation supplied from the parallel key modulation estimation unit 113 and the lyric data supplied from the lyrics extraction unit 115, and outputs the results to the UI generation unit 119. The detailed configuration of the chord analysis unit 114 will be described later in detail with reference to Figures 10 to 14.
[0057] The lyrics extraction unit 115 sets segments in the lyrics data based on line break information in the lyrics data, divides the lyrics data for each segment by line, and supplies the segments to the chord analysis unit 114 and the music structure analysis unit 116. In this case, the lyrics extraction unit 115 may set the segments and identify an intro, interlude, or outro depending on whether or not lyrics are present. Furthermore, for example, if one segment contains one line of lyrics and the number of lines is less than a predetermined number, the lyrics extraction unit 115 may combine the segment with the immediately preceding segment.
[0058] The lyrics extraction unit 115 also performs morphological analysis on the song title, calculates the degree of simplicity of the song title, and supplies the calculated degree of simplicity together with the lyrics data to the song structure analysis unit 116. At this time, for example, additional information such as "XXX ver." may be excluded from the song title.
[0059] The music composition analysis unit 116 calculates the distance between segments based on the lyric data in which the segments are set, and estimates (analyzes) the type of parts (composition divisions and change points) that make up the music, using the calculated distance between the segments as a unit. The type of parts (composition divisions and change points) that make up the music, using segments as a unit, may be, for example, an verse, a bridge, or a chorus. The music composition analysis unit 116 then supplies the UI generation unit 119 with information on the music composition, which is the estimated result and is made up of information on each part, as an analysis result. The detailed configuration of the music composition analysis unit 116 will be described later with reference to Figures 15 to 24.
[0060] The acoustic feature extraction unit 117 extracts acoustic features including basic features, instruments, moods (higher and lower order), genres, etc. based on the sound source data of the music to be analyzed, and supplies the extracted acoustic features to the beat analysis unit 118. The acoustic feature extraction unit 117 may use an existing music analysis function (for example, Deep12 (registered trademark)).
[0061] The basic features may be, for example, the shortness of the attack time when sound is pronounced, the probability of a major note, the number of notes, the sense of speed, BPM (Beats per Minute), whether it is a live performance or a programmed sound source, the score of a Hi-Fi audio, the sense of energy, the volume of high notes, the volume of low notes, the density of the frequency band, the ratio of stereo to mono, the attenuation of the sound, the length of the sound (long tones), volume, amplitude fluctuation, clarity of the tone, the frequency of pitch shifts, tempo fluctuations, the onset ratio of rhythm instruments, the probability of the presence of vocals, chord variation, the probability of using complex chords, the difficulty of key determination, and the probability of a live performance.
[0062] The instruments may be, for example, male vocals, female vocals, piano, organ, acoustic guitar, distorted guitar, wood bass, electric bass, synth bass, acoustic drums, electronic drums, breakbeats, orchestra, string instruments, brass instruments, and beeping sounds.
[0063] Moods (higher order) may be, for example, abstract, peaceful / quiet / tranquil / restful, soft / gentle, powerful / energy, and active / positive.
[0064] Mood (lower level) may be, for example, sad, calm, refreshing, happy, bright, fun, solemn, elegant, soothing, or upbeat.
[0065] Genres may be, for example, pop, rock, classical, orchestral, club / electronica, ballad, pop rock, rhythm and blues, rap, jazz, bossa nova, classical, techno, conversation, nursery rhymes, enka, piano solo, Asian (Okinawa), hard rock, and soul / funk.
[0066] The beat analysis unit 118 estimates the beat of the music piece to be analyzed in addition to surface feature quantities such as acoustic feature quantities, and supplies the beat analysis results (beat divisions for each structure) that are the estimated results to the UI generation unit 119. In this case, the beat analysis unit 118 may estimate the beat of the music piece to be analyzed based on statistics such as one measure per beat. The detailed configuration of the beat analysis unit 118 will be described later with reference to FIGS. 25 to 31.
[0067] The UI generation unit 119 generates UI images consisting of an analysis of part types based on the segments that make up the music, various types of music, and the characteristics of each fan or customer, based on the analysis results of chords supplied by the chord analysis unit 114, the analysis results of music composition supplied by the music composition analysis unit 116, and the analysis results of beats supplied by the beat analysis unit 118, and presents these on the display unit 104, which consists of a display or the like.
[0068] <Estimation of parallel keys and modulation by parallel key modulation estimator> Next, with reference to FIGS. 6 to 9, estimation of parallel keys and modulation by the parallel key modulation estimator 113 will be described.
[0069] The parallel key modulation estimation unit 113 first compares the chord data of the song with the available chords of each key in order, starting from chord number 0, and selects a key that contains all the chords up to that point as a candidate for the parallel key.
[0070] For example, consider the case where chords are arranged in order from number 0 as Bm, G, A, D, ..., D, Cm based on the chord data, as shown in Figure 6. In Figure 6, a guitar diagram for each chord is attached, and below it, based on a comparison with the relative keys of each of the 12 keys, the candidate keys included in the chords up to that point are marked with a circle, and the non-candidate keys that are not included are marked with a cross.
[0071] In the case of Figure 6, for example, for the 0th chord Bm, the parallel key modulation estimation unit 113 selects C|Am, G|Em, D|Bm, A|F♯m, E|C♯m, F♯|D♯m, and F|Dm as candidate keys for the parallel key.
[0072] Similarly, for example, for the first chord G, the parallel modulation estimator 113 selects C|Am, G|Em, D|Bm, A|F♯m, F♯|D♯m, and F|Dm as candidate keys.
[0073] Furthermore, for the second chord A and the third chord D, the parallel modulation estimator 113 selects C|Am, G|Em, D|Bm, A|F♯m, and F|Dm as candidate keys.
[0074] Next, when the number of selected keys is narrowed down to less than a predetermined number, for example, less than six of the twelve parallel keys, the parallel key modulation estimator 113 provisionally estimates the most likely key based on the following two conditions:
[0075] That is, in the example of FIG. 6, the number of candidate keys for the second chord A is narrowed down to five, and the parallel modulation estimator 113 estimates a provisional key.
[0076] The number of candidate keys selected to start estimating the provisional key does not have to be a majority, but here, an example in which it is a majority will be described as an example.
[0077] The parallel modulation estimation unit 113 measures the relative positional relationship of each chord, and sets the key in which there are chords two chords apart in the key scale as the provisional key, as the most likely key based on the relationship of the diatonic chords.
[0078] For example, in the case of Figure 6, if we look at the chords Bm, G, and A from 0th to 2nd, as shown in the circle of fifths in Figure 7, the chords Bm and G are adjacent to each other, so the distance is 1, and the chords Bm and A are also adjacent to each other, so the distance is 1. In contrast, the chords G and A are separated by the chord D, so the distance is 2.
[0079] Therefore, the parallel modulation estimator 113 sets D|Bm, which is adjacent to the chords G and A and whose distances from both are 1 and 2, as the provisional key.
[0080] As mentioned above, if it is not possible to set a temporary key based on the available chords, the parallel modulation estimation unit 113 will limit the available chords to the diatonic chords, rather than the available chords of each key, and will select the key that appears most frequently as the temporary key.
[0081] That is, as shown in Figure 8, for the 0th chord Bm, the chords D|Bm and A|F♯m are included among the diatonic chords C|Am, D|Bm, and A|F♯m; for the first chord G, the chords C|Am, D|Bm, and A|F♯m are included among the diatonic chords C|Am and D|Bm; and for the second chord A, the chords D|Bm and A|F♯m are included among the diatonic chords C|Am, D|Bm, and A|F♯m.
[0082] In this case, of the diatonic chords, keys C|Am, D|Bm, and A|F♯m, key C|Am is included only once in the first chord G, while key D|Bm is included in all three of the chords Bm, G, and A from the 0th to the 2nd, and key A|F♯m is included in the two chords Bm and A from the 0th to the 2nd.
[0083] For this reason, the parallel key modulation estimation unit 113 sets the key D|Bm, which is included in all three of the diatonic chords Bm, G, and A from the 0th to the 2nd chords, of the diatonic chord keys C|Am, D|Bm, and A|F♯m, as the temporary key.
[0084] Furthermore, even if the available chords are limited, if a provisional key cannot be set due to, for example, a tie, the parallel modulation estimator 113 may set a simple key or a key with few key signatures as the provisional key.
[0085] Next, the parallel key modulation estimation unit 113 determines whether the modulation is a borrowed chord (temporarily borrowing a chord from another key) based on whether a chord not included in the available chords of each selected key is used.
[0086] That is, for example, for the 177th code Cm in FIG. 6, a code that is not included in the available codes of each key for which a candidate key is selected is used.
[0087] In this case, since chord progressions are generally grouped into three or four cadences, the parallel key modulation estimation unit 113 looks ahead four chords and determines that a modulation has occurred if a chord that does not belong to the available chords of the key estimated by the previous chord is included.
[0088] Conversely, if all the chords belong to the available chords, the parallel key modulation estimator 113 determines that it is not a modulation but a borrowed chord.
[0089] For example, if the first four chords read ahead are E♭, B♭ on D, Cm, and A♭, the chords Cm and A♭ do not belong to the available chords in the key D|Bm, so when the chord is Cm, the parallel key modulation estimation unit 113 determines that a modulation is occurring.
[0090] Furthermore, if it is determined that a modulation has occurred, the parallel modulation estimation unit 113 estimates the key again from the point determined to be a modulation, using the same procedure.
[0091] Then, after scanning all the chords and estimating the key, the parallel modulation estimation unit 113 performs reverse scanning in order from the last chord, and if there is a part that has not been estimated, it fills in the part that has not been estimated in the reverse scanning with the key immediately before the part that has not been estimated.
[0092] For example, consider the case where the keys are estimated in the following order from left to right in forward scanning, as shown in Figure 9. Here, "None" indicates a part that could not be estimated.
[0093] In this case, the parallel modulation estimator 113 performs a reverse scan in the order of E♭|Cm, E♭|Cm, None, ... D|Bm, D|Bm, None, as shown by the arrows in Figure 9, and for the first None (right side in the figure), the previous key E♭|Cm is set as the estimated result in the reverse scan, and for the second None (left side in the figure), the previous key D|Bm is set as the estimated result in the reverse scan.
[0094] Through the above-described processing, the parallel key modulation estimation unit 113 estimates the parallel key (including borrowed chords) and modulation for each segment, and outputs the results to the chord analysis unit 114 .
[0095] <Configuration Example of Code Analysis Unit> Next, a configuration example of the code analysis unit 114 will be described with reference to the block diagram of FIG.
[0096] The chord analysis unit 114 estimates the chord progression pattern, key, and chord group / chord function based on the information on the parallel key (key) and modulation supplied by the parallel key modulation estimation unit 113, and formats and outputs the estimation results.
[0097] The chord analysis unit 114 includes a chord progression estimation unit 131, a key estimation unit 132, a chord group function estimation unit 133, and an analysis result output unit .
[0098] The chord progression estimation unit 131 estimates a chord progression pattern based on the estimation results of the parallel key or borrowed chord and modulation for each segment.
[0099] More specifically, the chord progression estimation unit 131 converts chords into degree notation (degree notation) based on the relative key, and estimates a chord progression pattern based on the pattern of the converted degree notation.
[0100] For example, if the C major scale is expressed as C, Dm, Em, F, G, Am, Bm (♭5) as shown in the upper left of Figure 11, the degree notation of the major scale expressed in Roman numerals is I, IIm, IIIm, IV, V, VIm, VIIm (♭5) as shown in the lower left of Figure 11.
[0101] Also, for example, if the minor scale of C♯m is expressed as C♯m, D♯m(♭5), E, F♯m, G♯m, A, B as shown in the upper right part of Figure 11, the degree notation of the corresponding major scale expressed in Roman numerals is I, IIm(♭5), ♭III, IVm, Vm, ♭VI, ♭VII as shown in the lower right part of Figure 11.
[0102] Next, the chord progression estimation unit 131 estimates a chord progression pattern based on the pattern expressed in degrees.
[0103] For example, as shown in Figure 12, when the key is D and the 0th to 3rd chords have a progression of Bm, G, A, D, the degree notation recognizes the chord progression as having a pattern of VI, IV, V, I. A chord progression pattern consisting of degree chords VI, IV, V, I is known to be a chord progression pattern called a Komuro progression, and is stored in advance in chord-related master information consisting of chord progression patterns stored in chord-related DB 103.
[0104] Therefore, the chord progression estimation unit 131 compares the chord progression pattern consisting of degree chords VI, IV, V, and I with the master information read and supplied from the chord-related DB 103, and estimates that the chord progression pattern is a small-room progression.
[0105] More than 100 other chord progression patterns for degree chords are registered in the master information that is read out from the chord related DB 103 and supplied.
[0106] Another example is a chord progression pattern of degree chords consisting of I, V, VI, III, IV, I, IV, V, known as a chord progression pattern called a canon progression.
[0107] The chord progression estimation unit 131 estimates a chord progression pattern through the above-described processing, and outputs the estimation result to the analysis result output unit 134 as the analysis result.
[0108] The key estimation unit 132 estimates whether the key is major or minor based on the lyrics and the four chords that correspond to them.
[0109] More specifically, the key estimation unit 132 determines that the key is minor if either of the following two conditions is met, and determines that the key is major if neither is met.
[0110] The first condition is that three or more of the chords at the beginning of the song, the chord at the beginning of the lyrics, the chord at the end of the lyrics, and the chord at the end of the song must be minor chords.
[0111] For example, if the opening chord of a song is Bm, the opening chord of the lyrics is Bm, the closing chord of the lyrics is B♭onD, and the closing chord of the song is E♭, then the opening chord of the song and the opening chord of the lyrics are both minor chords, but the closing chord of the lyrics and the closing chord of the song are major chords, and therefore do not satisfy the condition.
[0112] The second condition is that the chord at the beginning of the song and the chord at the beginning of the lyrics, or the chord at the end of the lyrics and the chord at the end of the song, is a minor chord in the estimated key (Bm in the case of D|Bm) and is not part of a specific chord progression pattern.
[0113] For example, if the chord progression from the beginning of a piece is four chords in a pattern such as Bm, G, A, D, the first chord of the piece is a minor chord (Bm) itself, but as mentioned above, it does not meet the condition because it is part of a minor progression.
[0114] Similarly, if the first four chords in the lyrics follow a pattern such as Bm, G, A, D, the first chord of the song is a minor chord (Bm) itself, but it is still part of the Komuro progression and does not meet the condition.
[0115] Furthermore, if the chord at the end of the lyrics is B♭onD, the condition is not met because it is not a minor chord (Cm after modulation) itself and is also part of the Komuro progression.
[0116] Also, if the last four chords of a piece are Cm, A♭, B♭, and E♭, the last chord, E♭, is a major chord (not a minor chord (Cm) after modulation) and is also part of a Komuro progression, so it does not meet the condition.
[0117] As described above, if neither the first nor the second condition is met, the key estimator 132 estimates the key as major. Conversely, if at least one of the first and second conditions is met, the key estimator 132 estimates the key as minor.
[0118] The key estimation unit 132 estimates the key through the above-described processing, and outputs the estimation result as an analysis result to the chord group function estimation unit 133 and the analysis result output unit 134 .
[0119] The chord group function estimation unit 133 reads a table that lists available chord groups and functions for each key, which are stored in advance in the master information stored in the chord-related DB 103, and estimates the chord group and chord function for each chord based on the estimation results of the parallel key and modulation, and the estimation result of the key.
[0120] The table that lists groups of available codes and functions for each key is, for example, a table such as that shown in FIG.
[0121] 13 shows an example of a table for the key of D major, with chord names shown on the left, functions shown on the right, and degree notation in the center. Note that the degree notation is for reference only and will not be explained here.
[0122] The chords on the left side of the diagram are, from top to bottom, diatonic chords, related IIm and secondary dominants, and subdominant minors.
[0123] The diatonic chords in the top row are, from top to bottom, D, Em, F♯m7, G, A, Am, C♯m7(♭5), and their functions are registered as T (tonic), SD (subdominant), T, SD, D (dominant), T, (D). Note that (D) in the table in Figure 13 indicates that it is a dominant that has a weaker influence toward the tonic than a D without () and is therefore weaker than a V.
[0124] In the middle row, the related IIm and secondary dominants have F♯m7 and F♯m7(♭5) registered as the related IIm, and below that, B7 is registered as the secondary dominant. Similarly, from the top, the two related IIm notes and one secondary dominant are alternately registered as G♯m7, G♯m7(♭5), C♯7, Am7, Am7(♭5), D♯7, Bm7, Bm7(♭5), E♯7, C♯m7, C♯m7(♭5), and F♯7, and the functions of each are registered as B7, B7, Em7, C7, C7, F♯m7, D7, D7, GM7, E7, E7, A7, F♯7, F♯7, and Gm7. In addition, the Resolution column shows the chord resolution for resolving (stabilizing) the chord, and the function of a related IIm is always SD (subdominant), and the function of a secondary dominant is always D (dominant).
[0125] Furthermore, the subdominant minor notes in the lower row are, from top to bottom, registered as Gm7, GmM7, Em7(♭5), B♭M7, C7, B♭7, and E♭M7, and their respective functions are registered as SDm (subdominant minor), SDm, SDm, and SDm.
[0126] FIG. 13 shows an example in which the key is D, but similar tables are registered as master information for other keys as well.
[0127] The chord group function estimation unit 133 estimates the chord group and chord function based on the estimated results of the parallel key and modulation, and the estimated result of the key, by comparing them with the information in this table.
[0128] For example, if the key is D and the zeroth through second chords are Bm, G, and A as shown in Figure 14, the chord group function estimation unit 133 estimates the chord group of the zeroth chord, Bm, as a diatonic chord and its chord function as a tonic. Similarly, the chord group function estimation unit 133 estimates the chord group of the first chord, G, as a diatonic chord and its chord function as a subdominant. Furthermore, the chord group function estimation unit 133 estimates the chord group of the second chord, A, as a diatonic chord and its chord function as a dominant.
[0129] The analysis result output unit 134 generates information on the final form from the estimation results of the parallel key modulation estimator 113, the chord progression estimator 131, the key estimator 132, and the chord group function estimator 133, and supplies this information to the UI generation unit 119. The analysis result output unit 134 may generate information on, for example, the key, modulation, chord progression pattern, the number of times each chord appears, the number of times each chord group and chord function appears, the number of borrowed chords, and the complexity of the song from the estimation results, and may also generate information on the final form by compiling this generated information.
[0130] Here, the key information that may be included in the information on the final form may be, for example, the final key determined from the results of estimating the parallel key and key. Key modulation information may include, for example, the timing of key modulation, the number of key modulations, and the key after modulation. Chord progression pattern information may include, for example, the number of occurrences of each chord progression pattern, a representative chord progression for each segment, and information on chord progression patterns for the original chords, chords converted to C key, and simplified chords (simplified chords excluding tension chords, etc.). Song complexity information may be calculated, for example, from the number of key signatures, the degree to which tension chords are incorporated, and whether or not there are key modulations.
[0131] <Configuration Example of Musical Structure Analysis Unit> Next, a configuration example of the musical structure analysis unit 116 will be described with reference to FIG.
[0132] The music composition analysis unit 116 analyzes the composition of the music to be analyzed based on the lyrics data.
[0133] More specifically, the music structure analysis unit 116 includes a music structure feature extraction unit 151 , a similarity calculation unit 152 , and a part estimation unit 153 .
[0134] The music structure feature extraction unit 151 formats the lyric data and extracts music structure feature quantities. More specifically, the music structure feature extraction unit 151 determines segments and divides and stores the lyric data for each segment by line. At this time, the music structure feature extraction unit 151 may determine segments based on, for example, line break information. The music structure feature extraction unit 151 may also determine an intro, interlude, or outro based on the presence or absence of lyrics. Furthermore, the music structure feature extraction unit 151 may also adjust a segment by combining it with the previous segment if the lyrics in one segment are less than one line and do not exceed a certain number of characters.
[0135] The music composition feature extraction unit 151 performs a morphological analysis of the song title and calculates the degree of simplicity of the song title. Here, the music composition feature extraction unit 151 may, for example, perform the morphological analysis after excluding additional information (such as xxx.ver) from the song title. Also, for example, if the title consists of only one pronoun, the degree of simplicity may be determined to be simple.
[0136] The music composition feature extraction unit 151 performs various distance calculations as music composition feature quantities. The music composition feature extraction unit 151 may perform various distance calculations as music composition feature quantities, for example, using the intro, interlude, and outro numbers of each segment, the number of times the music title appears in each segment, and on a part-by-part basis.
[0137] When performing various distance calculations, the music structure feature extraction unit 151 may, for example, perform a morphological analysis on the lyrics of each line in each segment, calculate the number of moras and the length of phonemes in each line based on the morphological analysis results, calculate the respective distances, and calculate the standardized edit distance of the lyrics of each segment. More specifically, the music structure feature extraction unit 151 may, for example, calculate the standardized edit distance for three patterns: all lyrics of each segment, the first word, and the last word. The standardized edit distance will be described in detail later.
[0138] More specifically, when lyrics data is supplied, the music structure feature extraction unit 151 allocates segments to the lyrics data. For example, the music structure feature extraction unit 151 may format the lyrics data and allocate segments based on line break information, as shown in FIG.
[0139] In FIG. 16, segment numbers RART0 to 11 are assigned to the segments, and segment sub-numbers are assigned after "-" on a row-by-row basis. Note that hereinafter, segment sub-numbers will also be referred to as segment numbers as needed. Furthermore, when expressing a specific segment number, for example, the "X"th segment, it will also be simply referred to as segment RART "X".
[0140] As shown in the top row of FIG. 16, PART0 is assigned as the segment number without lyrics.
[0141] Furthermore, the lyrics are abbreviated as "....", but the segment PART1, which consists of four lines of lyrics, is assigned segment numbers PART1-1 to PART1-4.
[0142] Below the segment PART1, segment PART2, which does not include lyrics, is assigned segment numbers PART2-1 to PART2-4.
[0143] Below segment PART2, the lyrics are omitted from the first to fourth lines, and segment numbers PART3-1 to PART3-5 are assigned to segment PART3, which includes "song title A" in the lyrics on the fifth line.
[0144] A segment PART4 without lyrics is assigned below the segment PART3.
[0145] Below the segment PART4, segments PART5 and PART6, which do not contain lyrics, are assigned segment numbers PART5-1 to PART5-4 and PART6-1 to PART6-4, respectively.
[0146] Below segment PART6, the lyrics are omitted from the first to fourth lines, and segment numbers PART7-1 to PART7-5 are assigned to segment PART7, which includes "song title A" in the lyrics on the fifth line.
[0147] Below the segment PART7, segment PART8, in which the lyrics are omitted from the first to sixth lines, is assigned segment numbers PART8-1 to PART8-6.
[0148] Below the segment PART8, segment PART9, in which the lyrics are omitted from the first to fourth lines, is assigned segment numbers PART9-1 to PART9-4.
[0149] Below segment PART9, the lyrics are omitted from the first to fourth lines, and segment numbers PART10-1 to PART10-5 are assigned to segment PART10, which includes "song title A" in the lyrics on the fifth line.
[0150] A segment PART11 without lyrics is assigned below the segment PART10.
[0151] Here, the music structure feature extraction unit 151 assigns indexes to the lyrics data after assigning segment numbers to them. For example, as shown in Fig. 16, the music structure feature extraction unit 151 may format the lyrics data, assign segment numbers to them, and infer the intro, interlude, and outro parts depending on whether lyrics are present and assign indexes to them.
[0152] 16, for example, since segments PART0, PART4, and PART11 have no lyrics, the music structure feature extraction unit 151 estimates the first part, PART0, as an intro. Similarly, the music structure feature extraction unit 151 estimates the last part, PART11, which has no lyrics, as an outro. Furthermore, the music structure feature extraction unit 151 estimates segment PART4, which is located within the intro and outro and has no lyrics, as an interlude.
[0153] Next, the music composition feature extraction unit 151 registers index information for each segment. At this time, the music composition feature extraction unit 151 may, for example, count the number of times that the title of each segment is included, and register this as index information together with the identified intro, outro, and interlude, as shown in FIG.
[0154] In the index information of Fig. 17, the "X" of segment PART "X" is set from left to right in the figure to 0 to 11, and the number of times c that the title name is included for each segment is recorded as "X:c" for segment PART "X". That is, in Fig. 16, the title name consisting of "Music Title A" is included in segment numbers PART3-5, PART7-5, and PART10-5.
[0155] 17, the number of times the song title is included is assigned to each of segments PART 3, 7, and 10 as 1. Segments PART 0, 4, and 11 are enclosed in solid lines, indicating that they are set as the intro, interlude, and outro parts, respectively.
[0156] Furthermore, the music composition feature extraction unit 151 converts the lyric data into katakana. For example, the music composition feature extraction unit 151 may perform morphological analysis on the lyric data and convert the results of the morphological analysis into katakana.
[0157] Next, the music structure feature extraction unit 151 generates music structure features for each row for each segment and outputs them to the similarity calculation unit 152. At this time, the music structure feature extraction unit 151 may, for example, calculate the number of moras (a unit of phonetic length) for each row for each segment excluding the intro, interlude, and outro, which have no lyrics, and create a matrix for each segment, and further generate music structure features by applying patterning using 0 to make the number of elements uniform.
[0158] For example, if the number of moras in each row of segment numbers PART1-1 to PART1-4 of segment PART1 is 12, 18, 18, and 20, the music structure feature extraction unit 151 converts the number of moras into a matrix like [12, 18, 18, 20] as shown in the top row on the left side of Figure 18.
[0159] Similarly, the music structure feature extraction unit 151 creates a matrix for each segment. In the case of the lyrics data of Fig. 16, the segments PART0, 4, and 11 corresponding to the intro, interlude, and outro are excluded as shown in Fig. 17, and therefore the left side of Fig. 18 shows an example in which a matrix is generated consisting of values corresponding to the number of moras for each row of the segments PART1 to PART3 and PART5 to PART10 from the top.
[0160] More specifically, in the left part of Figure 18, from top to bottom, the matrix is [12,18,18,23], [20,22,20,17], [22,26,26,9,15], [15,16,21,19], [20,24,21,16], [22,27,26,9,15], [15,15,24,13,30,26], [24,29,24,25], [22,26,26,9,15]. Note that the values that make up this matrix are fictitious values generated for the sake of explanation, and do not correspond to actual lyrics.
[0161] Furthermore, the music composition feature extraction unit 151 performs patterning using 0s as shown in the right part of FIG. 18 in order to align each matrix element.
[0162] That is, in the right part of FIG. 18, the elements are arranged from top to bottom as follows, [12,18,18,23,0,0], [20,22,20,17,0,0], [22,26,26,9,15,0], [15,16,21,19,0,0], [20,24,21,16,0,0], [22,27,26,9,15,0], [15,15,24,13,30,26], [24,29,24,25,0,0], and [22,26,26,9,15,0], so that the maximum number of elements is 6.
[0163] The similarity calculation unit 152 calculates the similarity for each segment based on the music structure feature supplied from the music structure feature extraction unit 151. At this time, the similarity calculation unit 152 may, for example, calculate a standardized edit distance of lyrics based on the number of moras for each segment based on the music structure feature supplied from the music structure feature extraction unit 151, and calculate the similarity according to the calculated standardized edit distance.
[0164] More specifically, for example, in the case of FIG. 18, when calculating the standardized edit distance between the segment PART1 and the segment PART2, the similarity calculation unit 152 first calculates the following Euclidean distance d.
[0165] That is, since the matrix consisting of the mora numbers of the segments PART1 and PART2 is [12,18,18,23,0,0] and [20,22,20,17,0,0], the similarity calculation unit 152 calculates the Euclidean distance d between them as shown in the following equation (1).
[0166] d=√((12-20) 2 +(18-22) 2 +(18-20) 2 +(23-17) 2 )=10.954 ... (1)
[0167] Next, the similarity calculation unit 152 calculates the standardized Levenshtein distance from the lyric data of the segment PART1 and the segment PART2.
[0168] Here, the Levenshtein distance is a distance that is counted so that the distance increases by one each time an edit, insertion, deletion, or substitution is performed.
[0169] When calculating the Levenshtein distance, for example, "Tanakadesu" and "Nakadadesu" match when the "ta" is deleted and the "da" is inserted as the third character, so two operations are required and the distance is calculated as 2.
[0170] However, in this case, the importance of a difference between two characters in a sentence of 10,000 characters is different from that of a difference between two characters in a sentence of 10 characters. Therefore, the similarity calculation unit 152 further standardizes the edit distance by dividing it by the length of the longer standardized string, thereby obtaining the standardized Levenshtein distance.
[0171] For example, if lyric 1, "I'll be with you tomorrow morning," is the lyric of segment PART1, and lyric 2, "The stars shine in the night sky," is the lyric of segment PART2, the Levenshtein distance r is 13. Note that lyric 1 and 2 have 17 and 18 characters, respectively, including spaces.
[0172] Note that, although the number of moras in lyrics 1 and 2 differs from that used when calculating the Euclidean distance d by referring to the above-mentioned formula (1), they will be used here as an example of the Levenshtein distance for the purpose of explanation.
[0173] Therefore, in this case, the standardized Levenshtein distance sr is expressed by the following equation (2).
[0174] Standardized Levenshtein distance sr = Levenshtein distance r / max(length of lyric 1 (17), length of lyric 2 (18)) = 13 / 18 = 0.72222 ... (2)
[0175] Furthermore, the similarity calculation unit 152 calculates the standardized Levenshtein distance not only for all lyrics but also for the beginning and end of lyrics.
[0176] For example, lyrics 1 and 2 begin with "Ashita" and "Hoshi," respectively, requiring two operations: replacing "A" with "Ho" and deleting "Ta." Therefore, the Levenshtein distance rh is 2, and the maximum value of the string is 3. Therefore, the standardized Levenshtein distance srh of the beginning of the lyrics is calculated as 2 / 3 = 0.6666.
[0177] Furthermore, for lyrics 1 and 2, the endings are "aeru" and "mau," respectively, which require three operations: replacing "a" with "ma," replacing "e" with "u," and deleting "ru." Therefore, the Levenshtein distance er is 3, and since the maximum value of a character string is 3, the standardized Levenshtein distance sre at the end of the lyrics is calculated as 3 / 3 = 1.
[0178] The similarity calculation unit 152 calculates the standardized edit distance sed between the segment PART1 and the segment PART2 as the sum of the calculated Euclidean distance d and the standardized Levenshtein distances sr, srh, and sre for the entire lyrics, the beginning of the lyrics, and the end of the lyrics, respectively, as shown in the following equation (3).
[0179] Standardized edit distance sed = Euclidean distance d + whole lyrics sr + beginning lyrics srh + end lyrics sre = 10.954 + 0.7222 + 0.66666 + 1 = 13.34286 ... (3)
[0180] The similarity calculation unit 152 calculates the standardized edit distance sed by the above-mentioned calculation in a round-robin manner for the segments PART1 to PART3 and PART5 to PART10. As a result, for example, a distance matrix such as that shown in Fig. 19 is obtained. In Fig. 19, the standardized edit distance sed between the segments PART1 and PART2 is set to 14.
[0181] In addition, in Figure 19, a gradation is applied according to the value of the square, and the smaller the standardized edit distance sed and the higher the similarity, the darker the color is, and the minimum distance of 0 is set to black.
[0182] Next, the part estimation unit 153 estimates segments based on the standardized edit distance matrix calculated by the similarity calculation unit 152. The estimated segments include, for example, the A verse, B verse, C verse and subsequent parts, and the chorus.
[0183] Here, the estimation of the chorus by the part estimation unit 153 will be described.
[0184] The chorus, literally meaning "chorus" with more voices and grander, fuller parts, is the climax of the song, and is the part that makes the grandest statement both lyrically and musically in the entire song.
[0185] From a lyrical perspective, the chorus tends to repeat the same lyrics, so it is sometimes called the refrain, and can be said to be the part that unifies (brings together) the entire song.
[0186] As such, lyrics that are likely to be the chorus tend to be repeated, so the similarity calculation unit 152 estimates the segment that will be the chorus part by searching for clusters formed by segments whose standardized edit distance is smaller than a predetermined threshold.
[0187] Hereinafter, the threshold of the standardized edit distance at which segments can be considered to constitute a chorus will be referred to as the minimum distance threshold, and a cluster formed by segments whose standardized edit distance is smaller than the minimum distance threshold, i.e., segments that constitute a chorus, will also be referred to as the minimum distance cluster.
[0188] Therefore, the part estimation unit 153 sets a minimum distance threshold based on the standardized edit distance matrix, and identifies segments that form a minimum distance cluster.
[0189] More specifically, the similarity calculation unit 152 first sorts the values of each element in ascending order, removes duplicates, sets the values as reference values in ascending order, and repeatedly compares them with the next largest value to the reference value. When the next largest value increases by 30% or more compared to the reference value, the similarity calculation unit 152 sets the reference value as the minimum distance threshold.
[0190] However, since it is assumed here that each song contains at least three parts: an A melody, a B melody, and a chorus, the number of choruses is assumed to be, for example, the number of segments divided by 3, and this is set as the upper limit of the number of attempts.
[0191] Here, for example, a case will be considered in which the data are sorted in ascending order of standardized edit distance based on the standardized edit distance matrix, as shown in FIG.
[0192] In FIG. 20, the values are sorted in the following order from top to bottom: 0.0, 1.0159574468085106, 3.823421366714802, 8.94326015681513, 10.8590811369678697, 12.660556883514321, 13.663852859505031, 14.186263883000338, 14.732952603483653, and 15.339716594401313.
[0193] In the first trial, the part estimation unit 153 sets the first value, 0.0, as the reference value and compares it with the next value, 1.0159574468085106, to determine whether there is an increase of 30% or more from the reference value. In this case, since the reference value is 0.0, it is assumed that there is no increase of 30% or more.
[0194] Next, in a second trial, the part estimation unit 153 compares the second value, 1.0159574468085106, as a reference value with the next third value, 3.823421366714802, to determine whether there is an increase of 30% or more. In this case, the third value is 3.76 times the second value, which is the reference value, and is therefore considered to be an increase of 30% or more.
[0195] Therefore, the part estimation unit 153 sets 1.0159574468085106, which is the reference value in the second trial, as the minimum distance threshold for searching for a segment that will become the minimum distance cluster.
[0196] When the part estimation unit 153 sets the minimum distance threshold, it counts up the number of other segments in the standardized edit distance matrix that have a distance smaller than the minimum distance threshold of 1.0159574468085106 as the number of other segments with high similarity, for each segment.
[0197] At this time, if the title of the song described with reference to FIG. 17 is not simple, the part estimation unit 153 also counts up the number of times the title is included in each segment.
[0198] 19, segments RART3, 7, and 10 have distances of 0 and 1 that are smaller than the minimum distance threshold of 1.0159574468085106, so the count is increased by 2. In other words, segments RART3, 7, and 10 have standardized edit distances that are smaller than the minimum distance threshold, and can be considered to be similar to each other.
[0199] Also, as shown in FIG. 17, in segments RART3, 7, and 10, the title name, which is characteristic of the chorus, is counted once each.
[0200] Therefore, if the title of the song is not simple, based on FIGS. 17 and 19, segments RART3, 7, and 10 each count up by one, for a total of three.
[0201] The part estimation unit 153 determines that a segment with a count-up value greater than 2, that is, a segment with many other segments with high similarities, is a chorus.
[0202] That is, in the examples of FIGS. 17 and 19, the count-up values of the segments RART3, 7, and 10 are all 3, which is greater than 2, and therefore all are considered to be choruses.
[0203] Next, the part estimation unit 153 identifies a segment that is not similar to any other segment other than the three types of segments that are assumed to be included in a song, namely the A melody, B melody, and chorus, for example, a segment corresponding to a part below the C melody, as a segment that has an outlier value for the minimum distance.
[0204] That is, the part estimation unit 153 extracts the minimum value of the minimum distance in each segment, sets values equal to or greater than the third quartile multiplied by 1.5 times the interquartile range as outliers, and determines that the part is below the C melody.
[0205] More specifically, consider the case where the minimum values of the segments PART1 to PART3 and PART5 to PART10 are, in order from top to bottom, 8.94326015681513, 3.823421366714802, 0.0, 8.94326015681513, 3.823421366714802, 1.0159574468085106, 35.85475721493387, 14.186263883000338, and 0.0, as shown in FIG.
[0206] In the case of Figure 21, rearranging the minimum distances in ascending order results in 0.0 (3), 0.0 (10), 1.0159574468085106 (7), 3.823421366714802 (2), 3.823421366714802 (6), 8.94326015681513 (1), 8.94326015681513 (5), 14.186263883000338 (9), and 35.85475721493387 (8). Here, the numbers in parentheses are segment numbers.
[0207] In this case, as shown in Figure 22, the first quartile Q1 is 1.0159574468085106(7) and the third quartile Q3 is 8.94326015681513(5). Therefore, the interquartile range (IQR) is Q3-Q1 = 7.92730271.
[0208] As a result, the maximum value within the range (the maximum value that is not an outlier) is the third quartile Q3 + 1.5 x interquartile range (IRQ), which is 20.83421 (= 8.94326015681513 + 7.92730271 x 1.5).
[0209] Therefore, the minimum distance of segment PART8, 35.85475721493387, is greater than the maximum non-outlier value of 20.83421, and is therefore considered an outlier. For this reason, segment PART8 is estimated to be a part below the C melody.
[0210] In addition, in Figure 22, a definition is given for the minimum value that falls within the range (the minimum value that does not become an outlier), but in the present disclosure, 0 is the minimum value and also the minimum distance that is classified as a chorus, so it does not need to be taken into consideration.
[0211] Next, the part estimation unit 153 estimates the same part based on the distance for each segment that constitutes the standardized edit distance matrix.
[0212] More specifically, the part estimation unit 153 sorts the distances of each segment in ascending order, sets reference values in ascending order, and repeatedly compares the reference values with the next largest value. When the next largest value increases by a predetermined amount or more from the reference value, the part estimation unit 153 extracts the segment with the smallest distance to the reference value as a candidate for the same part. In this case, the part estimation unit 153 may extract the segment with the smallest distance to the reference value as a candidate for the same part when the next largest value increases by 20% or more from the reference value.
[0213] For example, consider the segment PART1 that makes up the top row in the normalized edit distance matrix of FIG. 19, with the minimum distances sorted as shown in the upper left corner of FIG.
[0214] When the values of segment PART1 constituting the top row in the standardized edit distance matrix of FIG. 19 are sorted, the results are, in order from smallest to largest, 8.94326015681513, 13.663852859505031, 15.339716594401313, 20.20498993731372, 28.062712448267185, 28.062712448267185, 28.399528992617242, and 44.332399120953454, as shown in the upper left corner of FIG. 23.
[0215] In this case, in the first trial, the part estimation unit 153 sets the first value, 8.94326015681513, as the reference value and compares it with the next largest value, 13.663852859505031, to determine whether there is an increase of 20% or more from the reference value. In this case, the reference value is 8.94326015681513, and the next value, 13.663852859505031, is 1.53 times the reference value, so it is considered to be an increase of 20% or more.
[0216] Therefore, the part estimation unit 153 sets the segment with the smallest distance to the reference value of 8.94326015681513 in the first trial as a candidate for the same part.
[0217] As shown in the right part of FIG. 23, the part estimation unit 153 determines that the segment PART5, which has a distance of 8.9 that is smaller than 8.94326015681513, is a candidate for the same part from the segment PART1 in the standardized edit distance matrix.
[0218] The part estimation unit 153 performs the same process on all segments to set identical part candidates.
[0219] As a result, the following description will proceed assuming that identical part candidates are set as shown in the lower left part of FIG.
[0220] In addition, in the lower left corner of Figure 23, it is shown that the same part candidate from the perspective of segment PART1 is segment PART5, the same part candidate from the perspective of segment PART2 is segment PART6, the same part candidate from the perspective of segment PART3 is segments PART7 and 10, the same part candidate from the perspective of segment PART5 is segments PART1, 2 and 6, the same part candidate from the perspective of segment PART6 is segment PART2, the same part candidate from the perspective of segment PART7 is segments PART3 and 10, there are no same part candidates from the perspective of segment PART8, the same part candidate from the perspective of segment PART9 is segments PART2 and 6, and the same part candidate from the perspective of segment PART10 is segments PART3 and 7.
[0221] The right part of Fig. 24 summarizes the relationships shown in the lower left of Fig. 23, where the numbers in the circles indicate segment numbers and the end points of the arrows represent segments that are considered to be candidates for the same part from the segment that is the starting point of the arrow. Note that the left part of Fig. 24 is the same as the lower left part of Fig. 23.
[0222] 24, segment PART5 is a candidate for the same part from the perspective of segment PART1, and at the same time, segment PART1 is also a candidate for the same part from the perspective of segment PART5, as indicated by the bidirectional arrows. The same is true for segments PART2 and PART6, PART3 and PART7, PART7 and PART10, and PART3 and PART10.
[0223] Also, from the perspective of segment PART5, segments PART2 and PART6 are both candidates for the same part, but from the perspective of segments PART2 and PART6, segment PART5 is not a candidate for the same part, so the arrow is unidirectional, from segment PART5 to segments PART2 and PART6.
[0224] Similarly, from the perspective of segment PART9, segments PART2 and PART6 are both candidates for the same part, but from the perspective of segments PART2 and PART6, segment PART9 is not a candidate for the same part, so the arrow is unidirectional, from segment PART9 to segments PART2 and PART6.
[0225] Furthermore, segment PART8 is expressed as a standalone segment since there are no identical part candidates in any of the segments.
[0226] The part estimation unit 153 groups segments with bidirectional arrows as being in the same part. In this case, although segment PART9 is in one direction with respect to segments PART2 and PART6 in Fig. 24, segments PART2 and PART6 are in two directions, so segment PART9, which recognizes both segments PART2 and PART6 as being in the same part, is also estimated to be in the same part as segments PART2 and PART6.
[0227] As a result, the part estimation unit 153 places segments PART1 and 5 in the same part group, segments PART2, 6, and 9 in the same part group, segments PART3, 7, and 10 in the same part group, and segment PART8 in a single group.
[0228] Then, the part estimation unit 153 checks the segments of each group based on the part information estimated by another method, and identifies the part.
[0229] That is, as described above, the part estimation unit 153 estimates that segments PART3, 7, and 10 are chorus parts by using a minimum distance threshold based on a standardized edit distance matrix, and therefore identifies segments PART3, 7, and 10 as chorus parts here.
[0230] Furthermore, the part estimation unit 153 assigns the remaining groups to the A verse and the B verse in ascending order of segment numbers. For example, if the minimum distance between segments PART1 and 5 is smaller than the minimum distance between segments PART2, 6, and 9, the part estimation unit 153 may assign segments PART1 and 5, which include the smallest segment number, PART1, to the A verse, and assign segments PART2, 6, and 9, which have lower segment numbers, to the B verse. Furthermore, the part estimation unit 153 may assign the single outlier segment PART8 to the C verse. Note that if there are multiple single outlier segments, the part estimation unit 153 may assign them in ascending order of segment numbers within the range from the C verse to the F verse.
[0231] Then, for segments without lyrics identified when the music composition features are extracted by the music composition feature extraction unit 151, the part estimation unit 153 identifies each part, such as segment PART0 with the smallest segment PART number as the intro, segment PART11 with the largest segment PART number as the outro, and other interludes.
[0232] By combining these, the part estimation unit 153 estimates segment PART0 as the intro, segment PART1 as the A melody, segment PART2 as the B melody, segment PART3 as the chorus, segment PART4 as the interlude, segment PART5 as the A melody, segment PART6 as the B melody, segment PART7 as the chorus, segment PART8 as the C melody, segment PART9 as the B melody, segment PART10 as the chorus, and segment PART11 as the outro, and outputs these as the final estimation results.
[0233] <Configuration Example of Beat Analysis Unit> Next, a configuration example of the beat analysis unit 118 will be described with reference to FIG.
[0234] The beat analysis unit 118 includes a beat analysis feature extraction unit 171 , a beat belonging probability generation unit 172 , and a beat estimation unit 173 .
[0235] The beat analysis feature extraction unit 171 extracts beat analysis features based on the acoustic features supplied from the acoustic feature extraction unit 117 .
[0236] The beat analysis features referred to here may include, for example, danceability, energy, instrumentalness, key, liveness, loudness, mode, speechiness, tempo, time signature, type, and valence.
[0237] The beat analysis features may also include, for example, the number of audio samples (num_samples), the offset to the start position of the duration (offset_seconds), the sample rate (analysis_sample_rate), the number of analysis channels (analysis_channels), the end time of the fade-in period (in seconds) (end_of_fade_in), and the start time of the fade-out period (in seconds) (start_of_fade_out).
[0238] Additionally, the beat analysis features may include loudness (in decibels (dB)) (loudness), tempo (tempo), tempo confidence (tempo_confidence), estimated time signature (time_signature), estimated time signature confidence (time_signature_confidence), key (key), key confidence (key_confidence), major or minor mode (mode), mode confidence (mode_confidence), chord string (codestring), chord string version (code_version), echo print string (echoprintstring), echo print string version (echoprint_version), sync string (synchstring), sync string version (synch_version), rhythm string (Rhythmstring), and rhythm string version (rhythm_version).
[0239] The beat analysis features may also include the time interval (bar) of the bars of the entire track (including the start point, duration, and duration confidence), beats (including the start point, duration, and duration confidence), sections (including the start point, duration, and duration confidence, loudness, tempo, tempo confidence (tempo_confidence), key, key confidence (key_confidence), mode, mode confidence (mode_confidence), estimated time signature (time_signature), and estimated time signature confidence (time_signature_confidence)).
[0240] Additionally, the beat analysis features may include segment (start, duration), segmentation confidence (duration), loudness start level (loudness_start), peak loudness level (loudness_max), peak loudness time offset time (loudness_max_time), loudness end level (loudness_end), pitch, timbre, and tatums (the minimum time interval between successive notes in a rhythmic phrase, the lowest regular pulse train that a listener perceives and intuitively infers).
[0241] The beat analysis feature extraction unit 171 may apply existing provided technology, for example, Get Track's Audio Features (registered trademark) and Get Track's Audio Analysis (registered trademark).
[0242] Furthermore, the beat analysis feature extraction unit 171 may calculate, for example, the average, standard deviation, maximum, minimum, and maximum difference as features using the extracted beat analysis features as a unit of a predetermined period. In this case, since a feature such as a categorical variable, for example, a genre, takes the form of "Pop" or "Rock," it may be quantified as a dummy variable, such as 1 if it applies and 0 if it does not apply.
[0243] The beat affiliation probability generation unit 172 estimates the affiliation probability of each predetermined type of beat based on the beat analysis feature supplied from the beat analysis feature extraction unit 171, and outputs the estimated probability to the beat estimation unit 173. The predetermined types of beats may be, for example, seven types of beats, including four-on-the-floor, eight-beat, sixteen-beat, double beat (two-beat), shuffle beat, duple beat, and triple beat, as shown in Fig. 26. Note that duple beat and triple beat are originally concepts different from beats, but in the present disclosure, they are treated as beats because they can be distinguished from each other.
[0244] Figure 26 shows examples of the following belonging probabilities based on beat analysis features: 4 beats, 8 beats, 16 beats, double beat (2 beats), shuffle beat, duple beat, and triple beat, with probabilities of 7.9%, 37.5%, 30.1%, 5.3%, 66.3%, 23.1%, and 52.6%, respectively.
[0245] The beat affiliation probability generation unit 172 is, for example, a recognition model (estimation model) that analyzes the type of beat based on the beat analysis features described above, and is formed by machine learning based on learning data that pairs predefined beat correct flags with songs.
[0246] Here, learning of the recognition model that constitutes the beat belonging probability generation unit 172 will be described.
[0247] First, to explain the learning of the beat belonging probability generation unit 172 consisting of a recognition model, we will explain the rough definitions of seven types of beats: four-on-the-floor, eight-beat, sixteen-beat, double beat (two-beat), shuffle beat, two-beat, and three-beat.
[0248] A four-on-the-floor beat is a beat in which the bass drum (kick) is struck four times in one measure, as shown in the upper left section of Fig. 27. In Fig. 27, the timing of the hi-hat, snare drum, and bass drum (kick) strikes for each beat per measure is shown, with one square representing a sixteenth note.
[0249] An 8-beat is a beat in which the hi-hat is struck on eighth notes within one measure, or the snare drum is struck on the backbeats (the second and fourth beats), as shown in the upper right corner of Figure 27.
[0250] Also, a 16th beat is a beat in which the hi-hat is struck in sixteenth notes within one measure, as shown in the middle left of Figure 27.
[0251] Furthermore, a double beat (double time feel) (2 beats) is a beat in which the snare drum is struck four times on the offbeat, as shown in the middle right of FIG.
[0252] The shuffle beat is a beat based on triplets with the middle note omitted, as shown in the lower left of FIG.
[0253] Furthermore, triple time is a rhythm in which three beats form one unit of beat, as shown in the lower right of Figure 27, but the concept of dynamics, such as dynamics and weak / dynamics and weak, is introduced separately from the beat. Therefore, the concept of dynamics must be introduced when estimating triple time.
[0254] Similarly, duple time (not shown) is a rhythm that forms a beat with two beats as one unit, and like triple time, the concept of dynamics such as dynamics / dynamics is introduced.
[0255] However, the above-mentioned seven types of beats are only rough definitions, and there are many beats that have different targets, such as those based on the hi-hat, snare drum, or bass drum, etc. As a result, there may be beats that overlap and satisfy the definitions, making it impossible to distinguish them strictly, resulting in overlapping areas such as those shown in Figure 28.
[0256] For example, a four-on-the-floor beat is a beat in which the bass drum is struck four times, but a sixteenth beat is a beat in which the hi-hat is struck in sixteenth notes, so the two can coexist.
[0257] Therefore, the priority of beats that are easily perceived by humans is defined, and when multiple beats coexist, the beats are defined based on the priority. For example, as shown in the bottom of Figure 28, the priority may be as follows, in order of decreasing priority: 8 beats < 4 beats < double beats < 16 beats < shuffle < 2, 3 beats.
[0258] When constructing a recognition model that realizes the beat belonging probability generation unit 172 using machine learning, first, a plurality of songs are used so that learning pair data between beats that become correct flags and song data is generated for each beat.
[0259] Next, the generated training pair data is divided into training data and verification data for each beat, with the training data and verification data being divided at a predetermined ratio, for example, about 7:3.
[0260] Furthermore, the training data for each beat is combined with the training data for other beats. At this time, the training data for each beat may be combined with the training data for other beats so that the combined amount of the training data for each beat is approximately three times the amount of the training data for each beat. The remaining training pair data is combined with the verification data.
[0261] More specifically, for example, in the case of learning paired training data for four-on-the-floor beats, the paired training data for four-on-the-floor beats among the correct answer data for all songs is divided into training data and verification data at a predetermined ratio. The predetermined ratio may be, for example, a ratio of 7:3 as shown in FIG. 29 .
[0262] Furthermore, the divided four-on-the-floor learning data may be combined with a predetermined number of pairs of learning data for beats other than the four-on-the-floor beats, and the remaining data for the other beats may be combined with the verification data. As shown in Fig. 29, the divided four-on-the-floor learning data may be combined with a maximum of three times as many pairs of learning data for beats other than the four-on-the-floor beats as the divided four-on-the-floor learning data.
[0263] Then, through machine learning using the divided training data for the four-on-the-floor beats and training data for other beats to a predetermined extent, a recognition model that recognizes the four-on-the-floor beats is trained from among the models that make up the beat belonging probability generation unit 172.
[0264] This learning may be, for example, learning in which a predetermined number (e.g., about 1,500) of decision trees are set for the Random Forest Classifier algorithm provided in the scikit learn library.
[0265] Here, the random forest classification algorithm is an ensemble learning algorithm that combines multiple slightly different "decision trees."
[0266] By repeating the above-mentioned learning and verification using verification data for each beat, a beat affiliation probability generation unit 172 is formed as a recognition model for generating the affiliation probability of each beat based on the beat analysis features.
[0267] The beat estimation unit 173 estimates a beat based on the belonging probability for each beat supplied from the beat belonging probability generation unit 172, and outputs the estimation result.
[0268] That is, as shown in FIG. 30 , if the respective belonging probabilities of four beats, eight beats, sixteen beats, double beats (two beats), shuffle beats, duple beats, and triple beats supplied from the beat belonging probability generation unit 172 described above are, for example, 7.9%, 37.5%, 30.1%, 5.3%, 66.3%, 23.1%, and 52.6%, the beat estimation unit 173 estimates the beat of the song to be analyzed as a shuffle beat.
[0269] The beat estimation unit 173 is, for example, a recognition model (estimation model) that estimates beats based on the beat affiliation probability supplied from the beat affiliation probability generation unit 172, and is formed by machine learning based on learning pair data that pairs beat correct flags with the above-mentioned beat affiliation probabilities from multiple songs.
[0270] Here, learning of the recognition model that constitutes the beat estimation unit 173 will be described.
[0271] When constructing a recognition model that realizes the beat estimation unit 173 using machine learning, first, using a plurality of songs, learning pair data is generated between beats that are flagged as correct and beat affiliation probabilities generated by the beat affiliation probability generation unit 172 described above.
[0272] Next, the generated training pair data is divided into training data and verification data for each beat. At this time, the training data and the verification data are divided at a predetermined ratio. For example, as shown in FIG. 31 , the training data and the verification data may be divided at a ratio of about 7:3.
[0273] Then, by machine learning using the learning data, the beat estimation unit 173 is trained as a recognition model that estimates beats from the relationship between each beat and the beat belonging probability.
[0274] As with the beat belonging probability generation unit 172, this learning may involve setting a predetermined number (e.g., approximately 1,500) of decision trees for the Random Forest Classifier algorithm provided in the scikit learn library.
[0275] By repeating the above-mentioned learning based on the relationship between each beat and the beat belonging probability and verification using verification data, a beat estimation unit 173 is formed as a recognition model that estimates the beat of a song based on the beat belonging probability.
[0276] <Music Analysis Processing> Next, the music analysis processing will be described with reference to the flowchart of FIG.
[0277] In step S31, the input unit 111 receives an input from the user and determines whether or not at least one of the artist name and song name (song title) of the song to be analyzed, or the sound source data of the song, has been input. That is, in step S31, it is determined whether or not any information identifying the song to be analyzed has been input.
[0278] If it is determined in step S31 that at least one of the artist name and song name (song title) of the song to be analyzed and the sound source data of the song has been input, the process proceeds to step S32.
[0279] In step S32, the input unit 111 accepts input of at least one of the artist name and song name (song title) of the song to be analyzed and the sound source data of the song, and supplies the input to the data acquisition unit 112. At this time, when the input unit 111 accepts input of sound source data, for example, it may control the identification unit 111a to analyze the sound source data and identify the artist name and song title of the song, and supply the identification result to the data acquisition unit 112 as the artist name and song title of the song whose input was accepted.
[0280] In step S33, the data acquisition unit 112 accesses the music DB 102, acquires the chord data of the music, and supplies it to the parallel modulation estimation unit 113. If the chord data cannot be acquired, the data acquisition unit 112 notifies the parallel modulation estimation unit 113 that the chord data cannot be acquired.
[0281] In step S34, the parallel modulation estimator 113 determines whether or not the chord data has been acquired based on whether or not the chord data has been supplied from the data acquirer 112.
[0282] If it is determined in step S34 that the code data has been acquired, the process proceeds to step S35.
[0283] In step S35, the parallel key modulation estimator 113 and the chord analyzer 114 perform chord analysis processing and supply the analysis results to the UI generator 119. Details of the chord analysis processing will be described later with reference to the flowchart of FIG.
[0284] If the chord data cannot be acquired in step S34, the chord analysis process in step S35 is skipped.
[0285] In step S36, the data acquisition unit 112 accesses the music DB 102, acquires lyric data of the music, and supplies the acquired lyric data to the lyric extraction unit 115. If the lyric data cannot be acquired, the data acquisition unit 112 may notify the lyric extraction unit 115 that the lyric data cannot be acquired.
[0286] In step S37, the lyrics extraction unit 115 determines whether lyrics data has been acquired based on whether lyrics data has been supplied from the data acquisition unit 112 or not.
[0287] If it is determined in step S37 that the lyrics data has been acquired, the process proceeds to step S38.
[0288] In step S38, the lyrics extraction unit 115 and music structure analysis unit 116 execute music structure analysis processing and supply the analysis results to the UI generation unit 119. Details of the music structure analysis processing will be described later with reference to the flowchart in Fig. 34. At this time, the lyrics extraction unit 115 also outputs the acquired lyrics data to the chord analysis unit 114.
[0289] If the lyrics data cannot be acquired in step S37, the music composition analysis process in step S38 is skipped.
[0290] In step S39 , the data acquisition unit 112 accesses the music DB 102 to acquire the sound source data of the music, and supplies the data to the acoustic feature extraction unit 117 .
[0291] If the sound source data cannot be acquired, the data acquisition unit 112 notifies the acoustic feature extraction unit 117 that the sound source data cannot be acquired.
[0292] Furthermore, when the input unit 111 accepts input of sound source data, the data acquisition unit 112 may supply the accepted input sound source data as is to the acoustic feature extraction unit 117 .
[0293] In step S40, the acoustic feature extraction unit 117 determines whether or not the sound source data has been acquired based on whether or not the sound source data has been supplied from the data acquisition unit 112.
[0294] If it is determined in step S40 that the sound source data has been acquired, the process proceeds to step S41.
[0295] In step S41, the acoustic feature extraction unit 117 and the beat analysis unit 118 execute beat analysis processing and supply the analysis results to the UI generation unit 119. Details of the beat processing will be described later with reference to the flowchart in FIG.
[0296] If the sound source data cannot be acquired in step S40, the beat analysis process in step S41 is skipped.
[0297] In step S 42 , the UI generation unit 119 generates a UI image based on the analysis results supplied from the chord analysis unit 114 , the music structure analysis unit 116 , and the beat analysis unit 118 , and displays the UI image on the display unit 104 .
[0298] Examples of UI images generated by the UI generation unit 119 will be described in detail later with reference to Figures 36 to 40. The UI images are based on the premise that at least one of chord data, lyric data, and sound source data can be acquired. If none of the data can be acquired, the UI images cannot be generated, and the process of generating and presenting the UI images is skipped.
[0299] Through the above series of processes, it is possible to generate and present UI images for the chords, composition, and beats of a piece of music whose input has been accepted by the user operating the input unit 111 .
[0300] This allows for proper analysis of the chords, structure, and beat of a song, making it possible to analyze the characteristics of each fan or customer, which can then be used to develop and implement music production and marketing strategies.
[0301] In the flowchart of FIG. 32, the three processing blocks consisting of the processing block of steps S33 to S35, the processing block of steps S36 to S38, and the processing block of steps S39 to S41 are represented as being realized in series, but the order of these three processing blocks may be changed, or each may be processed in parallel.
[0302] In addition, the following description will assume that the processing of the three processing blocks performed in series in the flowchart of Figure 32 may be performed in a different order or in parallel.
[0303] <Code Analysis Processing> Next, the code analysis processing will be described with reference to the flowchart of FIG.
[0304] In step S51, the parallel key modulation estimator 113 shapes the chord data. For example, the parallel key modulation estimator 113 may shape the chord data by shaping notation fluctuations in the chord data supplied from the data acquisition unit 112, setting segments based on line break information, and setting an intro part, interlude, and outro part depending on whether lyrics are present.
[0305] In step S52, the parallel modulation estimator 113 reads chord-related master information. Here, the parallel modulation estimator 113 may, for example, request the chord-related master information from the data acquirer 112. In response, the data acquirer 112 may access the chord-related DB 103, acquire the chord-related master information, and supply it to the parallel modulation estimator 113. The parallel modulation estimator 113 may then acquire the chord-related master information supplied from the data acquirer 112.
[0306] In step S53, the parallel key modulation estimator 113 estimates the key (parallel key) and modulation. At this time, the parallel key modulation estimator 113 may, for example, sequentially scan the chord starting from the 0th chord and, as described with reference to FIG. 11 , estimate the key based on the relative positions of the chords and the number of diatonic chords that appear, and output the key to the chord analyzer 114. At this time, if a chord that deviates from the key estimated up to the middle chord appears, the parallel key modulation estimator 113 may, for example, estimate whether the key is a modulation or a borrowed chord, taking into account the length of the cadence. In the case of a modulation, the position of the modulation and the key after modulation may be estimated using a similar method.
[0307] In step S54, the chord progression estimation unit 131 of the chord analysis unit 114 estimates chord progression patterns and outputs them to the analysis result output unit 134. At this time, the chord progression estimation unit 131 may take into consideration the key, modulation, segments, etc., as described with reference to Fig. 12, and then compare the estimated chord progression patterns with the master information.
[0308] In step S55, the key estimation unit 132 estimates the key and outputs it to the analysis result output unit 134. At this time, the key estimation unit 132 may estimate whether the key of the piece of music to be analyzed is major or minor based on, for example, the chord data and lyric data, such as the first chord, the chord at the beginning of the lyrics, the chord at the end of the lyrics, and the final chord, the key, and chord progression pattern information.
[0309] In step S56, the chord group function estimation unit 133 estimates a chord group and a chord function for each chord based on the estimated relative key, and outputs the results to the analysis result output unit 134. The chord group function estimation unit 133 may estimate a chord group and a chord function for each chord based on the estimated relative key, as described with reference to Figs. 13 and 14, for example.
[0310] That is, the chord group function estimation unit 133 may estimate, for each chord, whether the chord group is, for example, diatonic, secondary dominant, related two minor, or subdominant minor. Furthermore, the chord group function estimation unit 133 may estimate, for each chord, whether the chord function is, for example, tonic, subdominant, or dominant.
[0311] In step S57 , the analysis result output unit 134 formats the estimation result based on the code data and outputs it to the UI generation unit 119 .
[0312] That is, the analysis result output unit 134 may estimate and output the final key determined from the key (parallel key) and key estimation results, for example.
[0313] Furthermore, the analysis result output unit 134 may output, for example, the timing of the key change, the number of changes, and information on the key after the key change as key change information.
[0314] Furthermore, the analysis result output unit 134 may output information on the chord progression pattern, such as the number of times each chord progression pattern appears, the representative chord progression pattern for each segment, and the chord progression patterns for the original chords, the chords converted to C key, and the simplified chords (simplified chords excluding tension chords, etc.).
[0315] Additionally, the analysis result output unit 134 may also output information on the complexity of the song calculated from, for example, the number of occurrences of each chord, the number of occurrences of chord groups and chord functions, the number of borrowed chords used, the number of key signatures, the degree to which tension chords are incorporated, the presence or absence of key modulation, etc.
[0316] Through the above process, various pieces of information relating to chords are analyzed and output based on the chord data of the piece of music to be analyzed.
[0317] <Music Structure Analysis Processing> Next, the music structure analysis processing will be described with reference to the flowchart of FIG.
[0318] In step S71, the music composition feature extraction unit 151 shapes the lyric data. For example, the music composition feature extraction unit 151 may determine segments based on line break information and determine whether the segments contain lyrics, such as intonation, interlude, or outro, and shape the lyric data by dividing and storing the lyric data for each segment by line.
[0319] In step S72, the music composition feature extraction unit 151 calculates the degree of simplicity of the music title. For example, the music composition feature extraction unit 151 may calculate the degree of simplicity of the music title by excluding additional information (such as xxx.ver) from the music title, performing morphological analysis, and converting it into katakana.
[0320] In step S73, the music structure feature extraction unit 151 generates music structure features for each row for each segment, and outputs them to the similarity calculation unit 152. Here, for example, as described with reference to Figures 16 to 18, the music structure feature extraction unit 151 may calculate the number of moras (a unit of phonetic length) for each row for each segment, excluding intros, interludes, and outros that do not have lyrics, and create a matrix (convert to a matrix) for each segment, and further generate music structure features by applying patterning using 0s to make the number of elements uniform.
[0321] In step S74, the similarity calculation unit 152 generates a standardized edit distance matrix. At this time, the similarity calculation unit 152 may calculate the standardized edit distance as the similarity in a round-robin manner for each segment from the lyric data of segments PART1 and PART2, for example, and generate a standardized edit distance matrix as described with reference to FIG. 19 .
[0322] In step S75, the part estimation unit 153 estimates segments that form the same part based on the standardized edit distance matrix. At this time, the part estimation unit 153 may estimate segments that form the same part based on the standardized edit distance matrix, for example, as described with reference to Figures 20 to 23.
[0323] In step S76, the part estimation unit 153 estimates the part for each segment with respect to the same part and the relationship between the segments, and outputs the estimation result to the UI generation unit 119. At this time, the part estimation unit 153 may estimate the part for each segment with respect to the same part and the relationship between the segments, as described with reference to FIG.
[0324] Through the above processing, identical parts are estimated based on the standardized edit distance, which is based on the standardized edit distance matrix, for each segment, and based on the relationship between the segments in the estimated identical part, it is possible to estimate whether the part type is an intro, interlude, chorus, A melody, B melody, C melody, or outro.
[0325] This makes it possible to estimate with high accuracy the types of parts that make up a piece of music, segment by segment.
[0326] As a result, it will be possible to analyze the characteristics of each of the users, fans and customers, and this information can be used for music production, the planning and implementation of marketing strategies, and more.
[0327] <Beat Analysis Processing> Next, the beat analysis processing will be described with reference to the flowchart in FIG.
[0328] In step S91, the acoustic feature extraction unit 117 extracts acoustic features based on the sound source data of the music piece to be analyzed, and supplies the extracted acoustic features to the beat analysis unit 118. At this time, the acoustic feature extraction unit 117 may extract acoustic features including, for example, basic features, instruments, moods (higher and lower orders), and genres based on the sound source data of the music piece to be analyzed.
[0329] In step S92 , the beat analysis feature extraction unit 171 extracts beat analysis features based on the acoustic features supplied from the acoustic feature extraction unit 117 .
[0330] In step S93, the beat affiliation probability generation unit 172 estimates the affiliation probability of each of predetermined types of beats based on the beat analysis feature supplied from the beat analysis feature extraction unit 171, and outputs the estimate to the beat estimation unit 173. At this time, the beat affiliation probability generation unit 172 may, for example, estimate the affiliation probability of each of seven types of beats, namely, four-on-the-floor, eight-beat, sixteen-beat, double beat (two-beat), shuffle beat, duple beat, and triple beat, based on the beat analysis feature supplied from the beat analysis feature extraction unit 171, as described with reference to Fig. 26 .
[0331] In step S94 , the beat estimation unit 173 estimates a beat based on the belonging probability for each beat supplied from the beat belonging probability generation unit 172 , and outputs the estimation result to the UI generation unit 119 .
[0332] Through the above processing, it is possible to estimate the beat of the music piece to be analyzed with high accuracy.
[0333] <UI image example (part 1)> Next, referring to Figure 36, we will explain an example of a UI image (part 1) generated by the UI generation unit 119 based on the chord progression pattern estimation result from the chord analysis process, the music structure estimation result from the music structure analysis process, and the beat estimation result from the beat analysis process.
[0334] As shown in FIG. 36, the UI image 201 may be provided with an artist column 211, a music content column 212, a hit trend column 213, a bookmark list column 214, and an analysis result display column 215.
[0335] The artist field 211 is a field that is operated (pressed) when selecting the artist of the song to be analyzed. For example, although not shown in Fig. 36, when the artist field 211 is operated, a drop-down list of selectable artists is displayed, and the user may select the desired artist using a pointer or the like.
[0336] In FIG. 36, artist A has already been selected, and it is displayed that songs related to artist A are being analyzed.
[0337] Musical piece content field 212 is a field that is operated (pressed) when selecting musical piece content to be analyzed. For example, as shown in Fig. 36, the musical piece content field 212 is displayed in an operated state, and a drop-down list showing songs A to E as selectable musical piece content is displayed, and the field 221 for song A may be selected.
[0338] The hit trend column 213 is a column that is operated (pressed) when a UI image summarizing hit trends is desired to be displayed.
[0339] The bookmark list field 214 is a field that is operated (pressed) when a user desires to display a UI image that has been registered as a bookmark.
[0340] The analysis result display field 215 is a field where the analysis results specified by the user are displayed, and in FIG. 36, the analysis results for song A by artist A are displayed.
[0341] In the analysis result display field 215 of FIG. 36, "Song A" is written at the top, indicating that the title of the song to be analyzed is "Song A."
[0342] Below the title of the song, an impression display field 231 and a characteristic composition display field 232 are provided.
[0343] The impression display field 231 displays typical impressions of the song to be analyzed, and in Fig. 36, "electronic," "sophisticated," and "delicate" are displayed from left to right. For details on impression estimation, see Patent Document 1.
[0344] The characteristic composition display field 232 is a field in which either the music characteristics, lyrics characteristics, or music composition are displayed, and for example, as shown in FIG. 36, the display may be switched by selecting a tab labeled "music characteristics," "lyrics characteristics," or "music composition."
[0345] In FIG. 36, "music characteristics" is selected, and the music characteristics are displayed in the characteristic configuration display field 232.
[0346] As shown in FIG. 36 , the characteristic configuration display field 232 may be provided with a BPM field 241, a song length field (Length) 242, a beat field (Beat) 243, a chord length field (Chord Length) 244, a key field (Key) 245, a modulation field (Modulation) 246, a chord progression field 247, a chord variation field 248, and a chord group field 249.
[0347] The BPM column 241 is a column that indicates the BPM of the song A to be analyzed, and in FIG. 36, it is indicated as "130", indicating that the BPM of the song A is 130.
[0348] The song length column (Length) 242 is a column in which the length of song A to be analyzed is displayed. In FIG. 36, "4:20" is displayed, indicating that the length of the song is 4 minutes and 20 seconds.
[0349] The beat column (Beat) 243 is a column in which the type of beat of the song A to be analyzed is displayed. In FIG. 36, "four on the floor" is displayed, indicating that the beat is four on the floor.
[0350] The chord length column 244 is a column in which the number of chords in the piece of music A to be analyzed is written. In FIG. 36, it is displayed as "279", indicating that the number of chords is 279.
[0351] The key column (Key) 245 is a column in which the type of key of the song A to be analyzed is displayed. In FIG. 36, the keys are displayed in the order of "E♭, D, F" from left to right, indicating that the key changes in the order of "E♭, D, F."
[0352] The modulation column 246 is a column that displays the number of modulations and the degree of modulation of the piece of music A being analyzed. In FIG. 36, it displays "2 times (-1 → +3)", indicating that there are 2 modulations and that the degree changes by -1 degree and then by +3 degrees.
[0353] The chord progression column 247 is a column that displays, in a pie chart, the usage ratio of the chord progression patterns used in the piece of music A that is the subject of analysis.
[0354] Chord variation 248 is a field that displays the chord usage ratio used in song A, which is the analysis target, in the circle of fifths.
[0355] The chord affiliation group column 249 is a column that displays the usage ratio of the chord affiliation groups used in the piece of music A to be analyzed in the form of a pie chart.
[0356] The analysis results are displayed on the UI image 201 in Figure 36, which allows you to visually check the chord progression usage ratio, chord usage ratio, and chord group usage ratio of the song being analyzed, in addition to information related to the beat, number of chords, key, and modulation.
[0357] This makes it possible to properly understand the composition of a piece of music in terms of chord progressions, chord usage ratios, and chord group usage ratios when studying the piece of music to be analyzed.
[0358] As a result, for example, if various hit songs are selected and presented, it will be possible to read the trends of hit songs from the perspective of chord progressions, chord usage ratios, and chord group usage ratios, which can be used for music production and the planning and implementation of marketing strategies.
[0359] <UI image example (part 2)> In the above, an example has been shown in which the tab labeled "music characteristics" is selected in the characteristic composition display field 232, but the song composition may also be displayed in the characteristic composition display field 232 by selecting "music composition."
[0360] Fig. 37 shows a UI image 201A in which a tab labeled "Music Composition" is selected in the characteristic composition display field 232A to display the composition. Note that the analysis result display field 215A in Fig. 37 is similar to the analysis result display field 215A in Fig. 36 except that, when the tab labeled "Music Composition" is selected, the characteristic composition display field 232 in Fig. 36 is replaced with the characteristic composition display field 232A, and therefore a description thereof will be omitted.
[0361] That is, as shown in FIG. 37, the characteristic structure display field 232A may be provided with, from the top, a music structure field 271 and a chord progression field 272.
[0362] The music composition column 271 is a column that displays the composition parts of the music piece, arranged from left to right in the drawing along with the progression of the music piece.
[0363] In Figure 37, parts classified into units of the above-mentioned segments are displayed sequentially from left to right in the figure, and as shown in Figure 37, the three boxes from the left are blank, but to the right of them, "Interlude," "A," "B," "Chorus," etc. may be written, indicating that the time is represented by the horizontal length of each box in the figure, in the order of interlude, A melody, B melody, and chorus from left to right.
[0364] The chord progression column 272 is a column showing the chord progressions used in each part of the music composition in the music composition column 271.
[0365] As shown in Figure 37, in the upper left corner, "Chorus" is written, and below that, "A♭ B♭ Gm Cm" is written, and below that, "4536 (standard progression)" is written, indicating that the corresponding part of the song is the chorus, the chord progression pattern is "A♭ B♭ Gm Cm", the degree notation in the chord progression can be expressed in Arabic numerals as "4536", and the name of the corresponding chord progression pattern is "standard progression".
[0366] Similar notations are arranged from the top left of the figure to the right in part order, and the next notation on the right side is arranged one row lower, from the left side, again to the right, repeating this arrangement.
[0367] That is, it is possible to display the music composition and chord progression patterns as shown in FIG.
[0368] This makes it possible to properly understand the composition of a piece of music from the perspective of its structure and chord progression patterns when studying the piece of music to be analyzed.
[0369] As a result, for example, by selecting and presenting a variety of hit songs, it becomes possible to discern trends in hit songs from the perspective of song structure and chord progression patterns, which can be used for song production and the formulation and implementation of marketing strategies.
[0370] <UI Image Example (Part 3)> Although UI image examples using songs as units have been shown above, analysis results for multiple songs may be displayed in artist units.
[0371] 38, when the artist field 211 is selected, a UI image 201B including an analysis result display field 215B is displayed. As shown in FIG. 38, the analysis result display field 215B may be configured to display the analysis results for songs O, P, Q, and R of the selected artist B.
[0372] More specifically, the analysis result display field 215B may be provided with, for example, a chord progression usage ratio display field 291, a chord function usage ratio display field 292, a chord group usage ratio display field 293, a song key ratio display field 294, a title name display field (Title) 295, a key display field (key) 296, a BPM display field (BPM) 297, a song length display field (Length) 298, and a release date display field (Release Date) 299, as shown in FIG.
[0373] As shown in the chord progression usage ratio display column 291, the chord function usage ratio display column 292, and the chord group usage ratio display column 293, the chord progression usage ratio, chord function usage ratio, and chord group usage ratio for the songs with the titles displayed in the title name display column 295 among the songs of the artist being analyzed may be displayed in pie charts.
[0374] As shown in the chord progression usage ratio display column 291, chord function usage ratio display column 292, and chord group usage ratio display column 293 in Figure 38, the chord progression usage ratio, chord function usage ratio, and chord group usage ratio for songs O, P, Q, and R by artist B being analyzed may be displayed in pie charts.
[0375] The song key ratio display field 294 displays, in circle of fifths, the key ratio of the song with the title displayed in the title name display field 295 among the songs of the artist being analyzed.
[0376] As shown in the song key ratio display field 294 in FIG. 38, the key ratios of songs O, P, Q, and R by artist B to be analyzed may be displayed in the circle of fifths.
[0377] The title display field 295 displays the titles of the songs set as the analysis target among the songs of the artist to be analyzed.
[0378] As shown in the title name display field 295 of Figure 38, the titles are displayed from left to right as Song O, Song P, Song Q, and Song R, and it may be displayed that the titles of the songs set as the analysis targets among the songs of artist B to be analyzed are, from left to right, Song O, Song P, Song Q, and Song R.
[0379] The key display field 296 displays the key of the song set as the analysis target among the songs of the artist to be analyzed.
[0380] As shown in the key display field 296 of Figure 38, the keys are written as B, E♭, B♭, and B from left to right in the figure, and it may be displayed that the keys of songs O, P, Q, and R set as the analysis targets among the songs by artist B to be analyzed are B, E♭, B♭, and B, respectively.
[0381] The BPM display field 297 displays the BPM of the song set as the analysis target among the songs of the artist to be analyzed.
[0382] As shown in the BPM display field 297 of Figure 38, from left to right in the figure, the BPMs are written as 86, 91, 103, and 103, and it may be displayed that of the songs by artist B to be analyzed, the BPMs of songs O, P, Q, and R set as the analysis targets are 86, 91, 103, and 103, respectively.
[0383] The song length display field 298 displays the length (length of time) of songs set as the analysis target among songs by the artist to be analyzed.
[0384] As shown in the song length display field 298 in Figure 38, from left to right in the figure, the lengths are displayed as 4 minutes 15 seconds, 4 minutes 25 seconds, 3 minutes 22 seconds, and 4 minutes 24 seconds, and it may be displayed that the lengths of songs O, P, Q, and R set as the analysis targets among the songs by artist B to be analyzed are 4 minutes 15 seconds, 4 minutes 25 seconds, 3 minutes 22 seconds, and 4 minutes 24 seconds, respectively.
[0385] The release date display field 299 displays the release dates of the songs set as the analysis target among the songs of the artist to be analyzed.
[0386] As shown in the release date display field 299 of Figure 38, from left to right in the figure, the dates are written as 2018-03-14, 2019-09-11, 2020-02-03, and 2020-08-05, and it may be displayed that the release dates of songs O, P, Q, and R set as the analysis targets among the songs by artist B to be analyzed are March 14, 2018, September 11, 2019, February 3, 2020, and August 5, 2020, respectively.
[0387] That is, as shown in FIG. 39, it is possible to display multiple chord progressions, chord groups / functions, and key usage ratios for each artist.
[0388] This makes it possible to properly grasp the style and tendencies of each artist's musical composition in terms of chord progressions and chord groups / functions when studying the songs of the artist being analyzed.
[0389] As a result, for example, by presenting the style and trends that can be obtained from the musical composition of multiple songs by an artist who has produced many different hit songs, it becomes possible to read the style and trends that could become hit songs, and this can be used for music production and the planning and implementation of marketing strategies.
[0390] <UI Image Example (No. 4)> In the above, examples of UI images based on songs or artists have been shown, but it is also possible to display trend analysis results for hit songs by decade.
[0391] As shown in FIG. 39, when the hit trend field 213 is selected, a UI image 201C including a trend display field 321 may be displayed.
[0392] As shown in the trend display field 321 of FIG. 39, a hit trend graph display field 331 and a song beginning part (music composition) display field 332 may be provided on the left side of the drawing.
[0393] The hit trend graph display field 331 displays, in graph form, the average values of various feature quantities for multiple songs classified as hit songs by decade, broken down by feature quantity type. As shown in FIG. 39 , the feature quantities handled in the hit trend graph display field 331 may be, from top to bottom, length (temporal length), BPM, chord progression pattern, chord pattern, number of chords per second, danceability, intensity, brightness, sound pressure, and acoustic feel. As shown in FIG. 39 , each graph may be displayed as a solid line graph with the horizontal axis representing the decade, and the time series trends of the line graphs for each normalized feature quantity may be represented by dotted arrows.
[0394] As shown in the song start part display field 332 in Fig. 39, the average song structure of the song start parts of multiple songs classified as hit songs for each decade may be displayed with the start time at the top of the drawing and the elapsed time at the bottom. As shown in Fig. 39, the song start part display field 332 may display the changes in the proportions of the intro, verse (A), and chorus that form the beginning of the song so that they can be compared by decade.
[0395] That is, as shown in the trend display field 321 of FIG. 39, it is possible to display the change in the characteristics of hit songs by decade and the change in the structure of the beginning of songs by decade.
[0396] This makes it possible to understand the time-series changes in the trends of songs that become hits from changes in the characteristics of hit songs.
[0397] As a result, for example, it will be possible to understand the characteristics of current hit songs and predict the characteristics of future hit songs from the evolution of the characteristics of songs that have the potential to become hits, which can be used for song production and the planning and implementation of marketing strategies.
[0398] <UI image example (part 5)> In the above, we have explained an example of presenting the changes in song characteristics and song structure at the beginning of a song over time as a result of trend analysis of hit songs by decade, but it is also possible to present only the changes in specific characteristics of hit songs, for example, it is possible to display the changes in the beat of songs that become hits over time.
[0399] As shown in FIG. 40, when the hit trend field 213 is selected, a UI image 201D including a trend display field 341 may be displayed.
[0400] In FIG. 40, a line graph is displayed that shows the change in beat content of hit songs over time.
[0401] That is, as shown in the trend display field 341 of FIG. 40, the time series changes in the proportions of each of the seven types of beats in hit songs, consisting of four-on-the-floor, eight-beat, sixteen-beat, double beat (two-beat), shuffle beat, two-beat, and three-beat, may be represented by a line graph with the horizontal axis representing the era.
[0402] That is, as shown in FIG. 40, it is possible to display the chronological changes in beats in hit songs.
[0403] This makes it possible to appropriately grasp the time series changes in the type of beat that makes a hit song.
[0404] As a result, for example, it will be possible to understand the beats of current hit songs and predict the beats of future hit songs from the changes in the beats of songs that have the potential to become hits, which can be used for song production and the planning and implementation of marketing strategies.
[0405] The UI image 201D in FIG. 40 is merely an example in which beats are selected and displayed from among the various feature quantities that express music, and other feature quantities may be selected and displayed.
[0406] <<4. Example of Execution by Software>> The above-described series of processes can be executed by hardware, but can also be executed by software. When the series of processes is executed by software, the program that constitutes the software is installed from a recording medium into a computer that is built into dedicated hardware, or into, for example, a general-purpose computer that can execute various functions by installing various programs.
[0407] 41 shows an example of the configuration of a general-purpose computer. This computer has a built-in CPU (Central Processing Unit) 1001. An input / output interface 1005 is connected to the CPU 1001 via a bus 1004. A ROM (Read Only Memory) 1002 and a RAM (Random Access Memory) 1003 are connected to the bus 1004.
[0408] The input / output interface 1005 is connected to an input unit 1006 including input devices such as a keyboard and a mouse through which a user inputs operation commands, an output unit 1007 that outputs a processing operation screen and images of processing results to a display device, a storage unit 1008 including a hard disk drive or the like that stores programs and various data, and a communication unit 1009 including a LAN (Local Area Network) adapter or the like that executes communication processing via a network typified by the Internet. Also connected is a drive 1010 that reads and writes data from / to a removable storage medium 1011 such as a magnetic disk (including a flexible disk), an optical disk (including a CD-ROM (Compact Disc-Read Only Memory) and a DVD (Digital Versatile Disc)), a magneto-optical disk (including an MD (Mini Disc)), or a semiconductor memory.
[0409] The CPU 1001 executes various processes in accordance with a program stored in a ROM 1002 or a program read from a removable storage medium 1011 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, installed in a storage unit 1008, and loaded from the storage unit 1008 into a RAM 1003. The RAM 1003 also stores data necessary for the CPU 1001 to execute various processes as appropriate.
[0410] In a computer configured as described above, the CPU 1001 performs the above-described series of processes by, for example, loading a program stored in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.
[0411] The program executed by the computer (CPU 1001) can be provided by being recorded on a removable storage medium 1011 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0412] In a computer, a program can be installed in the storage unit 1008 via the input / output interface 1005 by inserting a removable storage medium 1011 into the drive 1010. The program can also be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the storage unit 1008. Alternatively, the program can be installed in advance in the ROM 1002 or the storage unit 1008.
[0413] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0414] 41 realizes the functions of the music analysis device 101 in FIG.
[0415] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device with multiple modules housed in a single housing, are both systems.
[0416] Furthermore, the embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure.
[0417] For example, the present disclosure can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.
[0418] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0419] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0420] The present disclosure may also be configured as follows. <1> An information processing device comprising: a similarity calculation unit that calculates similarities between segments that make up a piece of music based on music data that is data of the piece of music; and a part estimation unit that estimates parts for each segment based on the similarities between the segments that make up the piece of music. <2> An information processing device described in <1>, wherein: a music composition feature extraction unit that extracts music composition features for each segment based on lyric data of the music data, and the part estimation unit estimates parts that make up the piece of music based on the music composition feature. <3> The information processing device described in <2>, wherein the music composition feature extraction unit sets the segments based on line break information in the lyric data and extracts the music composition feature for each segment, and the part estimation unit estimates parts that make up the piece of music for each segment based on the similarities between the segments that make up the piece of music. <4> The information processing device described in <3>, wherein the music structure feature extraction unit performs morphological analysis on the lyric data and extracts music structure features by converting the lyric data into a matrix for each segment based on the number of moras for each segment, and the similarity calculation unit calculates the similarity between the segments based on the distance between the matrices of the segments that make up the music structure feature, with a smaller distance indicating greater similarity. <5> The information processing device described in <4>, wherein the part estimation unit counts, for each segment, the number of other segments for which the distance is smaller than a predetermined minimum distance threshold and higher than a predetermined similarity, and estimates a segment for which the count is greater than a predetermined value as a chorus part that makes up the music. <6> The information processing device described in <5>, wherein the predetermined minimum distance threshold is a reference value at which, when a process of sorting the distances between all segments in ascending order, setting reference values in ascending order from the smallest value, and comparing the next largest value with the reference value is repeated, the next largest value becomes a value that is greater than a predetermined percentage of the reference value.<7> The information processing device described in <5>, wherein the part estimation unit estimates groups of segments that will be the same part based on a similarity corresponding to the distance between the segments, and estimates parts for each group of segments that will be the same part based on the distance between the segments. <8> The part estimation unit estimates other segments that will be the same part for each segment based on a similarity corresponding to the distance between the other segments, and estimates segments that are mutually estimated as the same part as a group of segments that will be the same part. <9> The information processing device described in <7>, wherein the part estimation unit estimates other segments that will be the same part for each segment based on a similarity corresponding to the distance between the other segments, and estimates the first, second, and third segments as a group of segments that will all be the same part when a second segment and a third segment are estimated as the same part from a first segment and neither the second segment nor the third segment is estimated as the same part from the first segment. <10> The information processing device described in <7>, wherein when there are multiple groups that constitute the same part and the chorus group is identified, the part estimation unit estimates the segments of groups other than the chorus group in order of decreasing similarity between the segments within each group, in the order of A verse, B verse, etc. <11> The information processing device described in <7>, wherein the part estimation unit estimates the segments of a group in which the similarity between the segments within each group is lower than a predetermined value as parts below the C verse. <12> The information processing device described in <11>, wherein the predetermined value is a value based on an outlier of the distance that is the similarity of each segment with other segments. <13> The information processing device described in <1>, further comprising a parallel key modulation estimation unit that estimates a parallel key and modulation of the song based on chord data in the song data.<14> The information processing device described in <13>, further comprising: a chord progression estimation unit that estimates a chord progression pattern based on the chord data and the parallel keys and modulations of the musical piece; a key estimation unit that estimates the key of the musical piece based on the chord data and the parallel keys and modulations of the musical piece; and a chord group function estimation unit that estimates a chord group and a chord function of each chord based on the chord data and the parallel keys and modulations of the musical piece. <15> The information processing device described in <1>, further comprising: a beat analysis feature extraction unit that extracts beat analysis features from sound source data of the musical piece data; and a beat estimation unit that estimates beats of the musical piece based on the beat analysis features. <16> The information processing device described in <15>, further comprising: a beat affiliation probability estimation unit that estimates affiliation probabilities of a predetermined number of types of beats based on the beat analysis features, wherein the beat estimation unit estimates the beats of the musical piece based on the affiliation probabilities of the predetermined number of types of beats estimated by the beat affiliation probability estimation unit. <17> The information processing device described in <16>, wherein the beat affiliation probability estimation unit is an affiliation probability estimation model that estimates the affiliation probability of the predetermined number of types of beats based on the beat analysis feature, and the beat estimation unit is a beat estimation model that estimates beats of the music piece based on the affiliation probability of the predetermined number of types of beats, and the affiliation probability estimation model and the beat estimation model are both formed by machine learning based on a random forest classification algorithm. <18> The information processing device described in <1>, further including a UI image generation unit that generates a UI image based on the parts that constitute the music piece estimated by the part estimation unit. <19> An information processing method including: performing a similarity calculation process that calculates a similarity between segments that constitute the music piece based on music data that is data of the music piece; and performing a part estimation process that estimates parts with the segments as units based on the similarity between the segments that constitute the music piece.<20> A program that causes a computer to function as: a similarity calculation unit that calculates similarities between segments that make up a piece of music, based on music data that is data of the piece of music; and a part estimation unit that estimates parts, with the segments as units, based on the similarities between the segments that make up the piece of music.
[0421] REFERENCE SIGNS LIST 101 Music analysis device, 102 Music DB, 103 Chord-related DB, 111 Input unit, 111a Recognition unit, 112 Data acquisition unit, 113 Chord data extraction unit, 114 Chord analysis unit, 115 Lyric extraction unit, 116 Music structure analysis unit, 117 Acoustic feature extraction unit, 118 Beat analysis unit, 119 UI generation unit
Claims
1. An information processing device comprising: a similarity calculation unit that calculates similarities between segments that make up a song based on song data, which is data of the song; and a part estimation unit that estimates parts using the segments as units based on the similarities between the segments that make up the song.
2. An information processing device as described in claim 1, comprising: a music composition feature extraction unit that extracts music composition features for each segment based on lyric data from the music data; and a part estimation unit that estimates the parts that make up the music based on the music composition feature values.
3. The information processing device described in claim 2, wherein the music composition feature extraction unit sets the segments based on line break information of the lyrics data and extracts the music composition feature values using the segments as units, and the part estimation unit estimates the parts that make up the music using the segments as units based on the similarity between the segments that make up the music.
4. The information processing device described in claim 3, wherein the music composition feature extraction unit performs morphological analysis on the lyric data and extracts music composition features by converting the lyric data into a matrix for each segment based on the number of moras for each segment, and the similarity calculation unit calculates the similarity between the segments based on the distance between the matrices of the segments that make up the music composition feature, with the smaller the distance, the more similar the segments are.
5. The information processing device described in claim 4, wherein the part estimation unit counts, for each segment, the number of other segments whose distance is smaller than a predetermined minimum distance threshold and whose similarity is higher than a predetermined value, and estimates the segments whose count is greater than a predetermined value as chorus parts that constitute the song.
6. The information processing device according to claim 5, wherein the predetermined minimum distance threshold is a reference value when, when the process of sorting all inter-segment distances in ascending order, setting reference values in ascending order from the smallest value, and comparing the next largest value to the reference value is repeated, the next largest value becomes a value that is greater than a predetermined percentage of the reference value.
7. The information processing device according to claim 5, wherein the part estimation unit estimates groups of segments that constitute the same part based on a similarity corresponding to the distance between the segments, and estimates the part for each group of segments that constitute the same part based on the distance between the segments.
8. The information processing device according to claim 7, wherein the part estimation unit estimates other segments that are the same part based on a similarity corresponding to the distance between each segment and other segments, and estimates segments that are mutually estimated as the same part as a group of segments that are the same part.
9. The information processing device described in claim 7, wherein the part estimation unit estimates other segments that are the same part based on a similarity corresponding to the distance from each segment to other segments, and when a second segment and a third segment are estimated to be the same part from the perspective of a first segment and neither the second segment nor the third segment is estimated to be the same part as the first segment, and when the second segment and the third segment are estimated to be the same part as each other, the information processing device estimates the first segment, the second segment, and the third segment as a group of segments that are all the same part.
10. The information processing device of claim 7, wherein when there are multiple groups that are the same part and the chorus group is identified, the part estimation unit estimates the segments of groups other than the chorus group in order of decreasing similarity between the segments within each group, in the order of A melody, B melody, etc.
11. The information processing device according to claim 7, wherein the part estimation unit estimates segments of a group in which the similarity between the segments within each group is lower than a predetermined value as parts below the C melody.
12. The information processing device according to claim 11, wherein the predetermined value is a value based on an outlier of the distance representing the similarity of each segment with other segments.
13. The information processing device according to claim 1, further comprising a parallel key modulation estimation unit that estimates the parallel key and modulation of the music piece based on chord data in the music piece data.
14. The information processing device according to claim 13, further comprising: a chord progression estimation unit that estimates a chord progression pattern based on the chord data and the parallel keys and modulations of the music; a key estimation unit that estimates the key of the music based on the chord data and the parallel keys and modulations of the music; and a chord group function estimation unit that estimates the chord group and chord function of each chord based on the chord data and the parallel keys and modulations of the music.
15. The information processing device according to claim 1, further comprising: a beat analysis feature extraction unit that extracts beat analysis features from sound source data in the music data; and a beat estimation unit that estimates the beat of the music based on the beat analysis features.
16. The information processing device of claim 15, further comprising a beat affiliation probability estimation unit that estimates the affiliation probability of a predetermined number of types of beats based on the beat analysis features, wherein the beat estimation unit estimates the beat of the music piece based on the affiliation probability of the predetermined number of types of beats estimated by the beat affiliation probability estimation unit.
17. The information processing device described in claim 16, wherein the beat affiliation probability estimation unit is an affiliation probability estimation model that estimates the affiliation probability of the predetermined number of types of beats based on the beat analysis features, and the beat estimation unit is a beat estimation model that estimates the beat of the music piece based on the affiliation probability of the predetermined number of types of beats, and the affiliation probability estimation model and the beat estimation model are both formed by machine learning based on a random forest classification algorithm.
18. The information processing device according to claim 1, further comprising a UI image generation unit that generates a UI image based on the parts that constitute the music piece estimated by the part estimation unit.
19. An information processing method comprising: performing a similarity calculation process to calculate the similarity between segments that make up a song based on song data, which is data of the song; and performing a part estimation process to estimate parts using the segments as units based on the similarity between the segments that make up the song.
20. A program that causes a computer to function as: a similarity calculation unit that calculates the similarity between segments that make up a song based on song data; and a part estimation unit that estimates parts using the segments as units based on the similarity between the segments that make up the song.
Citation Information
Patent Citations
Playing signal processor, storage medium, and method and program for playing signal processing
JP2002244657A
Method and device to detect chorus segment in music acoustic data and program to execute the method
JP2004233965A
Musical piece structure analyzing method, program and device
JP2008065153A
Electronic musical instrument, method for controlling the electronic musical instrument, and program for the electronic musical instrument
JP2018159831A