Karaoke equipment

The karaoke device uses a learning model to predict song difficulty by analyzing pitch and vocal duration patterns, enhancing singing evaluation accuracy and song recommendations.

JP7818469B2Active Publication Date: 2026-02-20DAIICHI KOSHO COMPANY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022088224
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2026-02-20
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

Existing karaoke devices lack an effective method to predict the difficulty level of songs based on the frequency of pitch differences, which affects the accuracy of singing evaluation and song recommendations.

Method used

A karaoke device equipped with a learning model that calculates the number of times adjacent notes appear with specific pitch differences and vocal durations to predict the difficulty level of a song, using machine learning to associate these occurrences with the song's difficulty.

Benefits of technology

Enables accurate prediction of song difficulty, allowing for improved singing evaluation corrections and personalized song recommendations based on the user's singing ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007818469000001
    Figure 0007818469000001
  • Figure 0007818469000002
    Figure 0007818469000002
  • Figure 0007818469000003
    Figure 0007818469000003
Patent Text Reader

Abstract

To provide a Karaoke device capable of performing a predetermined process using a difficulty level of music concerned, which is predicted based on number of occurrences of each difference in pitch of the music.SOLUTION: A Karaoke device includes: a learning model storage unit for storing a machine-learned learning model using number of occurrences of adjacent notes in a piece of music for each difference in pitch and a degree of difficulty of the piece as teacher data; a calculation unit for calculating the number of occurrences of adjacent notes in a piece of music for each difference in pitch based on reference data of a certain piece of music; a prediction unit for predicting the degree of difficulty of a certain piece of music by inputting the calculated number of occurrences of notes for each difference in pitch into the learning model; and a difficulty processing unit performing prescribed processing based on the predicted degree of difficulty.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a karaoke machine. [Background technology]

[0002] There are known techniques for determining the difficulty level of songs that can be karaoke-played by a karaoke device in order to present the difficulty level of songs to users, correct the karaoke singing evaluation results according to the difficulty level, and recommend songs with a difficulty level that matches the user's singing ability. The difficulty level of a song indicates how difficult it is to sing the song. For example, Patent Document 1 discloses a technique for automatically determining the difficulty level based on song data. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-107333 Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present invention is to provide a karaoke device that can perform predetermined processing using the difficulty level of a song that is predicted based on the number of times each pitch difference appears in the song. [Means for solving the problem]

[0005] The inventor analyzed the pitch of each note that constitutes the singing melody of a song and found that the number of times adjacent notes appear with a pitch difference of ±100 cents (semitones) affects the difficulty of karaoke singing. Specifically, the inventor found that the fewer the number of times adjacent notes appear with a pitch difference of ±100 cents (semitones), the lower the difficulty level, while the more the number of times adjacent notes appear with a pitch difference of ±700 cents (fifths) or more, the higher the difficulty level. The present invention was developed based on this discovery and is a technology that can predict the difficulty level of a song by calculating the number of times adjacent notes appear with a pitch difference of ±700 cents (fifths) or more.

[0006] Specifically, one invention for achieving the above object is a karaoke device having a learning model memory unit that stores a learning model that has been machine-trained using training data of the number of times adjacent notes appear for each pitch difference in a song and the difficulty level of singing the song at karaoke; a calculation unit that calculates the number of times adjacent notes appear for each pitch difference in a song selected by a user based on reference data for the song; a prediction unit that predicts the difficulty level of the song by inputting the calculated number of times adjacent notes appear for each pitch difference into the learning model; and a difficulty level processing unit that performs predetermined processing based on the predicted difficulty level. [Effects of the Invention]

[0007] According to the present invention, it is possible to perform predetermined processing using the difficulty level of a piece of music that is predicted based on the number of times each pitch difference in the piece of music appears. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram showing a karaoke device according to a first embodiment. [Figure 2] 1 is a diagram showing a karaoke main unit according to a first embodiment. FIG. [Figure 3] FIG. 2 is a diagram showing teacher data according to the first embodiment. [Figure 4] 4 is a flowchart showing the process of the karaoke device according to the first embodiment. [Figure 5]FIG. 10 is a diagram showing the number of times adjacent notes appear for each pitch difference in a piece of music selected by a user in the first embodiment. [Figure 6] FIG. 10 is a diagram showing predicted difficulty levels according to the first embodiment. [Figure 7] FIG. 10 is a diagram showing teacher data according to the second embodiment. [Figure 8] 10 is a flowchart showing the process of the karaoke device according to the second embodiment. [Figure 9] FIG. 11 is a diagram showing the number of times each note appears for each vocalization duration of a piece of music selected by a user in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] First Embodiment A karaoke device according to a first embodiment will be described with reference to FIGS.

[0010] ==Karaoke Equipment== The karaoke device K is a device for playing karaoke music and for users to sing karaoke. As shown in Fig. 1, the karaoke device K includes a karaoke main unit 10, a speaker 20, a display device 30, a microphone 40, and a remote control device 50.

[0011] The karaoke machine main unit 10 performs various controls related to karaoke performance and singing, such as controlling the karaoke performance of the selected song, controlling the display of lyrics and background images, and processing audio signals input through the microphone 40. The speaker 20 is configured to emit sound based on the sound emission signal from the karaoke machine main unit 10. The display device 30 is configured to display videos and images on a screen based on the signal from the karaoke machine main unit 10. The microphone 40 is configured to convert the singing voice of the user singing karaoke into an analog audio signal and input it to the karaoke machine main unit 10. The remote control device 50 is a device for performing various operations on the karaoke machine main unit 10.

[0012] 2, the karaoke machine 10 according to this embodiment includes a storage unit 10a, a communication unit 10b, an input unit 10c, a performance unit 10d, and a control unit 10e. Each component is connected to a bus B via an interface (not shown).

[0013] [Storage means] The storage means 10a is a large-capacity storage device that stores various types of data, such as a hard disk drive, etc. The storage means 10a stores music data.

[0014] The song data is provided with song identification information for identifying each song. The song identification information is information unique to each song, such as a song ID for identifying the song. The song data includes accompaniment data, reference data, etc. The accompaniment data is data that forms the basis of the karaoke performance sound. The reference data is data that indicates the singing melody of the song performed karaoke, and is used when scoring the user's karaoke singing. The reference data is made up of multiple notes. Each note has a predetermined pitch, vocal duration, etc.

[0015] The storage means 10a stores lyric telop data for displaying lyric telops corresponding to each piece of music on the display device 30 or the like in sync with the karaoke performance, and background image data such as background images to be displayed on the display device 30 or the like during the karaoke performance.

[0016] (Learning model memory section) In this embodiment, a part of the storage area of ​​the storage means 10a functions as a learning model storage unit 100. The learning model storage unit 100 stores a learning model obtained by machine learning training of training data.

[0017] The training data in this embodiment is data that associates the number of times adjacent notes appear for each pitch difference in a song with the difficulty level of singing the song at karaoke. The number of times adjacent notes appear for each pitch difference in a song is input data, and the difficulty level of singing the song at karaoke is correct answer data.

[0018] The number of occurrences of adjacent notes with different pitch differences in a piece of music can be calculated, for example, by performing the same calculation process as that performed by the calculation unit 200 (described later) on the reference data for that piece of music. The number of occurrences of adjacent notes with different pitch differences can be divided into predetermined ranges (for example, every 100 cents).

[0019] The difficulty level for singing a song in karaoke is a value that is set arbitrarily by the karaoke company, and can be assigned a predetermined value (for example, 1 (difficult) to 10 (easy)).

[0020] As mentioned above, the fewer the occurrences of pitch differences of ±100 cents (semitones), the lower the difficulty level, while the more the occurrences of pitch differences of ±700 cents (fifths) or more, the higher the difficulty level.

[0021] Specifically, melodies with a pitch difference of ±200 cents (a whole tone) or ±300 cents (three semitones) are easy to remember and sing. Furthermore, a pitch difference of 0 cents makes it easy to maintain pitch accuracy. In other words, songs with frequent occurrences of these pitch differences are less difficult.

[0022] On the other hand, vocal melodies composed of pitch differences of ±100 cents (semitones) are difficult to remember and sing. Furthermore, the greater the pitch difference, the more difficult it is to maintain pitch accuracy. In other words, songs with many occurrences of pitch differences of ±100 cents or large pitch differences (especially those of ±700 cents or more) are more difficult to sing.

[0023] Figure 3 shows an example of training data. Here, the frequency of occurrence of each pitch difference between adjacent notes is divided into 25 levels. Each training data is also assigned a difficulty level from 1 (difficult) to 10 (easy).

[0024] For example, in the training data TD1, the frequency of occurrence PT1 of each pitch difference between adjacent notes calculated based on the reference data of song X1 (assuming the number of notes is 400) shows that the frequency of occurrence of a pitch difference of 0 cent is higher than the frequency of occurrence of other pitch differences. Furthermore, the frequency of occurrence of pitch differences of ±200 cents and ±300 cents is also relatively high. In other words, song X1 has a low level of difficulty, so a difficulty level of (8) is associated with the frequency of occurrence PT1.

[0025] On the other hand, in the training data TD2, the frequency of occurrence PT2 of each pitch difference between adjacent notes calculated based on the reference data of song X2 (assuming the number of notes is 400) shows that the frequency of occurrence of a pitch difference of -1000 cents is higher than that of other pitch differences. Also, the frequency of occurrence of a pitch difference of 0 cent, a pitch difference of ±200 cents, and a pitch difference of ±300 cents is relatively low. In other words, song X2 is difficult, so a difficulty level of (2) is associated with the frequency of occurrence PT2.

[0026] Karaoke machine K generates a learning model based on multiple training data. Karaoke machine K performs deep learning to learn feature values ​​of the frequency of occurrence of adjacent notes for each pitch difference in an input song, thereby generating a neural network as the learning model. The feature values ​​are, for example, the values ​​and distributions of the frequency of occurrence of adjacent notes for each pitch difference in a song and the difficulty level of the song as output. The neural network is, for example, a CNN, and has an input layer that accepts input of the frequency of occurrence of adjacent notes for each pitch difference in a song, an output layer that outputs the difficulty level, and a middle layer that extracts feature values.

[0027] The karaoke device K stores the generated learning model in the learning model storage unit 100. The learning model may be generated by an external server device (not shown) separate from the karaoke device K. In this case, the external server device can provide the learning model to the karaoke device K. In addition to a neural network, the learning model may be an SVM, a Bayesian network, a regression tree, or the like.

[0028] [Communication means / input means] The communication means 10b provides an interface for communicating with the remote control device 50. The input means 10c is configured to allow the user to input various instructions. The input means 10c is a button or the like provided on the karaoke main unit 10. Alternatively, the remote control device 50 may function as the input means 10c.

[0029] [Means of performance] Based on the control of the control means 10e, the performance means 10d performs karaoke performance of music pieces and processes signals based on singing voices input through the microphone 40. The performance means 10d includes a sound source, a mixer, an amplifier, etc. (none of which are shown).

[0030] [Control means] The control means 10e performs various controls in the karaoke device K. The control means 10e includes a CPU and a memory (neither of which is shown). The CPU executes programs stored in the memory to realize various functions.

[0031] In this embodiment, the CPU executes a difficulty prediction program stored in the memory, and the control means 10e functions as a calculation unit 200, a prediction unit 300, and a difficulty level processing unit 400.

[0032] (Calculation section) The calculation unit 200 calculates the number of occurrences of adjacent notes for each pitch difference in a certain piece of music selected by a user, based on reference data of the certain piece of music.

[0033] For example, a user of the karaoke device K operates the remote control device 50 to select a piece of music that he or she wishes to sing.

[0034] The calculation unit 200 reads out the reference data of the selected song from the storage means 10a. The calculation unit 200 calculates the pitch difference between adjacent notes that make up the reference data. For example, assume that the reference data consists of notes N1 to Nn (n is the total number of notes). The calculation unit 200 calculates the pitch difference Pd1 by subtracting the pitch value of note N1 from the pitch value of note N2. Similarly, the calculation unit 200 calculates the pitch difference Pd2 by subtracting the pitch value of note N2 from the pitch value of note N3. The calculation unit 200 repeats the same process for all adjacent notes to calculate multiple pitch differences Pd1 to Pdn-1 (the number of calculated pitch differences is "the number of all notes - 1"). The calculation unit 200 counts the calculated multiple pitch differences for each value to calculate the number of occurrences of each pitch difference.

[0035] (Prediction Department) The prediction unit 300 predicts the difficulty level of a certain piece of music by inputting the calculated number of occurrences for each pitch difference into a learning model.

[0036] The prediction unit 300 inputs the occurrence counts of adjacent notes for each pitch difference in a piece of music calculated by the calculation unit 200 into the learning model stored in the learning model storage unit 100. The learning model outputs a difficulty level associated with the occurrence counts of each pitch difference that match or are similar to the input occurrence counts of each pitch difference. The prediction unit 300 predicts the difficulty level of the piece of music based on the difficulty level output from the learning model.

[0037] (Difficulty processing section) The difficulty level processing unit 400 performs a predetermined process based on the predicted difficulty level.

[0038] There are various types of predetermined processing based on the difficulty level. For example, the difficulty level processing unit 400 can present the predicted difficulty level to the user. Specifically, the difficulty level processing unit 400 can display the predicted difficulty level on the display device 30 or the remote control device 50. Alternatively, the difficulty level processing unit 400 can emit the predicted difficulty level as sound via the speaker 20.

[0039] The difficulty level processing unit 400 can also correct the karaoke singing evaluation results according to the difficulty level. The karaoke singing evaluation results can be obtained using known scoring techniques. The difficulty level processing unit 400 can also recommend songs with a difficulty level according to the user's singing ability. The user's singing ability can be obtained based on the singing history stored for each user. Known techniques can be used for the song recommendation process.

[0040] ==About the operation of the Karaoke device K== Next, a specific example of the operation of the karaoke device K in this embodiment will be described with reference to Figs. 4 to 6. Fig. 4 is a flowchart showing an example of the operation of the karaoke device K. Fig. 5 is a diagram showing the frequency of occurrence of adjacent notes for each pitch difference in the song Y1 selected by the user. Fig. 6 is a diagram showing the predicted difficulty level. In this example, it is assumed that a user U uses the karaoke device K. It is also assumed that the learning model storage unit 100 stores a learning model obtained by machine learning multiple pieces of training data, including the training data shown in Fig. 3.

[0041] The user U operates the remote control device 50 to select the song Y1 (selecting a song; step 10).

[0042] The calculation unit 200 calculates the number of occurrences of adjacent notes for each pitch difference in the song Y1 based on the reference data of the song Y1 selected by the user U (calculating the number of occurrences for each pitch difference; step 11).

[0043] Specifically, the calculation unit 200 reads reference data (assuming the number of notes is 400) of the song Y1 from the storage unit 10a. The calculation unit 200 calculates the pitch differences for all adjacent notes that make up the reference data. The calculation unit 200 counts the calculated pitch differences for each value to calculate the number of occurrences PTy1 for each pitch difference shown in FIG. 5.

[0044] The prediction unit 300 predicts the difficulty level of the piece of music Y1 by inputting the number of occurrences of each pitch difference calculated in step 11 into the learning model (predicting the difficulty level of the piece of music; step 12).

[0045] Specifically, the prediction unit 300 inputs the occurrence count PTy1 for each pitch difference calculated by the calculation unit 200 into the learning model stored in the learning model storage unit 100. The learning model extracts the occurrence count for each pitch difference that matches or is similar to the input occurrence count PTy1 for each pitch difference, and outputs the difficulty level associated with the extracted occurrence count. For example, the learning model compares the occurrence count for each pitch difference for each section and finds the degree of agreement for each section (for example, 100% for a perfect match). The learning model outputs the difficulty level associated with the occurrence count for each pitch difference that results in the highest average degree of agreement for each section.

[0046] In this example, the learning model extracts the occurrence count PT2 as the occurrence count for each pitch difference that is similar to the input occurrence count PTy1 for each pitch difference, and outputs the difficulty level "2" associated with the extracted occurrence count PT2. In this case, the prediction unit 300 predicts the difficulty level "2" output from the learning model as the difficulty level of the song Y1.

[0047] The difficulty level processing unit 400 presents the difficulty level of the piece of music Y1 predicted in step 12 to the user U (presenting the difficulty level of the piece of music; step 13).

[0048] Specifically, the difficulty level processing unit 400 displays a message on the display screen of the display device 30 saying, "The difficulty level of the song Y1 you selected is '2'. It's a pretty difficult song, but try your best to sing it at karaoke!" (see FIG. 6).

[0049] Presenting the difficulty level of song Y1 to user U is an example of a predetermined process based on difficulty. The difficulty level processing unit 400 may correct the evaluation result of user U's karaoke singing based on the predicted difficulty level. For example, assume that user U's karaoke singing evaluation result of song Y1 is 90 points out of 100. If the difficulty level of the song is 1-2 (very difficult), the difficulty level processing unit 400 may add 2 points to the evaluation result to 92 points. If the difficulty level is 9-10 (very easy), the difficulty level processing unit 400 may subtract 2 points from the evaluation result to 88 points. Otherwise, the evaluation result may remain 90 points. Furthermore, when recommending songs using known techniques, the difficulty level processing unit 400 may recommend songs with a difficulty level that corresponds to the user's singing ability. For example, the difficulty level processing unit 400 may calculate the difficulty level of songs that user U often sings karaoke based on user U's singing history and recommend songs with a difficulty level close to the calculated difficulty level.

[0050] As is clear from the above, the karaoke device K of this embodiment has a learning model storage unit 100 that stores a learning model that has been machine-learned using training data that includes the number of times adjacent notes appear for each pitch difference in a song and the difficulty level of singing the song karaoke; a calculation unit 200 that calculates the number of times adjacent notes appear for each pitch difference in a song selected by a user based on reference data for the song; a prediction unit 300 that predicts the difficulty level of a song by inputting the calculated number of times adjacent notes appear for each pitch difference into the learning model; and a difficulty level processing unit 400 that performs predetermined processing based on the predicted difficulty level.

[0051] The karaoke device K can predict the difficulty of a song selected by a user using a learning model generated using training data based on the frequency of occurrence of each pitch difference between adjacent notes in a song and the difficulty of singing the song. The karaoke device K can then present the predicted difficulty to the user, allowing the user to understand the difficulty of the selected song. In other words, the karaoke device K according to this embodiment can perform a predetermined process using the difficulty of the song predicted based on the frequency of occurrence of each pitch difference in the song.

[0052] Second Embodiment Next, a karaoke device according to a second embodiment will be described with reference to Fig. 7 to Fig. 9. In this embodiment, an example will be described in which the difficulty level of a certain song is predicted using a learning model that has been machine-learned to learn the number of times each note appears in a song for each vocal duration, in addition to the number of times each pitch difference appears and the difficulty level of the song as described in the first embodiment. Note that detailed description of the same configuration as in the first embodiment will be omitted.

[0053] The inventor analyzed the pitch and duration of each note that constitutes the singing melody of a song and found that the number of times adjacent notes appear for each pitch difference and each vocal duration affect the difficulty of karaoke singing. Specifically, when the difficulty level based on the number of times adjacent notes appear for each pitch difference is the same, the difficulty level tends to increase as the number of times vocal durations of 180 msec or less appear. The present invention was developed based on this discovery, and it is possible to predict the difficulty of a song by calculating the number of times adjacent notes appear for each pitch difference and each vocal duration for that song.

[0054] [Storage means] (Learning model memory section) The learning model storage unit 100 according to this embodiment stores, as training data, a learning model that has been machine-learned to estimate the number of times each note appears in a song for each vocal duration. That is, the training data according to this embodiment is data that associates the number of times each pitch difference between adjacent notes appears in a song, the difficulty level of singing the song at karaoke, and the number of times each note appears in a song for each vocal duration.

[0055] The duration of a note corresponds to its length. Therefore, if the note length is long (short), the duration will also be long (short).

[0056] The number of times adjacent notes appear for each pitch difference in a song and the number of times adjacent notes appear for each vocal duration in a song are input data, and the difficulty level of singing the song at karaoke is correct answer data.

[0057] The number of times each note occurs in a piece of music for each vocal duration can be calculated, for example, by performing the same calculation process as that performed by the calculation unit 200 on reference data for the piece of music. The number of times each note occurs in each vocal duration can be divided into predetermined ranges (for example, every 60 msec).

[0058] As described above, when the difficulty level based on the number of occurrences of each pitch difference is the same, the difficulty level increases as the number of occurrences of utterance durations of 180 msec or less increases.

[0059] Specifically, when singing adjacent notes in succession, the shorter the note length of the subsequent note (the shorter the vocal duration), the faster the pitch of the subsequent note must be stabilized. Adjusting the vocal pitch becomes particularly difficult when the vocal duration of the subsequent note is 180 msec or less. In other words, songs with many notes with vocal durations of 180 msec or less are more difficult to sing.

[0060] On the other hand, even if there are many notes with vocalization durations of 180 msec or less, if there are many occurrences of pitch differences of 0 cents, it will be easier to maintain pitch accuracy, and the difficulty level will be lower.

[0061] FIG. 7 shows an example of training data according to this embodiment. Here, the frequency of occurrence of each pitch difference between adjacent notes is divided into 25 levels. The frequency of occurrence of each vocalization duration of the notes is also divided into 9 levels. Furthermore, each training data is assigned a difficulty level from 1 (difficult) to 10 (easy).

[0062] For example, in the training data TD1, the frequency of occurrence of adjacent notes for each pitch difference PT1 calculated based on the reference data for song X1 (assuming the number of notes is 400) shows that the frequency of occurrence of a pitch difference of 0 cent is higher than the frequency of occurrence of other pitch differences. The frequency of occurrence of pitch differences of ±200 cents and ±300 cents is also relatively high. Furthermore, the frequency of occurrence of each note vocalization duration VT1 shows that there are many notes with long vocalization durations. In other words, song X1 has a low overall difficulty level, and is therefore assigned a difficulty level of (8).

[0063] Furthermore, in the training data TD2, the frequency of occurrence of adjacent notes for each pitch difference PT2 calculated based on the reference data for song X2 (assuming the number of notes is 400) is such that the frequency of occurrence of a pitch difference of -1000 cents is higher than the frequency of other pitch differences. The frequency of occurrence of pitch differences of 0 cents, ±200 cents, and ±300 cents is also relatively low. Furthermore, the frequency of occurrence of each note vocalization duration VT2 is also higher for notes with short vocalization durations. In other words, song X2 is generally difficult, and therefore is assigned a difficulty level of (2).

[0064] Furthermore, the teacher data TD6 is calculated based on the reference data for song X6 (assuming the number of notes is 400), and the frequency PT6 of adjacent notes appearing for each pitch difference is almost the same as that for song X2. On the other hand, the frequency VT6 of adjacent notes appearing for each vocal duration shows that there are more notes with longer vocal durations than in song X2. In other words, song X6 is easier to play than song X2, and so is assigned a difficulty level of (4).

[0065] The learning model can be generated in the same manner as in the first embodiment.

[0066] [Control means] (Calculation section) The calculation section 200 according to this embodiment calculates the number of occurrences of each vocalization duration of a note in a certain piece of music based on reference data of the certain piece of music.

[0067] For example, a user of the karaoke device K operates the remote control device 50 to select a piece of music that he or she wishes to sing.

[0068] The calculation unit 200 reads out the reference data of the selected song from the storage means 10a. The calculation unit 200 calculates the voicing duration set for each note that makes up the reference data. For example, assume that the reference data is made up of notes N1 to Nn (n is the total number of notes). The calculation unit 200 calculates the voicing duration VD1 set for note N1. The calculation unit 200 repeats the same process for all n notes to calculate multiple voicing durations VD1 to VDn (the number of voicing durations calculated is the "number of all notes"). The calculation unit 200 counts each of the calculated multiple voicing durations to calculate the number of occurrences of each voicing duration.

[0069] Note that, similar to the first embodiment, the calculation section 200 according to this embodiment also calculates the number of occurrences of adjacent notes for each pitch difference in a piece of music.

[0070] (Prediction Department) The prediction unit 300 according to this embodiment predicts the difficulty level of a certain piece of music by inputting the calculated number of occurrences of pitch differences and the number of occurrences for each vocalization duration into a learning model.

[0071] The prediction unit 300 inputs the frequency of occurrence of adjacent notes for each pitch difference and each vocalization duration calculated by the calculation unit 200 into a learning model stored in the learning model storage unit 100. The learning model outputs a difficulty level associated with the frequency of occurrence for each pitch difference and each vocalization duration that matches or is similar to the input frequency of occurrence for each pitch difference and each vocalization duration. The prediction unit 300 predicts the difficulty level of a certain piece of music based on the difficulty level output from the learning model.

[0072] ==About the operation of the Karaoke device K== Next, a specific example of the operation of the karaoke device K in this embodiment will be described with reference to Figs. 8 and 9. Fig. 8 is a flowchart showing an example of the operation of the karaoke device K. Fig. 9 is a diagram showing the number of occurrences of adjacent notes for each pitch difference and each vocalization duration of the notes in the song Y2 selected by the user. In this example, it is assumed that a user U uses the karaoke device K. It is also assumed that the learning model storage unit 100 stores a learning model obtained by machine learning multiple pieces of training data including the training data shown in Fig. 7.

[0073] The user U operates the remote control device 50 to select the song Y2 (selecting a song; step 20).

[0074] Based on the reference data of the song Y2 selected by the user U, the calculation unit 200 calculates the number of times adjacent notes appear for each pitch difference in the song Y2 and the number of times notes appear for each vocal duration in the song Y2 (calculating the number of times each pitch difference and the number of times each vocal duration).

[0075] Specifically, the calculation unit 200 reads reference data for song Y2 (assuming the number of notes is 400) from the storage unit 10a. The calculation unit 200 calculates the pitch differences for all adjacent notes that make up the reference data. The calculation unit 200 also calculates the vocalization durations that have been set for all notes that make up the reference data. The calculation unit 200 counts the calculated pitch differences for each value and the calculated vocalization durations for each value, thereby calculating the number of occurrences PTy2 for each pitch difference and the number of occurrences VTy2 for each vocalization duration, as shown in FIG. 9.

[0076] The prediction unit 300 predicts the difficulty level of the piece of music Y2 by inputting the number of occurrences for each pitch difference and the number of occurrences for each vocalization duration calculated in step 21 into the learning model (predicting the difficulty level of the piece of music; step 22).

[0077] Specifically, the prediction unit 300 inputs the number of occurrences PTy2 for each pitch difference and the number of occurrences VTy2 for each utterance duration calculated by the calculation unit 200 into the learning model stored in the learning model storage unit 100. The learning model extracts the number of occurrences for each pitch difference and the number of occurrences for each utterance duration that match or are similar to the input number of occurrences PTy2 for each pitch difference and the number of occurrences for each utterance duration, and outputs the difficulty level associated with the extracted number of occurrences. For example, the learning model compares the number of occurrences of pitch differences for each category and determines the degree of agreement for each category (e.g., 100% for a perfect match). The learning model also compares the number of occurrences of utterance durations for each category and determines the degree of agreement for each category (e.g., 100% for a perfect match). The learning model outputs the difficulty level associated with the number of occurrences for each pitch difference and the number of occurrences of utterance duration that results in the highest average degree of agreement for each category, for each of the number of occurrences of pitch differences and the number of occurrences of utterance duration.

[0078] In this example, the learning model extracts the number of occurrences VT6 of the vocalization duration that is closest to the number of occurrences VTy2 for the input vocalization duration from the numbers of occurrences VT2 and VT6 of occurrences PT2 and PT6 identified as the number of occurrences for each pitch difference similar to the number of occurrences PTy2 for the input pitch difference, and outputs the difficulty level "4" associated with the extracted number of occurrences VT6. In this case, the prediction unit 300 predicts the difficulty level "4" output from the learning model as the difficulty level of the song Y2.

[0079] The difficulty level processing section 400 presents the difficulty level of the piece of music Y2 predicted in step 22 to the user U (presenting the difficulty level of the piece of music; step 23).

[0080] As is clear from the above, in the karaoke device K according to this embodiment, the learning model storage unit 100 stores, as training data, a learning model that has been machine-learned to learn the number of times each note appears for each vocal duration in a song. The calculation unit 200 calculates the number of times each note appears for each vocal duration in a certain song based on reference data for the song. The prediction unit 300 predicts the difficulty level of a certain song by inputting the calculated number of times each pitch difference appears and the number of times each note appears for each vocal duration into the learning model.

[0081] This karaoke device K can predict the difficulty of a song selected by a user using a learning model generated using training data including the number of times adjacent notes appear for each pitch difference in the song, the number of times adjacent notes appear for each vocal duration in the song, and the difficulty level of singing the song. The karaoke device K can then present the predicted difficulty level to the user, allowing the user to understand the difficulty level of the selected song. In other words, the karaoke device K of this embodiment can perform a predetermined process using the difficulty level of the song predicted based on the number of times adjacent notes appear for each pitch difference and each vocal duration in the song.

[0082] <Other> The above-described embodiments are presented as examples and do not limit the scope of the invention. The above configurations can be implemented in appropriate combinations, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. The above-described embodiments and their modifications are included in the scope and spirit of the invention, as well as in the inventions described in the claims and their equivalents. [Explanation of symbols]

[0083] 100 Learning model memory unit 200 Calculation Unit 300 Prediction Department 400 Difficulty Processing Unit K Karaoke equipment

Claims

1. a learning model storage unit that stores a learning model that has been machine-trained using training data that includes the number of occurrences of adjacent notes for each pitch difference in a song and the difficulty level of singing the song at karaoke; a calculation unit that calculates the number of occurrences of adjacent notes for each pitch difference in a certain piece of music selected by a user based on reference data of the certain piece of music; a prediction unit that predicts the difficulty level of the certain piece of music by inputting the calculated number of occurrences of each pitch difference into the learning model; a difficulty level processing unit that performs predetermined processing based on the predicted difficulty level; A karaoke device having:

2. the learning model storage unit stores, as the training data, the learning model that has been machine-learned to learn the number of times each note occurs for each vocalization duration in a piece of music; the calculation unit calculates the number of occurrences of each voicing duration of a note in a piece of music based on reference data for the piece of music; 2. The karaoke apparatus according to claim 1, wherein the prediction unit predicts the difficulty level of the certain song by inputting the calculated number of occurrences of the pitch difference and the number of occurrences for each vocalization duration into the learning model.

Citation Information

Patent Citations

  • Karaoke machine

    JP2005107333A