Karaoke equipment
The karaoke device adapts breath timing suggestions to individual singers by analyzing song data and real-time breathing patterns, improving performance through personalized breath cues.
Patent Information
- Application Number
- JP2022025428
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-22
- Publication Date
- 2025-10-16
- Estimated Expiration
- 2042-02-22
AI Technical Summary
Existing karaoke machines can only present predetermined breath timings for a single piece of music, lacking adaptability to individual singers' needs.
A karaoke device with an extraction unit to identify breath-taking timings based on song reference data, a detection unit to detect actual breathing timings from audio and video signals, and a presentation unit to inform the singer of appropriate breath timings.
Enables the presentation of appropriate breath timings tailored to individual singers, enhancing their performance by providing timely breath cues.
Smart Images

Figure 0007755512000001 
Figure 0007755512000002 
Figure 0007755512000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a karaoke machine. [Background technology]
[0002] One of the singing support functions provided by karaoke machines is to suggest breath timing to the singer.
[0003] For example, Patent Document 1 discloses a technology that enables intuitive understanding of breath timing by displaying discrete sounds before and after the breath timing based on information indicating the breath timing contained in reference data. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-031394 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the technology of Patent Document 1 can only present predetermined breath timings for a single piece of music.
[0006] An object of the present invention is to provide a karaoke machine that can suggest appropriate breathing timings to a singer. [Means for solving the problem]
[0007] One invention for achieving the above object is a karaoke device having an extraction unit that extracts timings at which breathing can be taken in a song based on reference data for the song; a detection unit that detects breathing information including timings at which a singer takes a breath based on an audio signal corresponding to the singing voice of the singer singing the song karaoke and / or a video signal obtained by filming the singer; an identification unit that identifies timings at which the singer can take the next breath based on the extracted timings at which breathing can be taken in the song and the detected breathing information; and a presentation unit that presents the identified timing at which the singer can take the next breath to the singer. Other features of the present invention will become apparent from the following description and drawings. [Effects of the Invention]
[0008] According to the present invention, it is possible to present appropriate breathing timing to a singer. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a diagram showing a karaoke device according to an embodiment; [Figure 2] 1 is a diagram showing a karaoke machine main body according to an embodiment; [Figure 3] 10A and 10B are diagrams illustrating an example in which the presenting unit presents the timing of the next breath according to the embodiment. [Figure 4] 4 is a flowchart showing the process of the karaoke device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] <Embodiment> A karaoke device according to an embodiment will be described with reference to FIGS.
[0011] ==Karaoke Equipment== The karaoke device K is a device for performing karaoke songs and for singers to sing karaoke. As shown in Fig. 1, the karaoke device K includes a karaoke main unit 10, a speaker 20, a display device 30, a microphone 40, a remote control device 50, and a photographing means 60.
[0012] The karaoke machine main unit 10 performs various controls related to karaoke performance and singing, such as controlling the karaoke performance of the selected song, controlling the display of lyrics and background images, and processing audio signals input through the microphone 40. The speaker 20 is configured to emit karaoke performance sounds and singing voices based on signals from the karaoke machine main unit 10. The display device 30 is configured to display videos and images on a screen based on signals from the karaoke machine main unit 10. The microphone 40 is configured to convert the singing voice of the singer into an analog audio signal and input it to the karaoke machine main unit 10. The remote control device 50 is a device for performing various operations on the karaoke machine main unit 10. The imaging means 60 is a camera for imaging the singer.
[0013] 2, the karaoke machine 10 according to this embodiment includes a storage unit 10a, a communication unit 10b, an input unit 10c, a performance unit 10d, and a control unit 10e. Each component is connected to a bus B via an interface (not shown).
[0014] [Storage means] The storage means 10a is a large-capacity storage device that stores various types of data, including music data.
[0015] The music data includes music identification information for identifying each individual song. The music identification information is information unique to each song, such as a music ID for identifying the song. The music data includes accompaniment data, reference data, etc.
[0016] The accompaniment data is the data that is the source of the karaoke performance sound. The reference data is data that indicates the main melody to be sung in the karaoke performance song. Each song (each song's reference data) is made up of multiple notes. Each note is set with a predetermined pitch, vocalization start timing, vocalization end timing, etc. The vocalization start timing is the timing at which vocalization of the character corresponding to a certain note should begin, and the vocalization end timing is the timing at which vocalization of the character corresponding to a certain note should end. The vocalization start timing and vocalization end timing can be expressed as elapsed time based on the timing at which the song starts to be played.
[0017] The storage means 10a also stores lyric data for displaying lyrics corresponding to a song on the display device 30 or the like in sync with the karaoke performance of the song. The lyric data includes information on characters that make up the lyrics of the song. Each character is associated with one note. The lyric data also includes timing information that indicates the timing of displaying each character. The timing of displaying each character can be indicated by the elapsed time relative to the timing when the song starts to be played.
[0018] Furthermore, the storage means 10a stores background image data, such as background images to be displayed on the display device 30 during karaoke performance, and song attribute information, such as the singer's name, music genre information, and tempo information of the song.
[0019] [Communication means, input means, performance means] The communication means 10b provides an interface for communicating with the remote control device 50. The input means 10c is configured to allow the singer to input various instructions. The input means 10c is a button or the like provided on the karaoke main unit 10. Alternatively, the remote control device 50 may function as the input means 10c. The performance means 10d performs karaoke performance of a song and processes signals based on the singing voice input through the microphone 40 under the control of the control means 10e.
[0020] [Control means] The control means 10e performs various controls in the karaoke device K. The control means 10e includes a CPU and a memory (neither of which is shown). The CPU executes programs stored in the memory to realize various functions.
[0021] In this embodiment, the CPU executes a program stored in the memory, whereby the control means 10e functions as an extraction unit 100, a detection unit 200, an identification unit 300, and a presentation unit 400.
[0022] (Extraction part) The extraction unit 100 extracts timings at which breaths can be taken in a piece of music based on reference data for the piece of music.
[0023] The timing at which a breath is permitted is the timing at which the singer can take a breath comfortably (with ease). For example, in a non-sung section (a section in a song where no lyrics to be sung are set, such as an introductory section, interlude, or postlude) or in a sung section (a section in a song where lyrics to be sung are set, such as an verse or chorus), if there is a certain amount of time between the end of uttering one character and the start of uttering the next character, the singer can take a breath comfortably. The timing at which a breath is permitted can be indicated by the elapsed time relative to the start of the song.
[0024] Specifically, the extraction unit 100 extracts the first note N from the reference data for the song selected by the singer. n Timing of end of utterance TE n and the next note N n+1 Timing of utterance start TS n+1 Time difference between n+1 -TE n ) is calculated. The extraction unit 100 checks whether the calculated time difference satisfies a first predetermined condition. If the first predetermined condition is satisfied, the extraction unit 100 calculates the time difference of the note N n Timing of end of utterance TE nas timings at which a breath can be taken. The extraction unit 100 extracts multiple timings at which a breath can be taken in a song by performing the above process for each note (except the last note of the song) included in the song selected by the singer in the order of the notes. The first predetermined condition is a condition that is set in advance assuming a time at which a breath can be taken comfortably, such as 250 msec or more, or longer than 300 msec.
[0025] (Detection unit) The detection unit 200 detects breathing information, including the timing of breathing taken by a singer, based on an audio signal corresponding to the singing voice of the singer singing a karaoke song and / or a video signal obtained by filming the singer. The timing of breathing taken by a singer is the timing of breathing actually taken while singing a karaoke song. The timing of breathing taken by a singer can be indicated by the elapsed time relative to the timing at which the song starts to be played.
[0026] Specifically, the detection unit 200 uses a known method to analyze the frequency characteristics of the audio signal corresponding to the singer's singing voice input through the microphone 40. If the detection unit 200 determines based on the analysis result that the audio signal does not contain an audio signal due to exhalation (an audio signal corresponding to harmonic components) but does contain an audio signal corresponding to inhalation noise specific to breathing, it detects the time at which the singing voice corresponding to the audio signal was input as the timing of the singer taking a breath.
[0027] In addition, the detection unit 200 applies the technology described in Patent Publication No. 2007-271977 to detect the timing of the singer's breathing based on the power of the singer's voice signal (corresponding to the frequency characteristics described above) and the musical score sound data (corresponding to the reference data described above).
[0028] On the other hand, the detection unit 200 uses a known method to analyze the video signal obtained by photographing the singer with the photographing means 60. If the detection unit 200 determines based on the analysis result that a movement specific to breathing (for example, up and down movement of the shoulders or head, or opening and closing of the mouth) is included, it detects the time at which the movement was photographed as the timing of the singer taking a breath.
[0029] Alternatively, the detection unit 200 may compare the timing of a breath taken by the singer detected based on the audio signal with the timing of a breath taken by the singer detected based on the video signal, and detect only the timings that match or are close to each other as the timing of a breath taken by the singer.The detection unit 200 may also compare the timing of a breath taken by the singer detected based on the audio signal with the timing of a breath taken by the singer detected based on the video signal, and detect the earlier one (the one with the shorter elapsed time from the timing when the performance of the song starts) as the timing of a breath taken by the singer.
[0030] (Specific part) The determination unit 300 determines the timing at which the singer can take the next breath based on the extracted timing at which a breath can be taken in the music piece and the detected breathing information. The timing at which the singer can take the next breath can be indicated by the elapsed time based on the timing at which the music piece starts to be played.
[0031] Specifically, the determination unit 300 determines the timing BT of the breath taken by the singer included in the detected breath information and the timing TE after the timing BT among the possible breath taking times in the extracted music piece. m Time difference (TE m The determination unit 300 determines whether the determined time difference satisfies a second predetermined condition. If the second predetermined condition is satisfied, the determination unit 300 determines that the timing TE mis identified as the timing when it is possible to take the next breath. The second predetermined condition is a condition that is set in advance assuming a time when it is highly likely that the breath will run out (a time when karaoke singing is possible without taking a breath), such as 10 seconds or less, or less than 15 seconds.
[0032] Note that karaoke singing can be continued without taking a breath for a certain period of time after taking a breath. In other words, among the times when it is possible to take a breath after time BT, there are times when it is not necessary to specifically specify the time when it is possible to take the next breath. Therefore, the second predetermined condition may be set not only for times when it is likely that the breath will run out (times when karaoke singing is possible without taking a breath), such as between 10 and 15 seconds, or between 9 and 13 seconds, but also for times when it is likely that taking a breath is not necessary.
[0033] (Presentation part) The presenting section 400 presents the singer with the identified timing at which the next breath can be taken.
[0034] Specifically, the presentation unit 400 can display the identified timing at which the next breath can be taken, in association with the display of the lyrics of the song.
[0035] The presentation unit 400 refers to timing information included in the lyric data of the song and extracts characters that are set with timing that matches the identified timing at which the next breath can be taken. When the extracted characters are displayed on the display screen of the display device 30, the presentation unit 400 displays them in a way that lets the singer know when the next breath can be taken. For example, as shown in FIG. 3, the presentation unit 400 can present the singer with the timing for the next breath by displaying the character "breath" after the extracted character "ga." Note that the presentation unit 400 may also display an icon (e.g., a breath symbol) that encourages the singer to take a breath instead of the character "breath."
[0036] On the other hand, the presentation unit 400 may measure the elapsed time from the start of the music performance, display a countdown according to the identified timing at which the next breath can be taken, and display the word "Take a breath" when the identified timing at which the next breath can be taken arrives.
[0037] Alternatively, the presentation unit 400 can present to the singer the timing at which the next identified breath can be taken by having the speaker 20 emit a sound corresponding to the countdown instead of displaying the countdown.
[0038] ==About the operation of the karaoke device K== Next, a specific example of the operation of the karaoke device K in this embodiment will be described with reference to Fig. 4. Fig. 4 is a flowchart showing an example of the operation of the karaoke device K.
[0039] Singer S operates remote control device 50 to select song X. Karaoke device K registers the song ID of song X in the reservation queue to reserve a karaoke performance of song X (reserving a karaoke performance of song X; step 10).
[0040] The extraction unit 100 extracts timings at which breathing is permitted in the music piece X based on the reference data for the music piece X read out from the storage means 10a (extracting timings at which breathing is permitted in the music piece X; step 11).
[0041] Specifically, the extraction unit 100 calculates the time difference (TS2-TE1) between the timing TE1 at which the first note N1 ends and the timing TS2 at which the next note N2 starts for the piece of music X. The extraction unit 100 then checks whether the calculated time difference satisfies a first predetermined condition. In this example, the first predetermined condition is 250 msec or more. That is, if the interval between the timing at which one note ends and the timing at which the next note starts is 250 msec or more, a comfortable breath can be taken.
[0042] The extraction unit 100 extracts multiple timings where a breath can be taken in the song X by repeating the above process for the order of the notes included in the song X. In this example, the timings where a breath can be taken are the timing TE4 (3,000 msec) when the vocalization of note N4 ends, the timing TE8 (7,500 msec) when the vocalization of note N8 ends, ..., note N n Timing of end of utterance TE n (38,000msec), Note N o Timing of end of utterance TE o (42,000msec), Note N p Timing of end of utterance TE p (46,500msec), Note N q Timing of end of utterance TE q (49,000msec), Note N r Timing of end of utterance TE r (52,500msec), Note N s Timing of end of utterance TE s (55,500msec), Note N t Timing of end of utterance TE t (61,500 msec), ... are extracted. Note that the above timing values are based on the timing when the performance of song X starts (0 msec).
[0043] Thereafter, the karaoke device K starts the karaoke performance of the song X (start karaoke performance of song X; step 12). The singer S uses the microphone 40 to sing the song X in karaoke form.
[0044] The detection unit 200 detects breathing information including the timing of breathing taken by the singer S based on the audio signal corresponding to the singing voice of the singer S who sings the song X in karaoke (detecting breathing information; step 13).
[0045] Specifically, the detection unit 200 analyzes the frequency characteristics of the audio signal corresponding to the singing voice of the singer S input through the microphone 40 using a known method. If the detection unit 200 determines based on the analysis result that the audio signal does not contain an audio signal due to exhalation but does contain an audio signal corresponding to inhalation noise specific to breathing, it detects the time at which the singing voice corresponding to the audio signal was input as the timing BT of the singer S taking a breath. In this example, it is assumed that the breath was taken 40 seconds after the timing (0 msec) at which the performance of the song X started. In this case, the detection unit 200 detects 40,000 msec as the timing BT.
[0046] The determination unit 300 determines the timing at which the singer S can take the next breath based on the timing at which breathing is possible in the song extracted in step 11 and the breathing information detected in step 13 (determines the timing at which the next breath can be taken; step 14).
[0047] Specifically, the determination unit 300 determines the time difference between the timing BT (40,000 msec) of the breath taken by the singer S, which is included in the detected breath information, and the timing after timing BT among the possible breath taking times in the song X extracted in step 11. The determination unit 300 determines whether the determined time difference satisfies a second predetermined condition. In this example, the second predetermined condition is set to be between 10 and 15 seconds. That is, it is assumed that karaoke singing is possible without taking a breath for 10 seconds after the actual breath taking, but that there is a high possibility that the singer will run out of breath if the breath taking time exceeds 15 seconds.
[0048] Here, among the possible breath-taking timings in music X extracted in step 11, note N n Timing of end of utterance TE n The timing before (38,000 msec) is a timing before the timing BT. o Timing of end of utterance TE oThe time difference between the timing BT (40,000 msec) of the breath taken by singer S and the timing (42,000 msec) of the breath taken by singer S is calculated in order, and it is determined whether or not the second predetermined condition is met.
[0049] Timing BT (40,000msec) and Note N o Timing of end of utterance TE n The time difference between the timing BT (40,000 msec) and the note N is 2,000 msec. Therefore, the identifying unit 300 determines that the second predetermined condition is not satisfied. p Timing of end of utterance TE p The time difference between the timing BT (40,000 msec) and the note N (46,500 msec) is 6,500 msec. q Timing of end of utterance TE q The time difference between this and (49,000 msec) is 9,000 msec. Therefore, the identifying unit 300 determines that the second predetermined condition is not satisfied.
[0050] On the other hand, timing BT (40,000 msec) and note N r Timing of end of utterance TE r The time difference between this and (52,500 msec) is 12,500 msec. Therefore, the identifying unit 300 determines that the second predetermined condition is satisfied.
[0051] In addition, timing BT (40,000 msec) and note N s Timing of end of utterance TE s The time difference between this and (55,500 msec) is 15,500 msec. Therefore, the identification unit 300 determines that the second predetermined condition is not satisfied. Since it is clear that the subsequent utterance end timings do not satisfy the second predetermined condition, the identification unit 300 ends the process.
[0052] The identifying unit 300 identifies the note N that is determined to satisfy the second predetermined condition. r Timing of end of utterance TE rThe determination unit 300 determines (52,500 msec) as the timing at which it is possible to take the next breath. The determination unit 300 outputs the determined timing at which it is possible to take the next breath to the presentation unit 400.
[0053] The presenting unit 400 presents to the singer S the timing at which the next breath can be taken, which has been identified in step 14 (presenting the timing at which the next breath can be taken; step 15).
[0054] Specifically, the presentation unit 400 refers to timing information included in the lyrics data of the song X, and presents the timing of note N r Timing of end of utterance TE r (52,500 msec) and the character (i.e., note N) r The presentation unit 400 extracts the extracted note N r When the character corresponding to the note N is displayed on the display screen of the display device 30, the singer S must wait until the next breath is possible (note N) r Timing of end of utterance TE r ) is displayed so that it is clear.
[0055] The karaoke device K repeats the processes from step 13 to step 15 until the karaoke performance of the song X is completed (if Y in step 16). If multiple timings for taking the next breath are identified, they are displayed in accordance with the lyrics.
[0056] As is clear from the above, the karaoke device K of this embodiment comprises an extraction unit 100 that extracts the timing at which a breath can be taken in a song based on reference data for the song; a detection unit 200 that detects breathing information including the timing at which a singer takes a breath based on an audio signal corresponding to the singing voice of the singer singing the song karaoke and / or a video signal obtained by filming the singer; an identification unit 300 that identifies the timing at which the singer can take the next breath based on the extracted timing at which a breath can be taken in the song and the detected breathing information; and a presentation unit 400 that presents the identified timing at which the singer can take the next breath to the singer.
[0057] According to this karaoke device K, the next breath timing can be identified from among the possible breath timings in a certain song based on the breath timings actually taken by the singer while singing the karaoke version of the song, and the identified breath timing can be presented to the singer. This allows the singer to grasp the appropriate timing for the next breath. In other words, the karaoke device K according to this embodiment can present the appropriate breath timing to the singer.
[0058] Furthermore, in the karaoke device K according to this embodiment, the presenting unit 400 can display the identified timing at which the next breath can be taken, in association with the display of the lyrics of the song. With this karaoke device K, the singer can be visually informed of the timing at which they can take a breath.
[0059] In the above example, note N is the next possible timing for taking a breath. r Timing of end of utterance TE r On the other hand, the second predetermined condition is not set for each singer. Therefore, depending on the singer's singing ability and singing style, there is a possibility that the singer will run out of breath before the timing for the next breath that is determined based on the second predetermined condition.
[0060] Therefore, the identification unit 300 may identify, among the timings at which a breath can be taken in the extracted music piece, a timing that is earlier than the second predetermined condition in addition to a timing that satisfies the second predetermined condition as a timing at which the next breath can be taken.
[0061] For example, in the above example, the identifying unit 300 identifies the note N that is determined to satisfy the second predetermined condition. r Timing of end of utterance TE r (52,500 msec) and a note N that is faster than the second predetermined condition range (10 seconds or more and 15 seconds or less). q Timing of end of utterance TE q (49,000 msec) is identified as the timing at which the next breath can be taken.
[0062] With this configuration, it is possible to present appropriate breathing timings, taking into account differences in the singer's singing ability, etc.
[0063] <Variation 1> The extraction unit 100 may extract timings at which breaths can be taken in a piece of music based on reference data for the piece of music and lyric data for the piece of music.
[0064] First, the extraction unit 100 extracts the timings at which notes end vocalization that satisfy a first predetermined condition by executing the process described in the embodiment. Next, the extraction unit 100 narrows down the timings at which breaths can be taken in the song from among the extracted timings at which notes end vocalization that satisfy a first predetermined condition are performed, based on the lyrics data.
[0065] Specifically, the extraction unit 100 uses known technology (for example, JP 2004-070634 A) to analyze lyrics corresponding to the lyrics data and identify phrases where a comma can be inserted. From among the timings at which the vocalization of notes that satisfy a first predetermined condition ends, the extraction unit 100 extracts, as a timing at which a breath can be taken in the music piece, a timing that matches the timing at which the last character of the identified phrase is displayed.
[0066] Even in typical conversational sentences, there is a gap in the speech between the end of the pronunciation of the last character in a phrase where a comma can be inserted and the start of the pronunciation of the first character in the next phrase, so a breath is often taken. In other words, a natural-sounding breath is possible between the end of the pronunciation of the note corresponding to the last character in a phrase where a comma can be inserted and the start of the pronunciation of the note corresponding to the first character in the next phrase. Therefore, the breath-taking timing extracted based on the lyric data in combination with the music reference data makes it easier for the singer to take a breath, and the breath is also a timing that sounds natural.
[0067] In the above example, the lyrics corresponding to the lyrics data are analyzed to identify phrases where commas can be inserted, but the lyrics data may contain information corresponding to the timing of commas in advance. In this case, the extraction unit 100 narrows down the timings where breaths can be taken in the music based on the timing of commas included in the lyrics data.
[0068] As is clear from the above, in the karaoke device K according to this modification, the extraction unit 100 extracts possible breath-taking timings in a song based on the lyric data of the song. This karaoke device K can provide the singer with more appropriate breath-taking timings.
[0069] <Variation 2> The breath information detected by the detection unit 200 may include the timing of the breath taken by the singer and the length of the breath taken by the singer.
[0070] For example, the detection unit 200 analyzes the frequency characteristics of the audio signal corresponding to the singer's singing voice input through the microphone 40 using a known method. If the detection unit 200 determines based on the analysis result that the audio signal does not contain an audio signal due to exhalation (an audio signal corresponding to harmonic components) but does contain an audio signal corresponding to inhalation noise specific to breathing, it detects the time at which the singing voice corresponding to that audio signal was input as the timing of the singer's taking a breath. In this case, the detection unit 200 calculates the duration during which the audio signal corresponding to the inhalation noise specific to breathing is included, and detects that duration as the length of the breath taken by the singer.
[0071] Furthermore, the detection unit 200 analyzes the video signal obtained by photographing the singer with the photographing means 60 using a known method. If the detection unit 200 determines based on the analysis result that a movement specific to breathing (for example, up and down movement of the shoulders or head, or opening and closing of the mouth) is included, it detects the time at which the movement was photographed as the timing of the singer taking a breath. In this case, the detection unit 200 calculates the duration of the movement specific to breathing and detects this duration as the length of the breath taken by the singer.
[0072] Alternatively, the detection unit 200 may compare the length of a breath taken by the singer detected based on the audio signal with the length of a breath taken by the singer detected based on the video signal, and detect only the length that matches or is close to the length as the length of a breath taken by the singer.The detection unit 200 may also compare the length of a breath taken by the singer detected based on the audio signal with the length of a breath taken by the singer detected based on the video signal, and detect the longer length as the length of a breath taken by the singer.
[0073] In this modification, the second predetermined condition referred to by the determination unit 300 is set to a plurality of conditions according to the length of the breath taken by the singer. If the breath taken by the singer is short, it is expected that the time during which the singer is likely to run out of breath (the time during which karaoke singing is possible without taking a breath) will also be short. On the other hand, if the breath taken by the singer is long, it is expected that the time during which the singer is likely to run out of breath (the time during which karaoke singing is possible without taking a breath) will also be long. Therefore, it is preferable that the second predetermined condition includes a plurality of patterns according to the length of the breath.
[0074] For example, the second predetermined condition may be 10 seconds or more but less than 15 seconds if the breath length is less than 500 msec, 11 seconds or more but less than 16 seconds if the breath length is 500 msec or more but less than 750 msec, 12 seconds or more but less than 17 seconds if the breath length is 750 msec or more but less than 1,000 msec, or 13 seconds or more but less than 18 seconds if the breath length is 1,000 msec or more.
[0075] As described in the first embodiment, the identification unit 300 o Timing of end of utterance TE o The time difference between the timing BT (40,000 msec) of the breath taken by singer S and the timing (42,000 msec) of the breath taken by singer S is calculated in order, and it is determined whether or not the second predetermined condition is satisfied.
[0076] Timing BT (40,000msec) and Note N o Timing of end of utterance TE o The time difference between the timing BT (40,000 msec) and the note N (42,000 msec) is 2,000 msec. Therefore, the determination unit 300 determines that the second predetermined condition is not met, regardless of the length of the breath. p Timing of end of utterance TE p The time difference between the timing BT (40,000 msec) and the note N (46,500 msec) is 6,500 msec. q Timing of end of utterance TE qThe time difference between this and (49,000 msec) is 9,000 msec. Therefore, the identification unit 300 determines that none of the breaths satisfies the second predetermined condition, regardless of the length of the breath.
[0077] On the other hand, timing BT (40,000 msec) and note N r Timing of end of utterance TE r The time difference between this and (52,500 msec) is 12,500 msec. Therefore, the identification unit 300 determines that the second predetermined condition is met when the breath length is less than 500 msec, when the breath length is 500 msec or more but less than 750 msec, or when the breath length is 750 msec or more but less than 1,000 msec, and determines that the second condition is not met when the breath length is 1,000 msec or more.
[0078] Also, timing BT (40,000 msec) and note N s Timing of end of utterance TE s The time difference between this and (55,500 msec) is 15,500 msec. Therefore, the identification unit 300 determines that the second predetermined condition is not met if the breath length is less than 500 msec, and determines that the second condition is met if the breath length is 500 msec or more but less than 750 msec, if the breath length is 750 msec or more but less than 1,000 msec, or if the breath length is 1,000 msec or more.
[0079] Furthermore, timing BT (40,000 msec) and note N t Timing of end of utterance TE t The time difference between this and (61,500 msec) is 21,500 msec. Therefore, the identification unit 300 determines that the second predetermined condition is not satisfied, regardless of the length of the breath. Since it is clear that the timing of the end of utterance thereafter does not satisfy the second predetermined condition, regardless of the length of the breath, the identification unit 300 ends the process.
[0080] In this way, the detection unit 200 according to this modification can detect breathing information including the length of the breath taken by the singer. By taking the length of the breath taken by the singer into consideration when determining the timing at which the next breath can be taken, it is possible to present the singer with a more appropriate timing for taking a breath.
[0081] <Other> The above-described embodiments are presented as examples and do not limit the scope of the invention. The above configurations can be implemented in appropriate combinations, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. The above-described embodiments and their modifications are included in the scope and spirit of the invention, as well as in the inventions described in the claims and their equivalents. [Explanation of symbols]
[0082] 100 Extraction part 200 Detector 300 Specific section 400 Presentation section K Karaoke equipment
Claims
1. an extraction unit that extracts possible breath-taking timings in a piece of music based on reference data for the piece of music; a detection unit that detects breathing information including timing of breathing taken by a singer based on an audio signal corresponding to the singing voice of the singer who sings the karaoke version of the song and / or a video signal obtained by photographing the singer; a determination unit that determines a timing at which the singer can take the next breath based on the extracted breath-taking timings in the piece of music and the detected breath-taking information; a presentation unit that presents the singer with the identified timing at which the next breath can be taken; A karaoke device having:
2. 2. The karaoke apparatus according to claim 1, wherein the presenting unit displays the identified timing at which the next breath can be taken in association with the display of the lyrics of the song.
3. 3. The karaoke apparatus according to claim 1, wherein the extraction unit extracts timings at which breaths can be taken in the song based on lyric data of the song.
4. 4. The karaoke apparatus according to claim 1, wherein the detection unit detects the breath information including the length of a breath taken by the singer.
Citation Information
Patent Citations
Musical sound generating device
JP1991296789A
Karaoke system
JP2006235427A
Reference display device, and program
JP2016031394A
Karaoke system
JP2021149064A