Karaoke equipment

JP2026144213APending Publication Date: 2026-09-09DAIICHI KOSHO COMPANY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025031375
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2026-09-09

Smart Images

  • Figure 2026144213000001_ABST
    Figure 2026144213000001_ABST
Patent Text Reader

Abstract

The present invention provides a karaoke device that can output singing audio with minimal unnaturalness after correcting the singing pitch extracted from the singing voice. [Solution] A karaoke device having a first extraction unit for extracting singing pitch, a second extraction unit for extracting reference pitch, a setting unit that sets a correction ratio at a certain timing if singing pitch and reference pitch are extracted at a certain timing, and sets the correction ratio at a certain timing to 1 if singing pitch is not extracted at a certain timing or if only singing pitch is extracted, a calculation unit that calculates a corrected pitch at a certain timing if singing pitch and reference pitch are extracted at a certain timing, and a sound emission processing unit that emits singing sound corresponding to the calculated corrected pitch from a sound emission means.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a karaoke apparatus.

Background Art

[0002] A karaoke apparatus is equipped with a function that allows karaoke singing to sound pleasant.

[0003] Patent Document 1 discloses a technique that corrects an audio signal of a karaoke singer based on frequency information of a singing melody, so that accompaniment harmonizes well with the singing voice.

Prior Art Literature

Patent Literature

[0004]

Patent Document 1

Summary of the Invention

Problem to be Solved by the Invention

[0005] Here, when correction is performed based on the technique of Patent Document 1, the output singing voice becomes monotonous. Accordingly, there is a possibility that the singer and the audience listening to the karaoke singing may feel a sense of discomfort.

[0006] An object of the present invention is to provide a karaoke apparatus capable of outputting a singing voice with little sense of discomfort when correcting a singing pitch extracted from the singing voice.

Means for Solving the Problem

[0007] One invention to achieve the above objective is a first extraction unit that extracts the singing pitch from the singer's singing voice at predetermined timings from the start of karaoke performance of a song, a second extraction unit that extracts a reference pitch from reference data of the song at predetermined timings, and a setting unit that sets a correction ratio for correcting the singing pitch extracted at predetermined timings, wherein when the singing pitch and reference pitch are extracted at a certain timing, a first value obtained by multiplying the pitch ratio, which is the ratio of the reference pitch to the singing pitch, by a first predetermined value, and a second value obtained by multiplying the correction ratio set at the timing immediately preceding that timing by a second predetermined value, thereby correcting the singing pitch at that timing The karaoke device includes: a setting unit that sets a positive ratio and, if the singing pitch is not extracted at a certain timing or if only the singing pitch is extracted, sets the correction ratio at that timing to 1; a calculation unit that calculates a corrected pitch by correcting the singing pitch extracted at a predetermined timing, wherein if the singing pitch and reference pitch are extracted at a certain timing, the calculation unit calculates the corrected pitch at that timing by multiplying the singing pitch by the set correction ratio at that timing; and a sound emission processing unit that emits the singer's singing voice from a sound emission means, wherein the sound emission processing unit emits the singing voice corresponding to the calculated corrected pitch from the sound emission means. Other features of the present invention will be revealed in the specification and drawings described below. [Effects of the Invention]

[0008] According to the present invention, when the singing pitch extracted from the singing voice is corrected, it is possible to emit singing voice with less unnaturalness. [Brief explanation of the drawing]

[0009] [Figure 1] This is a diagram showing a karaoke device according to the first embodiment. [Figure 2] This is a diagram showing a karaoke machine according to the first embodiment. [Figure 3]This is a flowchart showing the processing of the karaoke apparatus according to the first embodiment. [Figure 4] This figure shows examples of reference pitch and singing pitch at predetermined timing intervals according to the first embodiment. [Figure 5] This is a diagram showing a karaoke machine according to the second embodiment. [Figure 6] This figure shows a transition diagram displayed on the display device according to the second embodiment. [Modes for carrying out the invention]

[0010] <First Embodiment> A karaoke apparatus according to the first embodiment will be described with reference to Figures 1 to 4.

[0011] ==Karaoke Equipment== The karaoke device K is a device for playing karaoke versions of songs and for singers to perform karaoke. As shown in Figure 1, the karaoke device K comprises a karaoke unit 10, speakers 20, a display device 30, a microphone 40, and a remote control device 50.

[0012] The karaoke unit 10 performs various controls related to karaoke performance and singing, such as controlling the karaoke performance of the selected song, controlling the display of lyrics and background images, and processing audio signals input through the microphone 40. The speaker 20 is configured to emit sound based on the sound emission signal from the karaoke unit 10. The speaker 20 is an example of a "sound emission means." A speaker (not shown) installed in a room such as a karaoke room where the karaoke device K is installed may also be used as a "sound emission means." The display device 30 is configured to display video and images on a screen based on signals from the karaoke unit 10. The display device 30 is an example of a "display means." A display (not shown) installed in a room such as a karaoke room where the karaoke device K is installed, or the display screen of the remote control device 50 may also be used as a "display means." The microphone 40 is configured to input the singing voice of the singer's karaoke performance to the karaoke unit 10. The remote control device 50 is a device for performing various operations on the karaoke unit 10.

[0013] As shown in Figure 2, the karaoke main body 10 according to the present embodiment includes a storage means 10a, a communication means 10b, an input means 10c, a playing means 10d, and a control means 10e. Each component is connected to a bus B via an interface (not shown).

[0014] [Storage Means] The storage means 10a is a large-capacity storage device that stores various types of data. The storage means 10a stores music data.

[0015] The music data is assigned music identification information for specifying each individual piece of music. The music identification information is information unique to each music piece, such as a music ID for identifying the music. The music data includes accompaniment data, reference data, section information, and the like.

[0016] The accompaniment data is data that serves as a source for karaoke performance sounds. The reference data is data indicating the singing melody of a piece of music performed in karaoke, and is used when scoring the karaoke singing by a singer. The reference data is composed of a plurality of notes, i.e., musical notes. A reference pitch (hereinafter referred to as "reference pitch") is set for each note in the reference data. The section information indicates a performance section. The performance section is a section in which karaoke performance is performed. The performance section includes a singing section and a non-singing section. The singing section is a section in which lyrics to be sung are set in the music (for example, the A melody, B melody, and chorus of the first verse). The non-singing section is a section in which no lyrics to be sung are set in the music, such as a prelude, an interlude, and a postlude.

[0017] The storage means 10a stores: lyric data for displaying lyrics corresponding to each music piece on the display device 30 or the like in synchronization with the karaoke performance; background video data such as a background video displayed on the display device 30 or the like during karaoke performance; and attribute information of the music piece (music name, singer name, genre, performance time, etc.).

[0018] [Communication Means and Input Means] The communication means 10b provides an interface for communicating with the remote control device 50. The input means 10c is configured for a user of the karaoke apparatus K to input various operations. The input means 10c is a button or the like provided on the karaoke main body 10. Alternatively, the remote control device 50 may function as the input means 10c.

[0019] [Performance means] The performance means 10d performs karaoke performance of music and processes audio signals input through the microphone 40 based on the control of the control means 10e. The performance means 10d includes a sound source, a mixer, an amplifier, and the like (none of which are illustrated).

[0020] [Control means] The control means 10e performs various controls in the karaoke apparatus K. The control means 10e includes a CPU and a memory (none of which are illustrated). The CPU implements various functions by executing programs stored in the memory.

[0021] In the present embodiment, when the CPU executes a program stored in the memory, the control means 10e functions as a first extraction unit 100, a second extraction unit 200, a setting unit 300, a calculation unit 400, and a sound emission processing unit 500.

[0022] (First Extraction Unit) The first extraction unit 100 extracts a singing pitch from the singing voice of a singer at every predetermined timing after the start of karaoke performance of a music piece.

[0023] One value is set in advance for the predetermined timing, for example, 20 msec. The singing pitch is the pitch of the singing voice uttered by the singer at a certain timing.

[0024] For example, the singer operates the remote control device 50 to select a song to sing karaoke. The remote control device 50 transmits song identification information of the selected song to the karaoke device K. Based on the song identification information of the selected song, the karaoke device K reads accompaniment data from the storage means 10a. The karaoke device K outputs the read accompaniment data to the performance means 10d to perform karaoke. The singer sings karaoke using the microphone 40. The first extraction unit 100 sets the start time of the karaoke performance of the song as 0 and extracts the singing pitch from the singer's singing voice at predetermined timings. Known techniques can be used for extracting the singing pitch from the singing voice. The first extraction unit 100 outputs the extracted singing pitch to the setting unit 300, the calculation unit 400, and the sound output processing unit 500.

[0025] On the other hand, for example, during non-singing sections, the singer does not perform karaoke singing. That is, there may be times when there is no input of singing voice. In such cases, the first extraction unit 100 outputs a signal to the setting unit 300 indicating that the singing pitch was not extracted.

[0026] (Second extraction section) The second extraction unit 200 extracts a reference pitch from the song's reference data at predetermined intervals.

[0027] In the example described above, the second extraction unit 200 reads the reference data of the selected song from the storage means 10a and extracts the reference pitch of the note to be pronounced at the same timing as when the first extraction unit 100 extracts the singing pitch. The second extraction unit 200 outputs the extracted reference pitch to the setting unit 300 and the calculation unit 400.

[0028] On the other hand, there may be times when there is no note to be pronounced at a given time. In such cases, the second extraction unit 200 outputs a signal to the setting unit 300 and the sound output processing unit 500 indicating that the reference pitch was not extracted.

[0029] (Settings section) The setting unit 300 sets the correction ratio. Specifically, if the singing pitch and reference pitch are extracted at a certain timing, the setting unit 300 sets the correction ratio at that timing by adding a first value obtained by multiplying the pitch ratio by a first predetermined value and a second value obtained by multiplying the correction ratio value set at the timing immediately preceding that timing by a second predetermined value. On the other hand, if the singing pitch is not extracted at a certain timing or if only the singing pitch is extracted, the setting unit 300 sets the correction ratio at that timing to 1.

[0030] The correction ratio is a percentage used to correct the singing pitch extracted at a predetermined timing. In other words, the correction ratio is set for each predetermined timing. The pitch ratio is the ratio of the reference pitch to the singing pitch. The first predetermined value and the second predetermined value are values ​​used to set the correction ratio. The first predetermined value and the second predetermined value are each set to one value in advance.

[0031] If the singing pitch and reference pitch are extracted at a certain point in time, the correction ratio for that point in time is set by referring to the correction ratio value set at the point immediately preceding that point in time. In other words, a specific value is set as the correction ratio for the current point in time, reflecting the correction ratio at the point immediately preceding that point in time. On the other hand, if the singing pitch is not extracted at a certain point in time, or if only the singing pitch is extracted, the correction ratio for that point in time is set to 1. In other words, regardless of the correction ratio at the point immediately preceding that point in time, the correction ratio for the current point in time is set to 1.

[0032] The first predetermined value is preferably a value greater than 0 and less than 1, and the second predetermined value is preferably a value obtained by subtracting the first predetermined value from 1. Furthermore, the first predetermined value is more preferably a value close to 0.

[0033] The first predetermined value is set as a correction amount that indicates how much the singing pitch is corrected relative to the reference pitch when the singing pitch and reference pitch are extracted at a certain timing. If the first predetermined value is set to 0, the extracted singing pitch is used as is. On the other hand, if the first predetermined value is set to 1, the extracted singing pitch is corrected to the extracted reference pitch. In other words, the closer the first predetermined value is to 0, the smaller the correction amount at a certain timing, and the more corrections are needed until the singing pitch matches the reference pitch. Conversely, the closer the first predetermined value is to 1, the larger the correction amount at a certain timing, and the fewer corrections are needed until the singing pitch matches the reference pitch. Therefore, in order to correct the singing pitch gradually, it is more preferable to use a value close to 0 as the first predetermined value. By using such a first predetermined value and a second predetermined value, the singing pitch can be corrected gradually, so that singing sounds with less unnaturalness can be emitted.

[0034] For example, suppose that at a certain timing Tn, the singing pitch SPn extracted from the first extraction unit 100 is output, and the reference pitch RPn extracted from the second extraction unit 200 is output. In this case, the setting unit 300 calculates the pitch ratio PRn (=RPn / SPn) from the output singing pitch SPn and reference pitch RPn. The setting unit 300 multiplies the pitch ratio PRn by a first predetermined value (preferably a value close to 0) to obtain the first value FVn. The setting unit 300 also multiplies the correction ratio CRn-1 set at the timing Tn-1 immediately preceding a certain timing Tn by a second predetermined value (preferably a value close to 1) to obtain the second value SVn. The setting unit 300 sets the correction ratio CRn at a certain timing Tn by adding the first value FVn and the second value SVn. The setting unit 300 outputs the set correction ratio CRn to the calculation unit 400.

[0035] On the other hand, suppose that at a certain timing Tn, a signal is output from the first extraction unit 100 indicating that the singing pitch was not extracted, or that at a certain timing Tn, the singing pitch SPn extracted from the first extraction unit 100 is output, and a signal is output from the second extraction unit 200 indicating that the reference pitch was not extracted. In this case, the setting unit 300 sets the correction ratio CRn at a certain timing Tn to 1.

[0036] (Calculation section) The calculation unit 400 calculates the corrected pitch. Specifically, when the singing pitch and reference pitch are extracted at a certain timing, the calculation unit 400 calculates the corrected pitch at that timing by multiplying the singing pitch by the correction ratio set for that timing.

[0037] The corrected pitch is the singing pitch corrected at predetermined timings. In other words, the corrected pitch is calculated each time the singing pitch and reference pitch are extracted at predetermined timings.

[0038] For example, suppose that at a certain timing Tn, the singing pitch SPn extracted from the first extraction unit 100 is output, and the reference pitch RPn extracted from the second extraction unit 200 is output. In this case, the setting unit 300 outputs the set correction ratio CRn to the calculation unit 400. The calculation unit 400 calculates the corrected pitch CPn at a certain timing Tn by multiplying the singing pitch SPn by the set correction ratio CRn. The calculation unit 400 outputs the calculated corrected pitch CPn to the sound output processing unit 500.

[0039] (Sound emission processing unit) The sound emission processing unit 500 emits the singer's vocal voice from the sound emission means. Specifically, the sound emission processing unit 500 emits the vocal voice corresponding to the calculated corrected pitch from the sound emission means.

[0040] For example, suppose the calculation unit 400 outputs a corrected pitch CPn. In this case, the sound output processing unit 500 emits the singing voice corresponding to the corrected pitch CPn from the speaker 20. On the other hand, if the calculation unit 400 does not output a corrected pitch CPn, and the singing pitch SPn extracted from the first extraction unit 100 is output, and the reference pitch is not extracted in the second extraction unit 200, the sound output processing unit 500 emits the singing voice from the speaker 20 as is without correction.

[0041] ==Regarding the operation of karaoke machine K== Next, specific examples of the operation of the karaoke device K in this embodiment will be described with reference to Figures 3 and 4. Figure 3 is a flowchart showing an example of the operation of the karaoke device K. Figure 4 shows examples of the reference pitch and singing pitch at predetermined timing intervals according to this embodiment.

[0042] The singer operates the remote control device 50 to select song X to sing karaoke. The remote control device 50 transmits the song ID of song X to the karaoke device K. Based on the song ID of song X, the karaoke device K reads the accompaniment data from the storage means 10a. The karaoke device K outputs the read accompaniment data to the performance means 10d and starts karaoke performance of song X (start of karaoke performance of song. Step 10). The singer sings karaoke of song X using the microphone 40 in time with the karaoke performance.

[0043] The first extraction unit 100 extracts the singing pitch from the singer's vocals at predetermined intervals from the start of the karaoke performance of song X. The second extraction unit 200 extracts the reference pitch from the reference data of song X at predetermined intervals.

[0044] If the singing pitch and reference pitch are extracted at a certain timing (when both steps 11 and 12 are Y), the setting unit 300 sets the correction ratio at that timing by adding a first value obtained by multiplying the pitch ratio by a first predetermined value and a second value obtained by multiplying the correction ratio value set at the timing immediately preceding that timing by a second predetermined value (setting the correction ratio; step 13).

[0045] On the other hand, if the singing pitch is not extracted at a certain time, and the reference pitch is not extracted (when step 11 is N and step 14 is N), the setting unit 300 sets the correction ratio at that time to 1 (set the correction ratio to 1. step 15). Also, if only the reference pitch is extracted (when step 11 is N and step 14 is Y), the setting unit 300 sets the correction ratio at that time to 1, similar to step 15 (set the correction ratio to 1. step 16). Also, if only the singing pitch is extracted (when step 11 is Y and step 12 is N), the setting unit 300 sets the correction ratio at that time to 1, similar to steps 15 and 16 (set the correction ratio to 1. step 17).

[0046] Next, if the singing pitch and reference pitch are extracted at a certain timing (when both step 11 and step 12 are Y), the calculation unit 400 calculates the corrected pitch at a certain timing by multiplying the singing pitch extracted in step 11 by the correction ratio at a certain timing set in step 13 (calculate the corrected pitch; step 18).

[0047] The sound output processing unit 500 emits the singing voice corresponding to the corrected pitch calculated in step 18 from the speaker 20 (emits the singing voice corresponding to the corrected pitch; step 19). On the other hand, if only the singing pitch is extracted (when step 11 is Y and step 12 is N), the sound output processing unit 500 emits the singing voice extracted in step 11 as is from the speaker 20 (emits the singing voice; step 20).

[0048] The karaoke device K repeats the processes from step 11 to step 20 at predetermined intervals until the karaoke performance of the selected song X is finished (until step 21 is Y).

[0049] Specifically, suppose that karaoke machine K starts playing karaoke music for song X, and the singer starts singing karaoke music for song X using microphone 40 in time with the karaoke music. In this example, the predetermined timing is 20 msec, the first value is 0.2, and the second value is 0.8.

[0050] The first extraction unit 100 sets the start time of the karaoke performance of song X as 0 and extracts the singing pitch from the singer's vocals every 20 msec. The second extraction unit 200 also sets the start time of the karaoke performance of song X as 0 and extracts the reference pitch from the reference data of song X every 20 msec.

[0051] Referring to Figure 4, at timing T1 (15980 msec), the first extraction unit 100 does not extract the singing pitch. In this case, the first extraction unit 100 outputs a signal to the setting unit 300 indicating that the singing pitch was not extracted.

[0052] Also, referring to Figure 4, at timing T1 (15980 msec), the second extraction unit 200 does not extract the reference pitch. In this case, the second extraction unit 200 outputs a signal to the setting unit 300 and the sound emission processing unit 500 indicating that the reference pitch was not extracted.

[0053] The setting unit 300 sets the correction ratio CR1 at timing T1 to "1" in response to the output signals indicating that the singing pitch was not extracted and the reference pitch was not extracted.

[0054] Furthermore, since the singing pitch and reference pitch are not extracted, the calculation unit 400 does not calculate the corrected pitch at timing T1. Also, since the singer is not emitting singing voice, the sound emission processing unit 500 does not emit singing voice.

[0055] Next, at timing T2 (16000 msec), the first extraction unit 100 does not extract the singing pitch. In this case, the first extraction unit 100 outputs a signal to the setting unit 300 indicating that the singing pitch was not extracted.

[0056] Meanwhile, at timing T2 (16000 msec), the second extraction unit 200 extracts the reference pitch RP2 (440 Hz). In this case, the second extraction unit 200 outputs the reference pitch RP2 to the setting unit 300 and the calculation unit 400.

[0057] The setting unit 300 sets the correction ratio CR2 at timing T2 to "1" according to the output signal indicating that the singing pitch was not extracted and the reference pitch RP2.

[0058] Furthermore, since the singing pitch is not extracted, the calculation unit 400 does not calculate the corrected pitch at timing T2. Also, since the singer is not emitting singing voice, the sound emission processing unit 500 does not emit singing voice.

[0059] Next, at timing T3 (16020 msec), the first extraction unit 100 extracts the singing pitch SP3 (410 Hz). In this case, the first extraction unit 100 outputs the extracted singing pitch SP3 to the setting unit 300, the calculation unit 400, and the sound output processing unit 500.

[0060] Meanwhile, at timing T3 (16020 msec), the second extraction unit 200 extracts the reference pitch RP3 (440 Hz). In this case, the second extraction unit 200 outputs the reference pitch RP3 to the setting unit 300 and the calculation unit 400.

[0061] The setting unit 300 calculates the pitch ratio PR3 (1.073171) from the output singing pitch SP3 and reference pitch RP3. The setting unit 300 multiplies the pitch ratio PR3 by a first predetermined value (0.2) to obtain the first value FV3 (0.214634). The setting unit 300 also multiplies the correction ratio "1" set at the timing T2, which is one timing before timing T3, by a second predetermined value (0.8) to obtain the second value SV3 (0.8). The setting unit 300 sets the correction ratio CR3 (1.014634146) at timing T3 by adding the first value FV3 and the second value SV3. The setting unit 300 outputs the set correction ratio CR3 to the calculation unit 400.

[0062] The calculation unit 400 calculates the corrected pitch CP3 (approximately 416Hz) at timing T3 by multiplying the singing pitch SP3 (410Hz) by the set correction ratio CR3 (1.014634146). The calculation unit 400 outputs the calculated corrected pitch CP3 to the sound output processing unit 500.

[0063] The sound output processing unit 500 emits singing voice corresponding to the corrected pitch CP3 from the speaker 20.

[0064] Next, at timing T4 (16040 msec), the first extraction unit 100 extracts the singing pitch SP4 (411 Hz). In this case, the first extraction unit 100 outputs the extracted singing pitch SP4 to the setting unit 300 and the calculation unit 400.

[0065] Meanwhile, at timing T4 (16040 msec), the second extraction unit 200 extracts the reference pitch RP4 (440 Hz). In this case, the second extraction unit 200 outputs the reference pitch RP4 to the setting unit 300 and the calculation unit 400.

[0066] The setting unit 300 calculates the pitch ratio PR4 (1.07056) from the output singing pitch SP4 and reference pitch RP4. The setting unit 300 multiplies the pitch ratio PR4 by a first predetermined value (0.2) to obtain the first value FV4 (0.214112). The setting unit 300 also multiplies the correction ratio "1.014634146" set at the timing T4 immediately preceding timing T4 by a second predetermined value (0.8) to obtain the second value SV4 (0.811707317). The setting unit 300 sets the correction ratio CR4 (1.025819239) at timing T4 by adding the first value FV4 and the second value SV4. The setting unit 300 outputs the set correction ratio CR4 to the calculation unit 400.

[0067] The calculation unit 400 calculates the corrected pitch CP4 (approximately 422Hz) at timing T4 by multiplying the singing pitch SP4 (411Hz) by the set correction ratio CR4 (1.025819239). The calculation unit 400 outputs the calculated corrected pitch CP4 to the sound output processing unit 500.

[0068] The sound output processing unit 500 emits the singing voice corresponding to the corrected pitch CP4 from the speaker 20.

[0069] The karaoke device K repeats the above process until timing T47 shown in Figure 4.

[0070] As is clear from the above, the karaoke device K according to this embodiment includes a first extraction unit 100 that extracts the singing pitch from the singer's singing voice at predetermined timings from the start of karaoke performance of a song, a second extraction unit 200 that extracts a reference pitch from the song's reference data at predetermined timings, and a setting unit 300 that sets a correction ratio for correcting the singing pitch extracted at predetermined timings, wherein when the singing pitch and the reference pitch are extracted at a certain timing, a first value obtained by multiplying the pitch ratio, which is the ratio of the reference pitch to the singing pitch, by a first predetermined value, and a second value obtained by multiplying the correction ratio set at the timing immediately preceding that timing by a second predetermined value, thereby, The system includes: a setting unit 300 that sets a correction ratio at a given timing, and if the singing pitch is not extracted at a given timing or if only the singing pitch is extracted, sets the correction ratio at that timing to 1; a calculation unit 400 that calculates a corrected pitch by correcting the singing pitch extracted at a predetermined timing, and if the singing pitch and reference pitch are extracted at a given timing, calculates the corrected pitch at that timing by multiplying the singing pitch by the set correction ratio at that timing; and a sound emission processing unit 500 that emits the singing sound corresponding to the calculated corrected pitch from the speaker 20.

[0071] With this karaoke device K, a correction ratio is calculated based on the extracted singing pitch and reference pitch to correct the singing pitch, and the singing voice based on the corrected pitch can be emitted. Therefore, compared to conventional methods of correcting the voice signal of a karaoke singer based on the frequency information of the singing melody, a more natural singing voice can be emitted without emitting an unnatural singing voice that is clearly identifiable as having been corrected. In other words, with the karaoke device K according to this embodiment, when the singing pitch extracted from the singing voice is corrected, a singing voice with less unnaturalness can be emitted.

[0072] Furthermore, the setting unit 300 in the karaoke device K according to this embodiment can use a value greater than 0 and less than 1 as the first predetermined value, and a value obtained by subtracting the first predetermined value from 1 as the second predetermined value. With such a karaoke device K, the singing pitch correction is performed gradually, so it is possible to emit singing sounds that sound less unnatural. Moreover, the setting unit 300 in the karaoke device K according to this embodiment can use a value close to 0 as the first predetermined value. With such a karaoke device K, it is possible to emit singing sounds that sound even less unnatural.

[0073] <Second Embodiment> Next, a karaoke device according to the second embodiment will be described with reference to Figures 5 and 6. In this embodiment, an example will be described in which a diagram showing the relationship between the extracted singing pitch, the extracted reference pitch, and the calculated correction pitch in chronological order will be displayed as in the first embodiment. Detailed explanations of configurations similar to those in the first embodiment will be omitted.

[0074] ==Karaoke Equipment== [Control means] The control means 10e performs various controls on the karaoke machine K. The control means 10e includes a CPU and memory (neither of which are shown in the figure). The CPU realizes various functions by executing programs stored in the memory.

[0075] In this embodiment, the CPU executes a program stored in memory, thereby enabling the control means 10e to function as a first extraction unit 100, a second extraction unit 200, a setting unit 300, a calculation unit 400, a sound emission processing unit 500, and a display processing unit 600 (see Figure 5).

[0076] (Display processing unit) The display processing unit 600 causes the display means to display a diagram showing the relationship between the extracted singing pitch, the extracted reference pitch, and the calculated correction pitch in chronological order.

[0077] In this embodiment, the first extraction unit 100 outputs the extracted singing pitch to the display processing unit 600, the second extraction unit 200 outputs the extracted reference pitch to the display processing unit 600, and the calculation unit 400 outputs the calculated correction pitch to the display processing unit 600.

[0078] The diagram showing the relationship between the extracted singing pitch, the extracted reference pitch, and the calculated corrected pitch in a time series is not particularly limited, as long as it directly or indirectly shows the time series relationship between the singing pitch, the reference pitch, and the corrected pitch.

[0079] For example, as shown in the example of the first embodiment, when the singing pitch, reference pitch, and corrected pitch are obtained from a predetermined timing T1 to a predetermined timing T47, the display processing unit 600 generates transition diagram data showing the transition state of each pitch, with the horizontal axis representing time every 20 msec and the vertical axis representing frequency (Hz), and can display a transition diagram based on this transition diagram data on the display screen of the display device 30 (see Figure 6).

[0080] Alternatively, the display processing unit 600 can show the degree of correction at each predetermined timing (the difference between the corrected pitch and the singing pitch) using a level meter or graph with "singing pitch as 0 and reference pitch as 10".

[0081] As is clear from the above, the karaoke device K according to this embodiment has a display processing unit 600 that displays a diagram showing the relationship between the extracted singing pitch, the extracted reference pitch, and the calculated correction pitch in chronological order on a display means. With such a karaoke device K, it is possible to visually understand how the singing pitch is being corrected.

[0082] <Other> It is also possible to supply the program to a computer using a non-transitory computer-readable medium (with an executable program thereon) on which the above program is stored. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), CD-ROMs (Read Only Memory), etc.

[0083] The above embodiments are presented as examples and do not limit the scope of the invention. The above configurations can be combined as appropriate, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. The above embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of symbols]

[0084] 100 First extraction section 200 Second extraction section 300 Setting section 400 Calculation Unit 500 Sound emission processing unit 600 Display Processing Unit K Karaoke machine

Claims

1. A first extraction unit extracts the singing pitch from the singer's vocals at predetermined intervals from the start of the karaoke performance of the song, A second extraction unit extracts a reference pitch from the reference data of the song at each predetermined timing, A setting unit for setting a correction ratio to correct the singing pitch extracted at a predetermined timing, wherein if the singing pitch and a reference pitch are extracted at a certain timing, the setting unit sets the correction ratio at that timing by adding a first value obtained by multiplying the pitch ratio, which is the ratio of the reference pitch to the singing pitch, by a first predetermined value, and a second value obtained by multiplying the correction ratio set at the timing immediately preceding that timing by a second predetermined value, and sets the correction ratio at that timing to 1 if the singing pitch is not extracted or if only the singing pitch is extracted at that timing, A calculation unit for calculating a corrected pitch obtained by correcting the singing pitch extracted at the predetermined timing, wherein when the singing pitch and reference pitch are extracted at a certain timing, the calculation unit calculates the corrected pitch at that timing by multiplying the singing pitch by a set correction ratio at that timing, A sound emission processing unit that emits the singing voice of a singer from a sound emission means, comprising a sound emission processing unit that emits the singing voice corresponding to the calculated corrected pitch from the sound emission means, A karaoke machine having the following features.

2. The karaoke apparatus according to claim 1, characterized in that the setting unit uses a value greater than 0 and less than 1 as the first predetermined value, and uses a value obtained by subtracting the first predetermined value from 1 as the second predetermined value.

3. The karaoke apparatus according to claim 2, characterized in that the setting unit uses a value close to 0 as the first predetermined value.

4. The karaoke device according to any one of claims 1 to 3, characterized in that it has a display processing unit that displays a diagram on a display means showing the relationship between the extracted singing pitch, the extracted reference pitch, and the calculated correction pitch in chronological order.

Citation Information

Patent Citations

  • Karaoke device

    JP1996234772A