Karaoke equipment
The karaoke device enhances user performance by recognizing and advising on optimal face directions for singing based on past scoring data, addressing the lack of vocalization ease guidance in existing systems.
Patent Information
- Application Number
- JP2022074092
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2042-04-28
AI Technical Summary
Existing karaoke devices fail to provide advice on vocalization ease based on facial direction, which varies among users and songs, limiting the potential for higher scores.
A karaoke device that recognizes the user's face direction during singing, evaluates the singing voice, and stores the associated score, then notifies the user of the optimal face orientation for better performance.
Enables users to recognize and replicate the facial orientation that leads to higher scores, thereby improving vocalization ease and scoring potential.
Smart Images

Figure 0007802604000002 
Figure 0007802604000003 
Figure 0007802604000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to a karaoke device. [Background technology]
[0002] Karaoke machines that give advice based on past evaluations of a user's singing have been proposed. One such karaoke machine is disclosed that visually informs a user of point-added sections and point-added trends based on past singing history (see, for example, Patent Document 1). The karaoke machine described in Patent Document 1 analyzes singing sections that have been added points in the past singing history, and displays advice information such as messages and marks based on the analysis results along with the singing sections. By visually informing the user of the existence of point-added sections and what singing techniques will result in added points, it becomes easier for the user to obtain a high score. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-68055 Summary of the Invention [Problem to be solved by the invention]
[0004] However, if the direction of the face changes while singing, the ease of vocalization changes depending on the degree of throat opening. While it is desirable to sing with an appropriate facial direction, the relationship between vocalization ease and facial direction differs for each user and song. The karaoke device described in Patent Document 1 can provide advice to users on singing techniques, but cannot provide advice on vocalization ease, so users cannot hope for a higher score.
[0005] An object of the present invention is to provide a karaoke machine that can inform a user of the face direction that makes it easier to sing. [Means for solving the problem]
[0006] The main invention to achieve the above object is to recognize the image of the user captured while singing a song and identify the face. Vertical angle level a scoring unit that evaluates the singing voice of the user input while singing a song and outputs a scoring value; and a user identifier of the user, a song identifier of the song, and a face of the user. Vertical angle level a storage unit that stores the score value of the singing voice in association with the singing voice score value; and a storage unit that stores the score value of the singing voice in association with the singing voice score value, and a storage unit that stores the score value of the singing voice in association with the singing voice score value. Highest or lowest score The face associated with Vertical angle level an extraction unit that extracts from the storage unit, and a face during singing extracted by the extraction unit Vertical angle level and a notification unit that notifies the user of the karaoke device. [Effects of the Invention]
[0007] According to the present invention, the facial orientation of a user captured while singing a song is stored together with a score. When the user attempts to sing the song again, the facial orientation at which the score meets a predetermined condition is extracted and notified to the user. This allows the user to recognize the facial orientation that is likely to earn a high score, i.e., makes it easier to produce a voice. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a configuration diagram of a karaoke system according to an embodiment of the present invention. [Figure 2] 1 is a functional block diagram of a karaoke device according to an embodiment of the present invention; [Figure 3] FIG. 4 is a flowchart showing an example of a processing operation of the karaoke apparatus according to the present embodiment. [Figure 4] FIG. 4 is a flowchart showing an example of a processing operation of the karaoke apparatus according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] The karaoke device of this embodiment will be described with reference to Figures 1 and 2. Figure 1 is a configuration diagram of the karaoke system of this embodiment. Figure 2 is a functional block diagram of the karaoke device of this embodiment. For the sake of convenience, the functional block diagram of Figure 2 shows functional blocks for realizing specific processing, but it is assumed that the device is equipped with the components that a karaoke device normally has.
[0010] As shown in Fig. 1, a karaoke system 1 includes a server device 2 installed in a data center or the like, and a karaoke device 10 installed in a karaoke store or the like. The server device 2 and the karaoke device 10 are connected to each other so that they can communicate with each other via a network 3. The server device 2 manages various pieces of information related to songs and users. A paired remote control device 14 is connected to the karaoke device 10, and when a user logs in to the karaoke device 10 by operating the remote control device 14, the server device 2 applies various pieces of information related to the user to the karaoke device 10.
[0011] In this case, by inputting a user ID and password into the login screen of the remote control device 14, the user ID and password are transmitted from the remote control device 14 to the server device 2 via the karaoke device 10. The server device 2 identifies the user information using a user database, and the server device 2 returns the user information to the karaoke device 10, thereby logging the user into the karaoke device 10. The karaoke device 10 stores the user information of the currently logged-in user. The user information includes a singing history including the score values for each song, the direction of the user's face, etc., which will be described later.
[0012] In addition to a remote control device 14, a monitor 11, a speaker 12, and a microphone 13 are connected to the karaoke device 10. The monitor 11 displays lyrics and captions along with background images in sync with the karaoke performance based on video signals and the like from the karaoke device 10. The microphone 13 converts the user's singing voice into a singing voice signal and inputs it to the karaoke device 10. The speaker 12 emits the user's singing voice along with the performance sound based on the performance sound signal from the karaoke device 10 and the singing voice signal from the microphone 13. The remote control device 14 accepts various operations for the karaoke device 10, such as reserving songs.
[0013] When the karaoke performance begins on the karaoke device 10, lyrics captions and background images are displayed on the monitor 11 in sync with the karaoke performance. In the karaoke device 10, the karaoke performance sound signal and the singing voice signal input from the microphone 13 are mixed by the mixer, and this mixed signal is amplified by the amplifier and output from the speaker 12. In this way, when the user sings along with the karaoke performance, the singing voice is output from the speaker 12 together with the performance sound. The karaoke device 10 also scores the user's singing ability, and the score is displayed on the monitor 11.
[0014] Furthermore, a camera 15 is connected to the karaoke device 10 to capture images of users while they are singing. If a singing stage is installed in the karaoke room, one camera 15 for capturing images of the singing stage may be installed in the karaoke room. Alternatively, multiple cameras 15 may be installed in the karaoke room to identify a singer from multiple users in the karaoke room. To identify a singer from multiple users included in the image captured by the camera 15, for example, a known technology can be used to identify the singer by recognizing the microphone 13 held by the user while singing.
[0015] Since the degree of throat opening changes depending on the direction of the face while singing, if the face is turned appropriately while singing, it becomes easier to produce sound and the user is more likely to receive a high score. However, the face direction that makes it easier to produce sound varies from person to person, and even for the same user, it varies from song to song. For this reason, the karaoke device 10 of this embodiment detects the face direction while singing a song from the image of the user captured by the camera 15. Then, the karaoke device 10 uses the face direction of the user when a high score is obtained to advise the user on the appropriate face direction when singing the same song again.
[0016] 2, karaoke device 10 is provided with registration unit 21, detection unit 22, scoring unit 23, storage unit 24, extraction unit 25, and notification unit 26. Registration unit 21 registers a song ID (song identifier) and a user ID (user identifier) in a reservation list. In this case, when a user selects a song using remote control device 14, a song details screen and an avatar of the logged-in user are displayed on remote control device 14. When the user selects their own avatar and presses the reservation button, the song ID and user ID are transmitted from remote control device 14 to karaoke device 10, and registration unit 21 registers the song ID and user ID in the reservation list.
[0017] For example, user U1 operates remote control device 14 to search for song X1, selects his or her own avatar as the singer, and presses the reservation button. Remote control device 14 transmits song ID "ID****X1" of song X1 and user ID "ID****U1" of user U1 to karaoke device 10. Registration unit 21 of karaoke device 10 registers song ID "ID****X1" of song X1 and user ID "ID****U1" of user U1 at the end of the reservation list. Song IDs are read out in order from the top of the reservation list, and karaoke performance begins based on the song data as the song IDs are read out.
[0018] The detection unit 22 recognizes images of the user captured while the song is being sung and detects the direction of the user's face. In this case, when a karaoke performance begins, images of the user captured by the camera 15 during the singing section are input to the detection unit 22 at predetermined time intervals based on information indicating singing and non-singing sections contained in the song data. The detection unit 22 analyzes the images of the users and detects the direction of the user's face. For example, when a karaoke performance begins based on the song data of song X1, images of user U1 are input from the camera 15 to the detection unit 22 every few seconds, and the detection unit 22 detects the direction of the user U1's face.
[0019] Alternatively, the image of the user may be input from camera 15 to detection unit 22 only during the period when the singing voice is being input from microphone 13, rather than during the singing section, and the facial orientation may be detected from the image of the user. The facial orientation detection method may be, for example, the method described in Japanese Patent Application Laid-Open No. 10-274516. In this detection method, the user's facial region is extracted from the image input from camera 15, and characteristic regions such as the user's eyes, nose, and mouth are also extracted. The facial orientation is then detected from the relationship between the center position of the facial region approximated to an ellipse, the area of the facial region approximated to an ellipse, and the center positions calculated from the local maximum points of each characteristic region.
[0020] The vertical angle level may be detected as the face orientation. An angle level of "0" is detected when the face is facing forward at "-7 degrees to 7 degrees." An angle level of "1" is detected when the face is facing upward at "8 degrees to 15 degrees," an angle level of "2" is detected when the face is facing upward at "16 degrees to 24 degrees," and an angle level of "3" is detected when the face is facing upward at "25 degrees or more." An angle level of "-1" is detected when the face is facing downward at "-8 degrees to -15 degrees," an angle level of "-2" is detected when the face is facing downward at "-16 degrees to -24 degrees," and an angle level of "-3" is detected when the face is facing downward at "-25 degrees or more." Furthermore, the detection unit 22 counts the number of detections for each face orientation in the count table shown in Table 1. [Table 1]
[0021] The scoring unit 23 evaluates the user's singing voice input while singing a song and outputs a score. In this case, the scoring unit 23 reads reference data from the song data, and the singing voice is input to the scoring unit 23 from the microphone 13. The pitch of the note and the pitch of the singing voice are compared during the note-on period (the period during which the note is sounded) of the reference data in the song data, and the singing voice for each note is judged as pass / fail depending on the degree of match between the pitch of the note and the pitch of the singing voice. The user's singing voice is scored based on the judgment results (evaluation results) for all notes.
[0022] More specifically, as described in JP 2005-049410 A, the singing frequency and the reference frequency are compared at each sample timing (e.g., every 20 milliseconds), and the difference (cent value) between the singing frequency and the reference frequency is calculated. A note whose difference falls within an acceptable range (e.g., within ±50 cents of the note pitch) a predetermined number of times or more (e.g., one or more times) during the note-on period is determined to be a passing note. On the other hand, a note whose difference falls within the acceptable range less than the predetermined number of times during the note-on period is determined to be a failing note. A score is calculated based on the ratio of the number of passing notes to the total number of notes. Note that if singing techniques such as vibrato or fisting are detected from the singing frequency, points may be added to the calculated score based on the number of times they are detected.
[0023] The storage unit 24 stores the user ID of the user, the song ID of the song, the user's facial direction, and the score value of the singing voice in association with each other. Here, the facial direction that is detected most frequently is stored. For example, when user U1 finishes singing, the song ID "ID****X1" and user ID "ID****U1" at the top of the reservation list, the angle level "1" (8 to 15 degrees upward) that is detected most frequently in the count table, and the score value "85 points" output from the scoring unit 23 are stored in association with each other. When user U1 logs out, each of these pieces of information is sent to the server device 2 and managed as user information.
[0024] When a user sings a song, the extraction unit 25 extracts from the storage unit 24 a facial direction associated with a score that satisfies a predetermined condition based on the user ID of the user and the song ID of the song. In this embodiment, the predetermined condition is set to be "the highest score." For example, when user U1 logs in to the karaoke device 10 again, the server device 2 transmits the user information of user U1 to the karaoke device 10. When user U1 reserves song X1 again, the server device 2 searches the user information for the angle level and score associated with the user ID "ID****U1" and song ID "ID****X1," and extracts the angle level "0" (-7 degrees to 7 degrees forward) that corresponds to the highest score "89" as the facial direction.
[0025] The notification unit 26 notifies the user of the facial orientation during singing that has been extracted by the extraction unit 25. For example, if the extraction unit 25 extracts angle level "0" as the facial orientation of user U1, the monitor 11 displays a message at the start of the karaoke performance saying, "You can get a high score for this song if you sing facing forward." Furthermore, if the scoring unit 23 judges the notes to be unsuccessful a predetermined number of times during singing, the monitor 11 displays a message saying, "Please face forward when singing." In this way, when user U1 sings song X1, the facial orientation that is likely to earn a high score is notified to user U1.
[0026] The processing of each part of the karaoke device 10 may be realized by software using a processor, or may be realized by a logic circuit (hardware) formed in an integrated circuit or the like. When a processor is used, the processor reads and executes programs stored in memory to perform various processes. As the processor, for example, a CPU (Central Processing Unit) is used. Furthermore, the memory is configured with one or more storage media such as ROM (Read Only Memory) and RAM (Random Access Memory) depending on the application.
[0027] The processing operation of the karaoke device of this embodiment will be described with reference to Figures 3 and 4. Figures 3 and 4 are flow diagrams showing an example of the processing operation of the karaoke device of this embodiment. Note that the symbols in Figures 1 and 2 will be used as appropriate in the description. Also, the following flow diagram is merely an example and can be modified as appropriate.
[0028] As shown in Fig. 3, when a user logs in, user information managed by the server device 2 is applied to the karaoke device 10 (step S01). When a song is reserved by the user, the registration unit 21 registers the user ID and song ID at the end of the reservation list (step S02). When the karaoke performance of the song with this song ID begins (step S03), the camera 15 captures an image of the user, and the detection unit 22 detects the direction of the user's face at predetermined time intervals during the singing section (step S04). As described above, the direction of the user's face is detected by an angle level indicating the angle of the face when facing vertically.
[0029] Each time the detection unit 22 detects the direction of the user's face, the number of detections for each direction of the face is counted and set in a count table (step S05). The processes of steps S04 and S05 are repeated until the karaoke performance is completed (No in step S06). When the karaoke performance is completed (Yes in step S06), the scoring unit 23 scores the user's singing voice (step S07). At this time, the singing voice is scored based on the comparison result between the pitch of the notes in the reference data for the song and the pitch of the singing voice.
[0030] The storage unit 24 stores the facial direction and the score value in association with the user ID and song ID (step S08). At this time, the facial direction that has been detected the most is stored in association with the user ID, song ID, and score value. The number of times the facial direction has been detected, which is set in the count table, is reset after the storage unit 24 stores the user ID, song ID, facial direction, and score value. When the user performs a logout operation, the user ID, song ID, facial direction, and score value are transmitted to the server device 2 (step S09). The server device 2 manages the user ID, song ID, facial direction, and score value as user information. In the karaoke device 10, the user information is deleted from the storage unit 24 after the server device 2 receives the user information.
[0031] As shown in Fig. 4, when the user logs in again, the karaoke device 10 receives user information from the server device 2 (step S11). The user information includes a user ID, a facial orientation associated with a song ID, and a score value, and the storage unit 24 stores the facial orientation and score value in association with the user ID and song ID (step S12). In this way, the user information managed by the server device 2 is applied to the karaoke device 10. When the user reserves the same song as last time, the registration unit 21 registers the user ID and song ID at the end of the reservation list (step S13).
[0032] When the karaoke performance of the same song as the previous one starts (step S14), the extraction unit 25 acquires the user ID and song ID from the reservation list, and extracts the facial direction and score value associated with the user ID and song ID from the storage unit 24 (step S15). At this time, the facial direction associated with the highest score value is extracted. The notification unit 26 then notifies the user of the facial direction using the monitor 11 or the like (step S16). In this case, the notification unit 26 may notify the user of the facial direction at the start of the karaoke performance, or may notify the user of the facial direction while the user is singing.
[0033] As described above, according to the karaoke device 10 of this embodiment, the facial orientation of the user captured while singing a song is stored together with the score. When the user attempts to sing the song again, the facial orientation at which the score that satisfies the predetermined condition is obtained is extracted and notified to the user. This allows the user to recognize the facial orientation that is likely to earn a high score, i.e., makes it easier to produce a good voice.
[0034] The notification unit may notify the user of the direction of the user's face by controlling the display position of the lyrics caption according to the direction of the face extracted by the extraction unit. When the direction of the face is angle level "3" (25 degrees or more upward), the lyrics caption is displayed at the top of the monitor, and the lyrics caption is moved closer to the center of the monitor as the direction of the face is angle level "2" (16 degrees to 24 degrees upward) and angle level "1" (8 degrees to 15 degrees upward) in this order. When the direction of the face is angle level "0" (-7 degrees to 7 degrees forward), the lyrics caption is displayed at the center of the monitor. When the direction of the face is angle level "-1" (-8 degrees to -15 degrees upward) and angle level "-2" (-16 degrees to -24 degrees upward), the lyrics caption is moved closer to the bottom of the monitor as the direction of the face is angle level "-3" (25 degrees or more upward). The display position of the lyrics caption displayed on the monitor changes depending on the direction of the face, thereby making it possible to inform the user of the direction of the face that is likely to receive a high evaluation.
[0035] In this embodiment, the extraction unit extracts the face orientation associated with the highest score, but the extraction unit may be configured to extract a face orientation associated with a score that satisfies a predetermined condition. For example, the extraction unit may extract a face orientation associated with the lowest score. In this case, the notification unit can notify the user that the face orientation is likely to result in a low score.
[0036] In the present embodiment, the detection unit detects the angle level in the vertical direction as the face orientation, but the detection unit may be configured to be able to detect the face orientation of the user. For example, the detection unit may detect the angle levels in the vertical direction and the horizontal direction as the face orientation.
[0037] In addition, in this embodiment, the notification unit notifies the face direction by displaying it on the monitor, but the notification unit may be configured to notify the face direction in any way. For example, before the singing starts, a message such as "You can get a high score by singing this song while facing forward" may be audibly notified by a speaker.
[0038] In the above embodiment, a face direction notification function may be added by installing a program in the karaoke machine. The program is stored in a storage medium. The storage medium is not particularly limited, and may be a non-transitory storage medium such as an optical disk, a magneto-optical disk, or a flash memory.
[0039] Although the present embodiment has been described, other embodiments may be obtained by combining the above-described embodiments and modifications in whole or in part.
[0040] Furthermore, the technology of the present invention is not limited to the above-described embodiments, and may be variously modified, substituted, or altered within the scope of the spirit of the technical idea. Furthermore, if the technical idea can be realized in a different way due to technological advances or other derived technologies, it may be implemented using that method. Therefore, the claims cover all embodiments that may fall within the scope of the technical idea. [Explanation of symbols]
[0041] 10: Karaoke equipment 22: Detection unit 23: Scoring Department 24: Storage section 25:Extraction part 26: Information Department
Claims
1. a detection unit that recognizes an image of the user captured while singing a song and detects the vertical angle level of the face; a scoring unit that evaluates the singing voice of the user input while singing the song and outputs a score; a storage unit that stores a user identifier of a user, a song identifier of a song, a vertical angle level of the user's face, and a score value of the singing voice in association with each other; an extraction unit that extracts, when a user sings a song, from the storage unit, a vertical angle level of the face that is associated with the highest or lowest score value based on the user identifier of the user and the song identifier of the song; a notification unit that notifies a user of the vertical angle level of the face during singing extracted by the extraction unit.
2. the detection unit detects the vertical angle level of the face of the user at predetermined time intervals and counts the number of detections for each vertical angle level of the face, 2. The karaoke device according to claim 1, wherein the storage unit stores the vertical angle level of the face that is most frequently detected in association with the user identifier of the user, the song identifier of the song, and the score value of the singing voice.
Citation Information
Patent Citations
Karaoke machine
JP2005242230A
Karaoke system having grading information display function
JP2005345555A
Karaoke device
JP2006251697A
Karaoke system
JP2012078526A
Karaoke system
JP2014071248A