Image generation device
The image generation device addresses the lack of motivation in elderly users by generating a virtual ideal dance video synchronized with their actual performance, enhancing engagement and skill development.
Patent Information
- Application Number
- JP2021542972
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-08-29
- Filing Date
- 2020-08-26
- Publication Date
- 2025-09-18
- Estimated Expiration
- 2040-08-26
AI Technical Summary
Existing karaoke systems fail to motivate elderly individuals with low physical abilities to continue dancing due to lack of engagement and feedback, leading to early discontinuation.
An image generation device that analyzes a user's dance movements, compares them to ideal skeletal poses, and generates a virtual ideal video synchronized with the user's actual performance, providing feedback and motivation through score calculation and emotional evaluation.
Enhances user motivation by presenting a virtual ideal dance video, improving dance skills and emotional engagement, thereby encouraging continued participation in dance activities.
Smart Images

Figure 0007741729000002 
Figure 0007741729000003 
Figure 0007741729000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image generation device. [Background technology]
[0002] Karaoke machines are known that acquire position information of a singer while singing karaoke and score the singer's choreography by comparing it with reference choreography data (see, for example, Patent Documents 1 and 2). Karaoke machines are also known that display animated images of the choreography of a karaoke song or that display images of people superimposed on the video being played while singing (see, for example, Patent Documents 3 to 5).
[0003] Patent document 6 describes a moving image generation system that obtains movement parameters for each part that makes up a 3D model of the human body from a person in an acquired 2D moving image, generates a 3D model, extracts texture data corresponding to each part, sets a viewpoint, and generates 2D moving images from the 3D model by interpolation using the texture data.
[0004] Non-Patent Document 1 describes a method for detecting a two-dimensional posture of a person in an image by estimating the positions of the joints using AI (artificial intelligence). Non-Patent Document 2 describes a method for generating a video of another person dancing from a video of one person dancing. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 11-212582 [Patent Document 2] International Publication No. 2014 / 162787 [Patent Document 3] Japanese Patent Application Publication No. 11-133987 [Patent Document 4] Japanese Patent Application Laid-Open No. 2000-209500 [Patent Document 5] Japanese Patent Application Laid-Open No. 2001-42880 [Patent Document 6] Japanese Patent Application Laid-Open No. 2002-269580 [Non-patent literature]
[0006] [Non-Patent Document 1] Zhe Cao, Tomas Simon, Shih-En Wei, Yaser Sheikh, “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”, [online], November 24, 2016, [Retrieved August 19, 2019], Internet<URL:https: / / arxiv.org / abs / 1611.08050> [Non-patent document 2] Caroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. Efros, “Everybody Dance Now”, [online], August 22, 2018, [Retrieved August 19, 2019], Internet<URL:https: / / arxiv.org / abs / 1808.07371> Summary of the Invention
[0007] As the population ages, there has been an increase in services such as dance karaoke and locomotion exercises that allow people to maintain their health while having fun. However, elderly people have low physical abilities, and many of them do not continue with dance karaoke for long because they cannot dance well or do not look cool, and so they quit after participating a few times. While devices that score dance, such as karaoke machines that score choreography, are known, there is an issue that simply scoring does not motivate participants and does not lead to continuation. In order to maintain health, it is necessary to come up with ways to increase participants' motivation to participate in dance karaoke and other activities.
[0008] The present invention aims to increase a subject's motivation to participate in dancing by presenting the subject dancing with an ideal dance video as a virtual video of the subject himself / herself.
[0009] The present invention provides an image generation device having a memory unit that stores ideal skeletal pose information indicating a series of skeletal poses corresponding to movements made at ideal predetermined timings in accordance with a predetermined rhythm for each predetermined rhythm; a recording unit that records actual video of a subject making movements at predetermined timings in accordance with the predetermined rhythm that is played back; a skeletal pose analysis unit that extracts skeletal pose information indicating a series of skeletal poses corresponding to movements made by the subject at the predetermined timings from a group of still images that constitute the actual video; a model generation unit that generates a learning model of an image of the subject corresponding to the skeletal poses based on the group of still images and the skeletal pose information; and an image generation unit that generates and outputs a virtual ideal video, which is an image of the subject making movements at predetermined timings that match the ideal skeletal pose information, based on the learning model and the ideal skeletal pose information.
[0010] The image generating device may further have a difference determination unit that determines the difference in joint angles in the skeleton poses between the skeleton pose information and the ideal skeleton pose information, and the image generating unit may output a virtual ideal image when the difference is equal to or greater than a reference value.
[0011] The image generation unit may select a section within a predetermined rhythm corresponding to an action performed by the subject at a predetermined timing based on the magnitude of the difference, and play back the actual image and the virtual ideal image in that section in synchronization.
[0012] The difference determination unit may take the difference for each of the multiple joints, and the image generation unit may select some of the joints based on the magnitude of the difference, enlarge the surrounding area of those joints, and play back the actual image and the virtual ideal image in synchronization.
[0013] The image generating device may further include a score calculation section that calculates a score for the action performed by the subject at a predetermined timing according to the magnitude of the difference.
[0014] The image generating device may further include a pulse wave analysis unit that extracts the subject's pulse wave signal from time series data that indicates the subject's skin color in the actual image and calculates an index that indicates the degree of fluctuation in the pulse wave interval, and an emotion determination unit that determines, based on the index, whether the subject's emotion is a negative emotion associated with brain fatigue, anxiety, or depression, or a positive emotion associated with no brain fatigue, anxiety, or depression, and calculates an emotional score for the subject according to the frequency of occurrence of the positive emotion during an action performed at a predetermined timing.
[0015] The image generation device may further have a physique correction unit that calculates the ratio of the torso and leg lengths of the subject in the actual image and corrects the torso and leg lengths of each skeletal pose in the ideal skeletal pose information to match that ratio, and the image generation unit may generate a virtual ideal image based on the learning model and the corrected ideal skeletal pose information.
[0016] The image generation device may further have an image extraction unit that extracts some still images from the group of still images based on predetermined criteria regarding the angles or positions of the joints in the skeletal poses corresponding to each still image, and the model generation unit may generate a learning model based on the skeletal poses corresponding to some of the still images from among the some of the still images and the skeletal pose information.
[0017] The image generating device may further include an image processing unit that executes generation of the virtual ideal image in the image generating unit.
[0018] The image generating device may further include a communication interface that receives an actual image of the subject from an external terminal device and transmits a virtual ideal image of the subject to the external terminal device.
[0019] In the image generating device, the predetermined rhythm may be generated by a piece of music.
[0020] In the image generating device, the movement performed at a predetermined timing may be a dance movement.
[0021] According to the image generating device described above, by presenting an ideal dance image to a subject dancing as a virtual image of the subject himself / herself, the subject's motivation to participate in dancing can be increased. [Brief explanation of the drawings]
[0022] [Figure 1] 1 is an overall configuration diagram of a karaoke system 1. FIG. [Figure 2] FIG. 1 is a functional block diagram of a karaoke system 1. [Figure 3] FIG. 2 is a diagram for explaining the function of an analysis unit 30. [Figure 4] FIG. 2 is a diagram for explaining the function of a generation unit 40. [Figure 5] FIG. 2 is a diagram for explaining the function of an image extraction unit 41. [Figure 6] FIG. 2 is a diagram for explaining the function of an image extraction unit 41. [Figure 7] FIG. 10 is a diagram illustrating an example of ideal skeleton pose information. [Figure 8] FIG. 2 is a diagram for explaining the function of a video generating unit 44. [Figure 9] Graph (A) shows an example of the waveform of the pulse wave signal PW, graph (B) shows an example of the degree of fluctuation in pulse wave intervals, and graph (C) shows an example of fluctuations in the positive and negative emotions of the subject. [Figure 10] 10A to 10C are diagrams showing examples of displaying dance scores and mentality scores. [Figure 11] 4 is a flowchart showing an example of the operation of the karaoke system 1. [Figure 12] FIG. 10 is a diagram for explaining the function of the generation unit 40 when there are multiple subjects. [Figure 13] FIG. 10 is an overall configuration diagram of an image generating device according to a second embodiment of the present disclosure. [Figure 14] FIG. 10 is an overall configuration diagram of an image generating device according to a third embodiment of the present disclosure. [Figure 15] FIG. 10 is an overall configuration diagram of an image generating device according to a fourth embodiment of the present disclosure. [Figure 16]10 is a flowchart for explaining an operation procedure of an image generation device according to a fourth embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0023] Hereinafter, an image generating device according to an embodiment of the present disclosure will be described with reference to the drawings. However, it should be understood that the present invention is not limited to the drawings or the embodiments described below. In the following description, a predetermined rhythm is generated by music, and a dance movement is described as an example of a movement performed at a predetermined timing in accordance with the predetermined rhythm.
[0024] This video generation device generates an ideal dance video of the subject dancing, in which the subject performs the same movements and poses as a dance instructor (or teacher), and presents it to the subject. This virtually generated ideal dance video is hereinafter referred to as a "virtual ideal video." In particular, the video generation device calculates a score for the subject's dance movements based on the subject's skeletal poses during the dance movements, and if the score is low, it determines that the dancer is not dancing well and presents the virtual ideal video. By showing the subject a video of themselves rather than a video of their instructor showing how they should dance, the subject's motivation to continue dancing is improved. Furthermore, the video generation device determines the subject's emotional category (whether they are feeling mental fatigue, anxiety, or depression) from facial images of the subject while dancing, and by evaluating the subject's emotional aspects as well, it enables the subject to identify songs that they can dance to enjoyably and well.
[0025] The following description will be given using an example of a dance performed in a karaoke scene, but this image generation device can be used in any scene where a dance is practiced, and the type of dance is not particularly limited.
[0026] [First embodiment] FIG. 1 is an overall configuration diagram of a karaoke system 1 constituting a video generation device according to a first embodiment. FIG. 2 is a functional block diagram of the karaoke system 1. The karaoke system 1 is composed of a karaoke device 2, an evaluation device 3, and a camera 4. The karaoke device 2 is composed of a main body 2A, a speaker 2B, a monitor 2C, and a remote control 2D. Furthermore, a projector (not shown) may be provided so that the video to be output to the monitor 2C is output to the projector.
[0027] The main body 2A has, as its functional blocks, a music database 11, a music selection unit 12, a video selection unit 13, and a playback unit 14. The music database 11 is a storage unit that stores music and videos of karaoke songs. The music selection unit 12 and the video selection unit 13 select music and videos stored in the music database 11 in response to the user's operation of the remote control 2D. The playback unit 14 outputs the music selected by the music selection unit 12 to the speaker 2B, and outputs the video selected by the video selection unit 13 to the monitor 2C, and plays them.
[0028] The remote control 2D is a terminal that allows the user to select songs and is also used to operate the evaluation device 3. The evaluation device 3 is an example of an image generation device, and the camera 4 captures images of the singer of the song and the subject performing dance movements to the song. FIG. 1 illustrates an example in which the evaluation device 3 and camera 4 are added to the karaoke device 2 as optional auxiliary systems, and in this case, the karaoke device 2 may be a commonly used one. However, all of the functions of the evaluation device 3 described below may be implemented as software and incorporated into the karaoke device 2, or the camera 4 may be one that is originally installed in the karaoke device.
[0029] The evaluation device 3 has, as its functional blocks, a recording unit 20, an analysis unit 30, a generation unit 40, an emotion evaluation unit 50, a display control unit 60, and a storage unit 70. Of these, the storage unit 70 is realized by a semiconductor memory or a hard disk, and the other functional blocks are realized by a computer program executed on a microcomputer including a CPU, a ROM, a RAM, and the like.
[0030] The recording unit 20 records the video by storing the data of the actual video captured by the camera 4 in the storage unit 70. When the subject (a participant in the dance karaoke) performs dance movements to the music played by the karaoke device 2, the actual video is recorded by the camera 4 and the recording unit 20.
[0031] The analysis unit 30 is composed of a skeleton pose analysis unit 31, a difference determination unit 32, and a score calculation unit 33. The analysis unit 30 analyzes the skeleton pose of the subject in the actual video captured by the camera 4, finds the difference between the ideal value and each joint angle in the skeleton pose, and calculates a score for the dance movement (hereinafter referred to as the dance score).
[0032] The skeletal pose analysis unit 31 recognizes the skeletal pose of the subject by applying, for example, a method described in Non-Patent Document 1 to each still image (for example, 30 still images per second for a 30 FPS video) that constitutes the actual video captured by the camera 4. The skeletal pose is a two-dimensional skeleton composed of multiple line segments representing the head, torso, right arm, left arm, right leg, and left leg, and information on their relative positions and relative angles is the skeletal pose information. Because the skeletal pose changes during dance movements, the skeletal pose information is defined in association with the elapsed time in the music being played. The skeletal pose analysis unit 31 extracts skeletal pose information indicating a series of skeletal poses corresponding to the subject's dance movements from the still images that constitute the actual video, and stores (preserves) still images that compare the actual video with the skeletal pose for each frame in the storage unit 70.
[0033] FIG. 3 is a diagram for explaining the function of the analysis unit 30. The symbol t represents elapsed time, symbols t0 and tN represent the start and end times of the music, respectively, and symbols t1 and t2 represent different times during the music playback. Symbol 80 represents an actual video of the subject, and the thick lines superimposed in the figure represent the skeletal pose recognized by the skeletal pose analysis unit 31. Symbols A to H represent the joint angles of the right shoulder, left shoulder, right elbow, left elbow, right hip, left hip, right knee, and left knee in the skeletal pose of the subject. Symbol 91 in FIG. 3 represents the skeletal pose of the instructor, and symbols A' to H' represent the joint angles in the skeletal pose of the instructor.
[0034] The difference determination unit 32 calculates the difference in joint angles between the skeletal pose information and the ideal skeletal pose information. The ideal skeletal pose information is linked to the music database 11 and indicates a series of skeletal poses corresponding to ideal dance movements by an instructor for the same music piece as the subject danced. The ideal skeletal pose information is defined in association with elapsed time in the music piece and is pre-stored in the storage unit 70 for each music piece. During music playback, the difference determination unit 32 calculates the difference between the subject's skeletal pose and the instructor's skeletal pose at the same time, for example, for eight joint angles indicated by symbols A through H in FIG. 3, at regular intervals, such as every second. Joint angles are the angles between line segments generated by skeletal recognition. For example, the angle difference for right shoulder A can be calculated as ΔA = A - A'. If video of the subject and instructor is captured from the same position and angle, the skeletal poses can be compared based on the angle difference in the two-dimensional image.
[0035] The score calculation unit 33 calculates the dance score of the subject according to the magnitude of the difference found by the difference determination unit 32, and displays the value on the monitor 2C of the karaoke device 2. For example, the score calculation unit 33 may calculate a dance score of the subject according to the magnitude of the difference found by the difference determination unit 32, and display the value on the monitor 2C of the karaoke device 2. ave +ΔB ave +ΔC ave +ΔD ave +ΔE ave +ΔF ave +ΔG ave +ΔG ave )" to calculate the dance score. ave ~ΔH aveis the average value of the differences ΔA to ΔH found for the joint angles of the right shoulder, left shoulder, right elbow, left elbow, right hip, left hip, right knee, and left knee, respectively, indicated by symbols A to H in Figure 3, from the start time t0 to the end time tN of the song, and k is an appropriate coefficient. However, this calculation formula is just an example, and the dance score can be appropriately defined so that the smaller the difference between each joint angle, the higher the dance score.
[0036] FIG. 4 is a diagram illustrating the function of the generation unit 40. In FIG. 4, reference numerals 80 and 81 represent the actual video and skeletal pose of the subject, and reference numerals 90 and 91 represent the actual video and skeletal pose of the instructor. When the dance score calculated by the score calculation unit 33 is equal to or lower than a reference value, the generation unit 40 learns the relationship between the actual video 80 and skeletal pose 81 of the subject and generates a learning model. The generation unit 40 then corrects pre-stored ideal skeletal pose information to match the subject's physique, and generates a virtual ideal video 85, which is a video of the subject performing the same dance movements as the instructor, based on the learning model and the corrected ideal skeletal pose information.
[0037] As shown in FIG. 2, the generation unit 40 is composed of an image extraction unit 41, a model generation unit 42, a physique correction unit 43, and a video generation unit 44. The image extraction unit 41 extracts (selects) some still images from the still images stored in the storage unit 70 to be used by the model generation unit 42 to generate a learning model. Considering use in a karaoke scene, it is desirable to complete video generation in approximately 30 minutes at most. However, the number of still images constituting the video, even if each image is only a few minutes long, can be several thousand or more, and using all of them for training would require an enormous amount of time. Furthermore, since many of the still images constituting the video are similar, it is not necessary to use all of them for training; several dozen images are sufficient. Therefore, to speed up training, the image extraction unit 41 selects several dozen still images with significantly different skeletal poses.
[0038] 5 and 6 are diagrams for explaining the function of the image extraction unit 41. FIG. 5 shows a group of still images 82 constituting an actual video 80 of a subject. For example, in the case of a 30 FPS video, there are 30 still images per second. The skeleton pose analysis unit 31 recognizes a skeleton pose 81 for each frame of the video and generates a group of image data that compares the still images with the skeleton pose. Then, as shown in FIG. 6, the image extraction unit 41 selects some still images 83 from the group of still images 82 based on predetermined criteria regarding the angles or positions of the joints in the skeleton poses corresponding to each still image.
[0039] For example, the image extraction unit 41 first selects an arbitrary image and then selects images whose coordinates or joint angles of the right arm or right shoulder or right elbow differ from the selected image by a reference value or more. Similarly, the image extraction unit 41 selects images whose coordinates or joint angles differ by a reference value or more for the left arm, right foot, and left foot until a total of several tens of images are selected. Alternatively, the image extraction unit 41 may calculate the average values of the coordinates and joint angles of the right arm, left arm, right foot, and left foot every second, and select several tens of images whose coordinates and joint angles differ from the average values by a reference value or more. Alternatively, the image extraction unit 41 may randomly select still images from a group of still images, regardless of the joint angles or positions. The number of still images extracted by the image extraction unit 41 is, for example, 20, 25, or 30, and is appropriately selected depending on the processing power of the hardware of the evaluation device 3 so that the model generation unit 42 can complete learning within several tens of minutes.
[0040] The model generation unit 42 performs deep learning on the relationship between the skeletal pose and the subject's image based on some of the still images extracted by the image extraction unit 41 and their corresponding skeletal poses, and generates a learning model of the subject's image corresponding to the skeletal pose. To do this, the model generation unit 42 uses the open-source algorithm Pix2Pix, for example, in accordance with the method described in Non-Patent Document 2. Pix2Pix is a type of image generation algorithm that uses generative adversarial networks. It learns the relationship between two paired images and generates a paired image from a single image by interpolating the relationship taking into account the relationship. The model generation unit 42 stores the generated learning model in the memory unit 70.
[0041] The physique correction unit 43 calculates the ratio of the length of the subject's torso to the length of the legs (thighs and shins) in the actual image 80, and corrects the torso and leg lengths of each skeletal pose in the ideal skeletal pose information used when the image generation unit 44 generates the virtual ideal image 85, according to that ratio. For example, if the subject has longer legs relative to his or her torso than the instructor, the physique correction unit 43 lengthens only the leg lengths by the same ratio as those of the subject, while leaving the arms and torso unchanged, for each skeletal pose in the ideal skeletal pose information. If the subject and instructor are significantly different in height, the image generated by the image generation unit 44 may appear unnatural, as if stretched or compressed, so physique correction is performed to prevent this.
[0042] The video generation unit 44 generates a virtual ideal video of the subject based on the learning model generated by the model generation unit 42 and the ideal skeleton pose information (after correction by the physique correction unit 43) stored in the storage unit 70 for the same music as the subject danced to. To achieve this, the video generation unit 44, like the model generation unit 42, uses Pix2Pix in accordance with the method described in Non-Patent Document 2, for example. The virtual ideal video generated by the video generation unit 44 is a video of the subject performing dance movements that match the ideal skeleton pose information (i.e., the arms and legs moving in the same way as the instructor) and having the same face, physique, and clothing as the subject (a video in which everything except the movements has been changed from the instructor to the subject).
[0043] FIG. 7 is a diagram showing an example of ideal skeletal pose information. FIG. 8 is a diagram for explaining the function of the video generation unit 44. As shown in FIG. 7, the ideal skeletal pose information is a group of still images (image data group) 92 that compares still images of an actual video 90 of an ideal dance movement by an instructor with a skeletal pose 91. The group of still images 92 includes all still images corresponding to the dance movements from the beginning to the end of a song. The video generation unit 44 inputs the ideal skeletal pose information for the target song into a learning model, and generates a group of virtual still images 84 of the same skeletal pose as the instructor, as shown in FIG. 8, and converts this into a video to generate a virtual ideal video 85.
[0044] For example, the video generation unit 44 outputs the virtual ideal video when the dance score calculated by the score calculation unit 33 is less than the reference value, i.e., when the subject dances poorly and the joint angle difference calculated by the difference determination unit 32 is equal to or greater than the reference value. The video generation unit 44 plays only the virtual ideal video, or a comparison of the virtual ideal video and the actual video captured by the camera 4, on the monitor 2C of the karaoke device 2. However, the video generation unit 44 may display the virtual ideal video on the monitor 2C regardless of the dance score, and in this case, calculation of the dance score by the score calculation unit 33 may be omitted.
[0045] The video generation unit 44 may play back the virtual ideal video only for the portion of the song where there is a large difference from the instructor's skeletal pose, rather than the entire song. In this case, the video generation unit 44 may, for example, select a section of the song where the dance score calculated by the score calculation unit 33 is below a reference value, and synchronously play back the actual video and the virtual ideal video for that section. Alternatively, the video generation unit 44 may play back the virtual ideal video by enlarging only the periphery of a portion such as the arms or waist where there is a large difference from the instructor's skeletal pose. In this case, the video generation unit 44 may, for example, select the average value ΔA of the differences calculated by the score calculation unit 33 for the eight joint angles indicated by symbols A to H in FIG. 3. ave ~ΔH aveThe largest one of these may be selected, and the area around the joint may be enlarged, and the actual image and the virtual ideal image may be played back in synchronization.
[0046] As shown in Fig. 2, the emotion evaluation unit 50 is composed of a face recognition unit 51, a pulse wave analysis unit 52, and an emotion determination unit 53. The emotion evaluation unit 50 performs face recognition on the actual video captured by the camera 4 and detects the pulse wave from the skin color of the subject's face. Furthermore, while the subject is dancing (while music is being played and while the recording unit 20 is recording), the emotion evaluation unit 50 determines the subject's emotion category based on the degree of fluctuation in the pulse wave interval at regular intervals, for example, every 30 seconds, and calculates an emotional score (hereinafter referred to as a mental score).
[0047] The face recognition unit 51 analyzes the facial appearance of each still image constituting the actual video of the subject captured by the camera 4 using a contour detection algorithm or a feature point extraction algorithm, and identifies exposed skin areas such as the forehead as measurement sites. The face recognition unit 51 outputs time-series data indicating the skin color at the measurement sites to the pulse wave analysis unit 52.
[0048] First, pulse wave analysis unit 52 extracts the subject's pulse wave signal from the time-series data acquired from face recognition unit 51. For example, since there are many capillaries inside a person's forehead, an image of the face contains brightness change components synchronized with the subject's blood flow, and in particular, the pulse wave (blood flow change) is most reflected in the green light brightness change component of the image. Therefore, pulse wave analysis unit 52 extracts the pulse wave signal from the green light brightness change component of the time-series data using a band-pass filter that passes frequencies of approximately 0.5 to 3 Hz that are contained in a person's pulse wave.
[0049] FIG. 9(A) is a graph showing an example waveform of pulse wave signal PW. The horizontal axis t represents time (milliseconds), and the vertical axis A represents the amplitude strength of the pulse wave. As shown in FIG. 9(A), pulse wave signal PW has a triangular waveform that reflects fluctuations in blood flow due to cardiac pulsation, and the intervals between peak points P1-Pn at which blood flow is strongest are defined as pulse wave intervals d1-dn. Pulse wave analysis unit 52 detects peak points P1-Pn of the extracted pulse wave signal PW, calculates pulse wave intervals d1-dn in milliseconds, and generates time series data of the pulse wave intervals from pulse wave intervals d1-dn.
[0050] Figure 9(B) is a graph showing an example of the degree of fluctuation in pulse wave intervals. This graph is called a Lorenz plot, and it plots time-series data of pulse wave intervals on the coordinates (dn, dn-1) for n=1, 2, ..., with the horizontal axis representing pulse wave interval dn and the vertical axis representing pulse wave interval dn-1 (both in milliseconds). It is known that the degree of variation in the dots in the graph in Figure 9(B) reflects the subject's positive or negative emotions.
[0051] The pulse wave analysis unit 52 further uses the generated time series data, i.e., the coordinates (dn, dn-1) in the Lorenz plot of Figure 9(B), to calculate the maximum Lyapunov exponent λ, which is an index showing the degree of fluctuation in the subject's heartbeat interval, at regular intervals, such as every 30 seconds or every 32 pulse wave beats, using the following equation 1:
number
[0052] Based on the maximum Lyapunov exponent λ calculated by the pulse wave analysis unit 52, the emotion determination unit 53 determines, at the same regular intervals as the pulse wave analysis unit 52, whether the subject is experiencing a negative emotion, such as brain fatigue, anxiety, or depression, or a positive emotion, such as not experiencing brain fatigue, anxiety, or depression. The maximum Lyapunov exponent λ in Equation 1 takes a large negative value when the subject is experiencing a negative emotion, and takes a negative value greater than or equal to 0 or a small absolute value when the subject is experiencing a positive emotion. Therefore, the emotion determination unit 53 determines that the subject is experiencing a negative emotion when the calculated maximum Lyapunov exponent λ satisfies the following Equation 2, and determines that the subject is experiencing a positive emotion when λ does not satisfy Equation 2. λ≦λt...Equation 2 λt (<0) is an appropriate threshold value set for emotion determination. The emotion determination unit 53 stores (preserves) the determination result in the storage unit 70.
[0053] Instead of being limited to two stages of positive and negative emotions, the emotion determination unit 53 may classify the subject's emotion into, for example, four stages of "positive emotion (stress-free)," "active," "mild negative emotion (mild fatigue state)," and "negative emotion (fatigue state)" depending on the value of the maximum Lyapunov exponent λ. The emotion determination unit 53 may sequentially display the determined emotion stages of the subject on the monitor 2C while the subject is dancing.
[0054] Figure 9(C) is a graph showing an example of fluctuations in a subject's positive and negative emotions. The horizontal axis t is time (seconds), and the vertical axis λ is the maximum Lyapunov exponent. In the graph, the region R+ (λ>0) corresponds to positive emotions, the region R- (λ<0) corresponds to negative emotions, and the region R0 (λ≒0) between them corresponds to emotions intermediate between positive and negative. Assume that the subject performed a dance movement during a period T from time t1 to t2. In the example shown, out of 10 emotion judgments, one judgment was made as an emotion intermediate between positive and negative, and 9 judgments were made as a positive emotion.
[0055] The emotion determination unit 53 calculates the mentality rating of the subject according to the frequency of positive emotions occurring during recording by the recording unit 20, and displays the calculated value on the monitor 2C of the karaoke device 2. For example, the emotion determination unit 53 refers to the memory unit 70 and calculates the rate of occurrence of positive emotions, defined as "(number of times determined to be positive emotions) / (number of times determinations have been made by the emotion determination unit 53)", as the mentality rating. In the example of FIG. 9(C), the rate of occurrence of positive emotions is 90%, so the mentality rating is 90 points.
[0056] Even if the dance score is high, a low mentality score indicates that the subject is not enjoying the dance. Therefore, by presenting both the dance score and the mentality score, the subject can obtain an indicator for identifying songs that suit them and allow them to dance well without pain. One possible method for determining emotion classification would be to attach a sensor to a karaoke microphone and detect the singer's pulse wave while holding the microphone. However, this method is inappropriate for dancing, as the singer cannot hold the microphone. The emotion evaluation unit 50 can calculate the mentality score hands-free, making it applicable to dancing as well. Furthermore, since the subject's face is captured in the video of the dance movements, pulse waves can be detected from that video, so only one camera is required. Therefore, the emotion evaluation unit 50's method is advantageous in that it does not require a separate sensor for pulse wave detection.
[0057] In response to instructions from the recording unit 20, the analysis unit 30, the generation unit 40 and the emotion evaluation unit 50, the display control unit 60 outputs the actual video of the subject captured by the recording unit 20, the dance score calculated by the score calculation unit 33, the virtual ideal video generated by the video generation unit 44 and the mental score calculated by the emotion determination unit 53 to the karaoke device 2 and displays them on the monitor 2C.
[0058] The storage unit 70 is mainly composed of an ideal data storage unit 71 and a learning data storage unit 72. The ideal data storage unit 71 stores the above-mentioned ideal skeletal pose information for each song. If the karaoke device 2 is an online karaoke device, the ideal data storage unit 71 may be provided on a cloud, together with the music database 11, separately from the evaluation device 3 and the karaoke device 2. The learning data storage unit 72 stores image data groups (comparisons between actual video and skeletal poses) generated by the skeletal pose analysis unit 31 and information on the learning model generated by the model generation unit 42. The storage unit 70 also stores information necessary for the operation of the evaluation device 3, such as actual video of the subject captured by the recording unit 20 and the subject's emotion category determined by the emotion determination unit 53.
[0059] 10(A) to 10(C) are diagrams showing examples of displaying dance scores and mentality scores. These scores may be displayed numerically, as shown in FIG. 10(A), or as a graphic, such as a bar that becomes longer as the value increases, as shown in FIG. 10(B). FIG. 10(B) shows an example in which, in addition to the dance score and mentality score, a singing score and a total score for dance, mentality, and singing are calculated and displayed. While FIGS. 10(A) and 10(B) show examples of display after the dance motion has ended, as shown in FIG. 10(C), intermediate data on the dance score and mentality score may be displayed in pitch format in real time while the subject is dancing. Alternatively, the dance score and mentality score may be superimposed on an actual video or a virtual ideal video of the subject. Furthermore, calculation and display of the dance score and mentality score are not essential, and one or both of the score calculation unit 33 and the emotion evaluation unit 50 may be omitted.
[0060] 11 is a flowchart showing an example of the operation of the karaoke system 1. First, in response to the user's operation of the remote control 2D, the music selection unit 12 of the karaoke device 2 selects a dance karaoke song (S1), and the playback unit 14 starts playing the song (S2). At this time, the playback unit 14 may also display on the monitor 2C, for example, actual video of an instructor demonstrating ideal dance movements. As playback starts, the subject begins dancing, and the camera 4 starts taking video, and the recording unit 20 starts recording (S3).
[0061] The skeleton pose analysis unit 31 begins analyzing the skeleton pose of the actual video captured by the camera 4 (S4), and stores still images comparing the actual video with the skeleton pose for each frame in the learning data storage unit 72 (S5). The difference determination unit 32 calculates the difference in each joint angle in the skeleton pose between the ideal skeleton pose information for the target song and the skeleton pose information extracted by the skeleton pose analysis unit 31 every second (S6). The face recognition unit 51 performs face recognition on the actual video captured by the camera 4, and the pulse wave analysis unit 52 begins detecting the pulse wave (S7). The pulse wave analysis unit 52 calculates the maximum Lyapunov exponent λ for every 32 pulse wave beats (S8), and the emotion determination unit 53 classifies the subject's emotion into four levels based on the value and stores the result in the storage unit 70 (S9).
[0062] After that, when playback of the selected karaoke song ends (S10), the camera 4 stops capturing video, the recording unit 20 stops recording, and the skeleton pose analysis unit 31 stops saving still images (S11). The image extraction unit 41 extracts a specified number of images from the group of still images (image data group) stored in the training data storage unit 72, in which the coordinates or joint angles of the arms and legs differ by more than a reference value, and the model generation unit 42 generates a training model through deep learning (S12). The score calculation unit 33 calculates a dance score from the average value of the differences in each joint angle during song playback and displays it on the monitor 2C (S13). The emotion determination unit 53 calculates the occurrence rate of positive emotions during song playback and displays it on the monitor 2C as a mental score (S14).
[0063] If the dance score calculated in S13 is equal to or greater than the reference value (No in S15), the operation of the karaoke system 1 ends. If the dance score is less than the reference value (Yes in S15), the image generation unit 44 generates a virtual ideal image (S16). The image generation unit 44 further plays, for example, a virtual ideal image of the subject and an actual image for comparison, or a virtual ideal image of the subject and an actual image of the instructor's (teacher's) ideal dance movements for comparison, or plays a virtual ideal image with only the areas around the joints that are significantly different from the instructor's skeletal pose enlarged (S17). This ends the operation of the karaoke system 1.
[0064] Contrary to the judgment in S15, if the dance score is below the reference value, i.e., if the subject practices the dance and the difference between the skeletal pose and the instructor's skeletal pose becomes smaller, the virtual ideal video may be displayed as a reward. The model generation unit 42 may generate a learning model not only based on a video of the subject dancing to an actual karaoke song, but also based on a video of the subject performing a reference movement such as radio calisthenics.
[0065] The evaluation device 3 improves the subject's motivation to continue dancing by showing the subject their own video of their ideal dance, rather than a video of an instructor. The subject can compare their own video with the video of their ideal dance and feel the difference between their own video and the ideal dance movements, making it easy to see what needs to be corrected, improving the efficiency of dance learning. Once a learning model is generated, virtual ideal videos for multiple songs can be easily generated by inputting ideal skeletal pose information. Furthermore, the evaluation device 3 can be attached externally to an existing karaoke device 2, and requires minimal modification to the karaoke device, offering the advantages of low implementation costs and flexibility.
[0066] Although the above description has been given for the case of one subject, a virtual ideal video can also be generated using the same processing as above when two or more subjects are dancing side by side (when the subjects are a group). When there are multiple subjects, the skeletal pose analysis unit 31 extracts the skeletal pose of each subject from the actual video of the dance movements of the multiple subjects captured by the camera 4. In this actual video, the subjects must be captured without overlapping each other in order to recognize the skeletal poses. The ideal data storage unit 71 pre-stores group ideal skeletal pose information indicating a series of skeletal poses corresponding to ideal dance movements for the same song by the same number of instructors as the subjects.
[0067] FIG. 12 is a diagram illustrating the function of the generation unit 40 when there are multiple subjects. When there are multiple subjects, the model generation unit 42 generates a group learning model by learning the relationship between the actual video 80' of the subjects and the skeletal pose 81' of each person based on some still images selected by the image extraction unit 41 from the group of still images constituting the actual video of the subjects. The image extraction unit 41 selects still images for each subject in the video that have different arm and leg coordinates or joint angles. The physique correction unit 43 corrects the torso and leg lengths of each instructor's skeletal pose in the group ideal skeletal pose information to match that of the corresponding instructor in the group (reference numerals 90' and 91' represent the actual video and skeletal pose of the instructor). The video generation unit 44 generates a group virtual ideal video 85', which is a video of the subjects themselves performing the same dance movements as the instructor, based on the generated group learning model and the corrected group ideal skeletal pose information.
[0068] The model generation unit 42 may generate a group learning model by combining learning models generated separately for each subject, and the image generation unit 44 may generate a group virtual ideal image 85' by combining virtual ideal images 85 generated separately for each subject. However, when combining learning models, the positions of the subjects in the image must be approximately the same as the positions of the instructors in the group ideal skeletal pose information. Alternatively, the memory unit 70 may store group ideal skeletal pose information and a group learning model instead of ideal skeletal pose information and a learning model, and the image generation unit 44 may use a portion of that data to generate virtual ideal images of individuals in the group.
[0069] [Second embodiment] Next, a video production device according to a second embodiment of the present disclosure will be described. FIG. 13 shows an overall configuration diagram of the video production device according to the second embodiment of the present disclosure. In the video production device according to the first embodiment described above, an example has been described in which an evaluation device is provided separately from the karaoke device and the evaluation device is configured as hardware. In the video production device according to the second embodiment, all of the functions of the evaluation device are implemented as software and incorporated into the main body 21A of the karaoke device 20A. Therefore, each component of the main body 21A according to the second embodiment performs the same functions as the components included in the evaluation device of the video production device according to the first embodiment described above.
[0070] The video production device according to the second embodiment is made up of a karaoke device 20A. The karaoke device 20A is made up of a main body 21A, a speaker 2B, a monitor 2C, a remote control 2D, and a camera 4. The main body 21A has, as functional blocks, a music selection unit 12, a video selection unit 13, a playback unit 14, a display control unit 61, a storage unit 700, and a control unit 300, which are connected by a bus 200.
[0071] The storage unit 700 is realized by a semiconductor memory or a hard disk, and the other functional blocks are realized by a computer program such as dedicated plug-in software executed on a microcomputer including a CPU, ROM, RAM, and the like.
[0072] The storage unit 700 has a music database 11, an ideal data storage unit 71, and a learning data storage unit 72. The music database 11 stores music and video of karaoke songs. The ideal data storage unit 71 stores, for each song, ideal skeletal pose information that indicates a series of skeletal poses corresponding to ideal dance movements that match the song. The learning data storage unit 72 stores image data groups generated by a skeletal pose analysis unit included in the analysis unit 30 and information on learning models generated by a model generation unit included in the generation unit 40.
[0073] The music selection unit 12 and the video selection unit 13 select music and video stored in the music database 11 in response to the user's operation of the remote control 2D. The playback unit 14 outputs the music selected by the music selection unit 12 to the speaker 2B, and outputs the video selected by the video selection unit 13 to the monitor 2C, and plays them. The camera 4 is provided in the karaoke device 2. The camera 4 captures video of the subject performing dance movements to the music.
[0074] The control unit 300 has a recording unit 20, an analysis unit 30, a generation unit 40, and an emotion evaluation unit 50. The recording unit 20 records the video by storing data of the actual video captured by the camera 4 in the storage unit 700. When the subject dances to the music played by the karaoke device 20A, the actual video is recorded by the camera 4 and the recording unit 20.
[0075] The analysis unit 30 analyzes the skeletal pose of the subject in the actual video captured by the camera 4, finds the difference between the ideal value and each joint angle in the skeletal pose, and calculates a dance score. The analysis unit 30 also has a skeletal pose analysis unit that extracts skeletal pose information indicating a series of skeletal poses corresponding to the dance movements of the subject from the group of still images that make up the actual video.
[0076] If the dance score is below the reference value, the generation unit 40 learns the relationship between the actual video of the subject and the skeletal pose to generate a learning model. Then, the generation unit 40 corrects pre-stored ideal skeletal pose information to match the subject's physique, and generates a virtual ideal video, which is a video of the subject performing the same dance movements as the instructor, based on the learning model and the corrected ideal skeletal pose information.
[0077] The emotion evaluation unit 50 performs face recognition on the actual video captured by the camera 4 and detects the pulse wave from the skin color of the subject's face. Furthermore, the emotion evaluation unit 50 determines the subject's emotion category based on the degree of fluctuation in the pulse wave interval at regular intervals, such as every 30 seconds, while the subject is dancing, and calculates a mental evaluation score.
[0078] In response to instructions from the recording unit 20, the analysis unit 30, the generation unit 40 and the emotion evaluation unit 50, the display control unit 61 displays on the monitor 2C the actual video of the subject captured by the recording unit 20, the dance score, the virtual ideal video and the mental score.
[0079] The monitor 2C that is normally provided in the karaoke device 20A can be a liquid crystal display device or the like. When displaying images or the like on a screen larger than the monitor 2C, the karaoke device 20A may be provided with a projector 5, and the images to be output to the monitor 2C may be output to the projector 5.
[0080] When extracting skeletal pose information from a captured dance video of a subject, it is preferable to use as many images as possible from the captured video. However, processing many images at high speed may place a heavy load on the video generation section of the generation unit 40 (see FIG. 2). Therefore, it is preferable to further include an image processing section that generates the virtual ideal video in the video generation section. The image processing section 6 can be connected to the main body 21A via a cable 7. A GPU (Graphics Processing Unit) can be used for the image processing section 6. By using the image processing section, image processing can be performed at high speed, many images can be processed at high speed, and the generated virtual ideal video can be displayed smoothly. For example, an RTX2070 manufactured by NVIDIA can be used as the GPU. Furthermore, the cable 7 can be connected to a Thunderbolt (registered trademark) It is preferable to use a high-speed general-purpose data transmission cable such as 3.
[0081] In the video generation device according to the second embodiment, all the functions of the evaluation device are implemented as software and incorporated into the main body of the karaoke device, which allows for a compact device. In addition, by using an external image processing unit, it is possible to select and use an image processing unit with the appropriate processing capacity according to the amount of image data to be handled.
[0082] [Third embodiment] Next, a video generation device according to a third embodiment will be described. Fig. 14 shows an overall configuration diagram of the video generation device according to the third embodiment. In the video generation device according to the first embodiment, an example in which an evaluation device is used in combination with a karaoke device has been described. In the video generation device according to the third embodiment, the evaluation device is used as a standalone video generation device, independent of the karaoke device.
[0083] The image generation device according to the third embodiment includes an evaluation device 3A, a camera 4, and a projector 5. The evaluation device 3A realizes the same functions as the evaluation device 3 constituting the image generation device according to the first embodiment. The evaluation device 3A can be realized by a personal computer (PC) or the like.
[0084] The evaluation device 3A has a music selection unit 121, a video selection unit 131, a playback unit 140, a display control unit 60, a storage unit 701, a control unit 301, and an image processing unit 6A, which are connected by a bus 201. The evaluation device 3A also has a speaker 21 and a monitor 22. The evaluation device 3A can record an actual video of the subject captured by the camera 4 and display a virtual ideal video generated by the evaluation device 3A on the monitor 22.
[0085] The storage unit 701 is realized by a semiconductor memory or a hard disk, and the other functional blocks are realized by a computer program executed on a microcomputer including a CPU, a ROM, a RAM, and the like.
[0086] The storage unit 701 has a music database 11, an ideal data storage unit 71, and a learning data storage unit 72. The music database 11 stores music and videos. The music selection unit 121 and the video selection unit 131 select music and videos stored in the music database 11 in response to a user's operation using a keyboard or mouse. The playback unit 140 outputs the music selected by the music selection unit 121 to the speaker 21 and outputs the video selected by the video selection unit 131 to the monitor 22, where they are played. The camera 4 is externally attached to the evaluation device 3A. The camera 4 captures video of the subject performing dance movements in time with the music.
[0087] The control unit 301 has a recording unit 20, an analysis unit 30, a generation unit 40, and an emotion evaluation unit 50. The recording unit 20 records the video by storing data of the actual video captured by the camera 4 in the storage unit 701. When the subject performs a dance movement to the music played by the evaluation device 3A, the actual video is recorded by the camera 4 and the recording unit 20.
[0088] The operations of the analysis unit 30, generation unit 40, emotion evaluation unit 50, and display control unit 60 are similar to the operations executed by each component in the image generation device according to the first embodiment, and therefore detailed description thereof will be omitted.
[0089] A liquid crystal display device or the like can be used for the monitor 22 that is normally provided in the evaluation device 3A. When displaying images or the like on a screen larger than the monitor 22, a projector 5 may be provided in the evaluation device 3A, and the image to be output to the monitor 22 may be output to the projector 5.
[0090] A GPU can be used for the image processing unit 6A. The image processing unit 6A can perform image processing at high speed, and the generated virtual ideal image can be displayed smoothly.
[0091] The image generating device according to the third embodiment does not use a karaoke device, so the configuration can be simplified and the device can be used in, for example, dance classes held in private rooms.
[0092] [Fourth embodiment] Next, an image generating device according to a fourth embodiment will be described. FIG. 15 shows an overall configuration diagram of the image generating device according to the fourth embodiment. In the image generating device according to the third embodiment described above, an example was shown in which the image generating device is installed in a classroom or the like and used by a subject. In this case, the subject needs to go to the classroom or the like where the image generating device is installed, but there is also a need for subjects to be able to use the image generating device regardless of location. For example, when elderly people undergo rehabilitation at home, there is a need to increase their motivation for rehabilitation by viewing videos of ideal rehabilitation movements. Therefore, the image generating device according to the fourth embodiment further includes a communication interface that receives a real image of the subject from an external terminal device and transmits a virtual ideal image of the subject to the external terminal device.
[0093] The image generation device according to the fourth embodiment has an evaluation device 3B and a terminal device 100, which are connected via the Internet 500. The terminal device 100 may be, for example, a smartphone, a tablet terminal, a mobile phone (feature phone), a personal digital assistant (PDA), a desktop PC, a notebook PC, or other terminal device.
[0094] The evaluation device 3B is configured by a server and includes a storage unit 702, a control unit 302, an image processing unit 6B, and a communication interface (I / F) 8, which are connected by a bus 202. The evaluation device 3B realizes the same functions as the evaluation device 3 described above.
[0095] The storage unit 702 is realized by a semiconductor memory or a hard disk, and the control unit 302 is realized by a computer program executed on a microcomputer including a CPU, a ROM, a RAM, and the like.
[0096] The storage unit 702 has a music database 11, an ideal data storage unit 71, and a learning data storage unit 72. The music database 11 stores music and video.
[0097] The control unit 302 has an analysis unit 30, a generation unit 40, and an emotion evaluation unit 50. The operations of the analysis unit 30, the generation unit 40, and the emotion evaluation unit 50 are similar to the operations executed by each component in the video generation device according to the first embodiment, and therefore detailed explanations thereof will be omitted.
[0098] A GPU can be used for the image processing unit 6B. The image processing unit 6B can perform image processing at high speed, and the generated virtual ideal image can be displayed smoothly.
[0099] The communication I / F 8 transmits and receives video data to and from the terminal device 100 via the Internet 500 in a wired or wireless manner.
[0100] The terminal device 100 has a music selection unit 122, a video selection unit 132, a control unit 303, a storage unit 703, a playback unit 141, a display control unit 62, a recording unit 20, and a communication I / F 800, which are connected by a bus 203. The terminal device 100 also has a speaker 23, a monitor 24, and a camera 4. The recording unit 20 records video by storing data of actual video captured by the camera 4 in the storage unit 703.
[0101] The terminal device 100 uses the music selection unit 122 and the video selection unit 132 to select music and video stored in the music database 11 of the evaluation device 3B.
[0102] The playback unit 141 outputs the music selected by the music selection unit 122 to the speaker 23, and outputs the video selected by the video selection unit 132 to the monitor 24, and plays them. The camera 4 is built into the terminal device 100. When the subject dances to the music played on the terminal device 100, the camera 4 and the recording unit 20 record the actual video.
[0103] The actual video image representing the dance movements of the subject stored in the storage unit 703 is transmitted to the server 3B via the communication I / F 800 and the Internet 500.
[0104] The real video data transmitted to the server 3B is received by the communication I / F 8. The control unit 302 generates a virtual ideal video from the received real video and transmits the generated virtual ideal video to the terminal device 100 via the Internet 500. The procedure for generating a virtual ideal video from the real video of the subject is the same as the procedure in the video generation device according to the first embodiment.
[0105] The display control unit 62 of the terminal device 100 displays the virtual ideal image received from the server 3B on the monitor 24.
[0106] Next, the operation procedure of the image generating device according to the fourth embodiment will be described. Fig. 16 shows a flowchart for explaining the operation procedure of the image generating device according to the fourth embodiment of the present disclosure. Here, the case where an elderly person or the like is undergoing rehabilitation or the like at home will be described as an example.
[0107] First, in step S21, a video of the subject is captured by the camera 4 of the terminal device 100. For example, a video of an elderly person or the like performing movements such as squats to music is captured.
[0108] Next, in step S22, the captured video is stored in the storage unit 703 of the terminal device 100.
[0109] Next, in step S23, the saved video is uploaded from the terminal device 100 to the server 3B via the Internet 500.
[0110] Next, in step S24, the server 3B generates a virtual ideal video based on the uploaded video. For example, the server 3B collects skeletal pose information of the subject from the video uploaded by the subject, and generates a virtual ideal video of the subject performing correct squat exercise using a video of ideal squat exercise etc. prepared in advance.
[0111] Next, in step S25, the generated virtual ideal image is downloaded from the server 3B to the terminal device 100 via the Internet 500.
[0112] Next, in step S26, the downloaded virtual ideal video is displayed on the monitor 24 of the terminal device 100. By viewing the virtual ideal video of the subject performing the correct squat exercise, the subject can easily understand how to move their body. Furthermore, differences between the video uploaded by the subject and the virtual ideal video may be highlighted. This allows the subject to easily recognize which parts of the movements they performed were incorrect. The video generation device according to the fourth embodiment can improve the motivation of elderly people and others to perform rehabilitation exercises. Note that, although squats have been used as an example of rehabilitation exercise, this is not limiting and the invention can be applied to other rehabilitation exercises as well.
[0113] In the above explanation, an example has been shown in which a virtual ideal video is generated from a video of a subject captured by the terminal device 100, but as with the video generation device of the first embodiment, a dance score and a mental score may also be calculated together.
[0114] According to the image generation device of the fourth embodiment, a subject can generate a virtual ideal image from a real image at any location.
[0115] In the above description, an example was given in which skeletal pose information indicating a series of skeletal poses corresponding to the dance movements of a human subject was extracted from a group of still images constituting an actual video of the human subject, but this is not a limited example. That is, as long as a skeleton can be extracted, skeletal pose information may be extracted from non-human objects, such as dolls or animations. That is, the term "subject" is not limited to humans, but includes anything from which skeletal pose information can be extracted. For example, when applied to animation, skeletal pose information of an animation character can be extracted, and based on the extracted skeletal pose information, a virtual ideal video can be generated in which the animation character performs dance movements that conform to the ideal skeletal pose information.
[0116] In the above explanation, a dance video is used as an example of a virtual ideal video, but the present invention is not limited to such an example, and the ideal video may be a video of a movement performed at a predetermined timing, such as a baseball pitching form. For example, when applied to a pitching form, a virtual ideal video can be generated in which a subject performs the same pitching form as the famous player by using ideal skeletal pose information generated from the pitching motion of a famous player.
[0117] Furthermore, although the example has been shown in which the dance movements are performed in time with music, the present invention is not limited to this example, and the dance movements may be performed in time with a predetermined rhythm for synchronization, such as a predetermined rhythm sound produced by a metronome.
[0118] In the above explanation, an example was given in which a video of a subject performing dance movements similar to those in the virtual ideal video was captured with a camera and used as the real video, but this is not limited to this example. As long as the subject's skeletal pose information can be extracted, a video of movements other than dance movements can also be used as the real video. With this configuration, for example, a virtual ideal video of a subject performing ideal dance movements can be generated from a video of a subject who is unable to perform dance movements but is performing other movements. In this case, in order to generate a smooth virtual video, it is preferable that the real video be similar to the movements in the virtual ideal video.
Claims
1. A playback unit that plays music or rhythm sounds; a storage unit that stores, for each predetermined rhythm, ideal skeletal pose information that indicates a series of skeletal poses corresponding to movements that are performed at ideal predetermined timings that are in sync with the predetermined rhythm; a recording unit that records actual video of a dance movement performed by a subject at a predetermined timing, the video being captured by a camera in time with a predetermined rhythm of the music piece or the rhythmic sound being played; a skeleton pose analysis unit that extracts, from a group of still images that constitute the actual video, skeleton pose information that indicates a series of skeleton poses corresponding to movements made by the subject at predetermined timings; a model generation unit that generates a learning model based on the group of still images and the skeleton pose information, from the relationship between the group of still images and the skeleton pose information, so that when a skeleton pose is input, an image of the subject corresponding to the input skeleton pose is output; an image generating unit that generates and outputs a virtual ideal image, which is an image of the subject performing a movement at the predetermined timing that matches the ideal skeleton pose information, by inputting a series of the ideal skeleton pose information into the learning model; a difference determination unit that determines the difference between the angles of the joints in the skeleton poses of the skeleton pose information and the ideal skeleton pose information, the image generation unit outputs the virtual ideal image to a display unit when the difference is equal to or greater than a reference value, and does not output the virtual ideal image to a display unit when the difference is less than the reference value; An image generating device characterized by:
2. The image generating device of claim 1, wherein when the magnitude of the difference is equal to or greater than a reference value, the image generating unit selects a portion of a predetermined rhythm corresponding to an action performed by the subject at a predetermined timing, and plays back the actual image and the virtual ideal image in that portion.
3. the difference determination unit calculates differences in angles of a plurality of joints in the skeleton poses between the skeleton pose information and the ideal skeleton pose information, 2. The image generating device according to claim 1, wherein the image generating unit selects some of the joints when the magnitude of the difference between the angles of the plurality of joints calculated respectively is equal to or greater than a reference value, and enlarges the surrounding area of the some of the joints to reproduce the actual image and the virtual ideal image.
4. a score calculation unit that calculates an action score, which is a score of an action performed by the subject at a predetermined timing, according to the magnitude of the difference; 4. The image generating device according to claim 1, wherein the score calculation section calculates the action score such that the smaller the difference is, the larger the action score becomes.
5. a pulse wave analysis unit that extracts a pulse wave signal of the subject from a time-varying change in a luminance change component synchronized with the blood flow of the subject, the time-varying change being extracted from time-series data indicating the skin color of the subject in the actual video, and calculates a maximum Lyapunov index indicating a degree of fluctuation of a pulse wave interval; an emotion determination unit that determines that the subject's emotion is a negative emotion, such as brain fatigue, anxiety, or depression, when the maximum Lyapunov index is equal to or less than a threshold value set for emotion determination, and determines that the subject's emotion is a positive emotion, such as no brain fatigue, anxiety, or depression, when the maximum Lyapunov index is greater than the threshold value; The image generating device according to claim 1, wherein the emotion determination unit calculates the emotional rating, which is an emotional rating of the subject, so that the higher the frequency of occurrence of the positive emotion within a predetermined period of time in the real image, the higher the emotional rating.
6. a physique correction unit that corrects the length of the trunk and legs of the skeleton pose in the ideal skeleton pose information in accordance with the length ratio of the trunk and legs of the subject in the actual video so that the length ratio of the trunk and legs of the skeleton pose matches the length of the trunk and legs of the subject; The image generating device according to any one of claims 1 to 5, wherein the image generating unit generates the virtual ideal image by inputting the ideal skeletal pose information corrected by a series of the physique correction units into the learning model.
7. The method further comprises an image extraction unit that extracts from the group of still images a portion of still images in which the difference between the angle or position of a joint in each of the skeletal poses corresponding to each still image and the angle or position of a joint in the skeletal pose corresponding to one image is equal to or greater than a reference value that is a threshold value for the angle or position of the joint, The image generating device according to any one of claims 1 to 6, wherein the model generating unit generates the learning model based on some of the still images extracted by the image extracting unit and the skeleton pose information.
8. A system comprising an image generation device and an image processing device, The image generating device is a playback unit that plays back music or rhythm sounds; a storage unit that stores, for each predetermined rhythm, ideal skeletal pose information that indicates a series of skeletal poses corresponding to movements that are performed at ideal predetermined timings that are in sync with the predetermined rhythm; a recording unit that records actual video of a dance movement performed by a subject at a predetermined timing, the video being captured by a camera synchronized with a predetermined rhythm of the music or rhythmic sound being played; a skeleton pose analysis unit that extracts, from a group of still images that constitute the actual video, skeleton pose information that indicates a series of skeleton poses corresponding to movements made by the subject at predetermined timings; a model generation unit that generates a learning model based on the group of still images and the skeleton pose information, from the relationship between the group of still images and the skeleton pose information, so that when a skeleton pose is input, an image of the subject corresponding to the input skeleton pose is output; a difference determination unit that determines the difference between the angles of the joints in the skeleton poses of the skeleton pose information and the ideal skeleton pose information, The image processing device includes: an image generating unit that generates and outputs a virtual ideal image, which is an image of the subject performing a movement at the predetermined timing that matches the ideal skeleton pose information, by inputting a series of the ideal skeleton pose information into the learning model; the image generation unit outputs the virtual ideal image to a display unit when the difference is equal to or greater than a reference value, and does not output the virtual ideal image to a display unit when the difference is less than the reference value; A system characterized by:
9. A playback unit that plays music or rhythm sounds; a storage unit that stores, for each predetermined rhythm, ideal skeletal pose information that indicates a series of skeletal poses corresponding to movements that are performed at ideal predetermined timings that are in sync with the predetermined rhythm; a communication interface that receives from a terminal device actual video of a dance movement of a subject performed at a predetermined timing, captured by a camera in time with a predetermined rhythm of the music or rhythmic sound being played; a skeleton pose analysis unit that extracts, from a group of still images that constitute the actual video, skeleton pose information that indicates a series of skeleton poses corresponding to movements made by the subject at predetermined timings; a model generation unit that generates a learning model based on the group of still images and the skeleton pose information, from the relationship between the group of still images and the skeleton pose information, so that when a skeleton pose is input, an image of the subject corresponding to the input skeleton pose is output; an image generating unit that generates and outputs a virtual ideal image, which is an image of the subject performing a movement at the predetermined timing that matches the ideal skeleton pose information, by inputting a series of the ideal skeleton pose information into the learning model; a difference determination unit that determines the difference between the angles of the joints in the skeleton poses of the skeleton pose information and the ideal skeleton pose information, the image generation unit transmits the virtual ideal image to the terminal device via the communication interface when the difference is equal to or greater than a reference value, and does not transmit the virtual ideal image when the difference is less than the reference value; An image generating device characterized by:
Citation Information
Patent Citations
Karaoke device having feature in arranged animation display
JP1999133987A
Karaoke device provided with choreography scoring function
JP1999212582A
Method for synthesizing portrait image separately photographed with recorded background video image and for outputting the synthesized image for display and karaoke machine adopting this method
JP2000209500A
Karaoke device and karaoke video software
JP2001042880A
Animation creating system
JP2002269580A