Motion evaluation system, motion evaluation method
The motion evaluation system accurately assesses movement similarity between subjects by generating motion information images without a time axis and using structural similarity indices, addressing the challenge of evaluating dance movements synchronized with music.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-01-28
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies are inadequate for accurately evaluating the similarity of movement between subjects based on sound information, particularly in dance contexts, as they do not account for displacement and synchronization with music.
A motion evaluation system and method that captures images of subjects, calculates representative motion information in a time series, generates motion information images without a time axis, and uses structural similarity indices to evaluate movement similarity, incorporating sound information for synchronization and correction based on subject size and posture.
Enables accurate evaluation of movement similarity between subjects, accounting for displacement and synchronization with music, even when starting positions or sizes differ, by using position, velocity, and acceleration information, and correcting for subject differences.
Smart Images

Figure 0007843583000002 
Figure 0007843583000003 
Figure 0007843583000004
Abstract
Description
Technical Field
[0001] The present invention relates to a motion evaluation system and a motion evaluation method for evaluating the motion of a subject, particularly the similarity of motion between two subjects.
Background Art
[0002] It is known that dance involving whole-body movement contributes to the improvement of cardiopulmonary function and the enhancement of muscle strength and balance sense. By incorporating dance into daily habits, it is expected to reduce the risk of falls and injuries, for example, among the elderly. Also, by learning new dance steps and moving the body in rhythm with music, it is expected to enhance the spatial recognition ability, memory, attention, etc. of children with developmental disabilities and the elderly.
[0003] In dance lessons, students gradually acquire movement by imitating the steps of the instructor. In this process, by evaluating the similarity (degree of coincidence) of the movement between the instructor and the student and reflecting it in the guidance, it is possible to improve the dance performance of the student and further improve the motor function and the like. Also, by incorporating dance into rehabilitation, evaluating the similarity of the movement between a caregiver such as a doctor or a physical therapist and the care recipient, and reflecting it in a care plan such as the way of applying load, it is considered possible to proceed with rehabilitation more efficiently.
[0004] As a technique for evaluating the similarity of subjects captured in an image, for example, a method is known that includes the steps of: acquiring a digital image capturing an environment including at least a first subject; segmenting a first portion of the digital image capturing the first subject into a plurality of superpixels; assigning a semantic label to each of the plurality of superpixels; extracting features of the superpixels; determining a similarity index between the features extracted from the superpixels and features extracted from a reference superpixel identified in a reference digital image, wherein the index includes the step of giving the reference superpixel a reference semantic label that matches the semantic label assigned to the superpixel; and determining whether the first subject is captured in the reference image based on a plurality of similarity indices associated with the plurality of superpixels. (Patent Document 1)
[0005] According to Patent Document 1, it is possible to determine whether a first subject is captured in a reference image based on multiple similarity indicators associated with multiple superpixels. [Prior art documents] [Patent Documents]
[0006] [Patent Document 1] Special Publication No. 2021-531539 [Disclosure of the Invention] [Problems that the invention aims to solve]
[0007] However, the technology disclosed in Patent Document 1 is intended to identify people in digital images and enables the identification of people in digital images by combining features such as clothing, accessories, hair, and face, but it does not suggest evaluating the similarity of movement between two subjects that are displaced in accordance with sound information.
[0008] The present invention was devised to solve the problems of the prior art, and its objective is to provide a motion evaluation system and motion evaluation method that can easily and accurately evaluate the similarity of movement between two objects or measurement targets that are displaced in accordance with sound information. [Means for solving the problem]
[0009] The present invention, made to solve the aforementioned problems, comprises an imaging unit for capturing images of a first subject and a second subject, and a control unit. The control unit calculates at least one representative motion information representing the movement of the first subject and the second subject in a time series based on the output of the imaging unit. Furthermore, it generates a first motion information image and a second motion information image by plotting each of the representative motion information as pixels in a coordinate space that does not include the time axis. Based on the first motion information image and the second motion information image, it derives the degree of similarity of the movement of the first subject and the second subject. This makes it possible to represent representative motion information representing the movement of each subject as an image and to easily derive the degree of similarity based on the differences between the images.
[0010] Furthermore, the present invention is a motion evaluation system comprising a motion detection unit for detecting the movement of a first measurement target and a second measurement target, and a control unit, wherein the control unit calculates at least one representative motion information in time series for the first measurement target and the second measurement target based on the output of the motion detection unit, and further generates a first motion information image and a second motion information image by plotting each of the representative motion information as pixels in a coordinate space that does not include the time axis, and derives the degree of similarity of the movement of the first measurement target and the second measurement target based on the first motion information image and the second motion information image. This makes it possible to represent representative motion information that represents the movement of each measurement target as an image and to easily derive the degree of similarity based on the differences between the images.
[0011] Furthermore, the present invention uses position information, velocity information, or acceleration information as the representative motion information. This makes it possible to evaluate the similarity of motion between the first subject and the second subject in more detail using not only position information but also velocity information and acceleration information.
[0012] Furthermore, the present invention renders the representative motion information in the first motion information image and the second motion information image as an object of a predetermined size exceeding the size of one pixel. This makes it possible to properly obtain similarity even when the number of representative position information is small.
[0013] Furthermore, in this invention, the control unit calculates a structural similarity (SSIM) index as the similarity based on the first motion information image and the second motion information image. This makes it possible to derive the similarity while taking into account the characteristics of the human visual system.
[0014] Furthermore, the present invention is a motion evaluation method that involves photographing a first subject and a second subject, calculating at least one representative motion information that represents the movement of the first subject and the second subject in a time series, generating a first motion information image and a second motion information image by plotting each of the said representative motion information as pixels in a coordinate space that does not include the time axis, and deriving the degree of similarity of the movement of the first subject and the second subject based on the first motion information image and the second motion information image. This makes it possible to represent the representative motion information that represents the movement of each subject as an image and to easily derive the degree of similarity based on the differences between the images.
[0015] Furthermore, the present invention is a motion evaluation method that detects the movement of a first measurement target and a second measurement target, calculates at least one representative motion information that represents the movement of the first measurement target and the second measurement target in a time series, generates a first motion information image and a second motion information image by plotting each of the said representative motion information as pixels in a coordinate space that does not include the time axis, and derives the similarity of the movement of the first measurement target and the second measurement target based on the first motion information image and the second motion information image. This makes it possible to represent the representative motion information that represents the movement of each measurement target as an image and to easily derive the similarity based on the differences between the images.
[0016] Furthermore, the present invention is a motion evaluation system comprising: a sound information output unit that outputs sound information; an imaging unit that photographs a first subject and a second subject that are displaced based on the sound information; and a control unit. The control unit calculates at least one representative position information representing the respective position information of the first subject and the second subject from the output of the imaging unit based on a predetermined timing extracted from the sound information, performs calibration to match the initial values of the representative position information of the first subject and the second subject, and derives the degree of similarity of movement between the first subject and the second subject based on the time-series change of motion information extracted from the representative position information of the first subject and the second subject. This makes it possible to evaluate movement with high accuracy even if the positions where the first subject and the second subject start dancing (movement, motion) are different, or even if their positions in the image are different.
[0017] Furthermore, in the present invention, the control unit pre-calculates a correction coefficient to match the sizes of the first and second subjects based on their height or width dimensions when the first and second subjects assume the same posture, and corrects the representative position information of the first or second subject using the correction coefficient. This makes it possible to evaluate movement with high accuracy even if the sizes of the first and second subjects are different.
[0018] Furthermore, in this invention, the control unit detects a beat or rhythm contained in the sound information and extracts the peak value of the representative position information that is temporally before or after the timing based on the detected beat or rhythm as motion information. This makes it possible to accurately acquire motion information even for the movement of a subject that is intentionally out of sync with the beat timing.
[0019] Furthermore, in this invention, the sound information is music, and the control unit determines a predetermined timing based on the change in sound pressure of the sound information. This makes it possible to unify the timing of acquiring movement information for a first subject and a second subject who are performing a dance in time with the music.
[0020] Furthermore, in the present invention, the control unit sets a predetermined timing when the sound pressure of the sound information exceeds a predetermined value, or when the change in the sound pressure of the sound information exceeds a predetermined value. This makes it possible to easily obtain the timing for acquiring motion information.
[0021] Furthermore, the present invention is a motion evaluation method that involves photographing a first subject and a second subject that displace based on sound information, calculating at least one representative position information that represents the respective position information of the first subject and the second subject based on predetermined timings extracted from the sound information, performing calibration to match the initial values of the representative position information of the first subject and the second subject, and deriving the degree of similarity of movement between the first subject and the second subject based on the time-series changes in motion information extracted from the representative position information of the first subject and the second subject. This makes it possible to evaluate movement with high accuracy even if the positions where the first subject and the second subject start dancing (movement, motion) are different, or even if their positions in the image are different.
[0022] In addition, the present invention calculates in advance a correction coefficient for matching the sizes of the first subject and the second subject based on the sizes of the first subject and the second subject in the height direction or the width direction when the first subject and the second subject are in the same posture, and corrects the representative position information of the first subject or the second subject using the correction coefficient. As a result, even if the sizes of the first subject and the second subject are different, it is possible to evaluate the movement with high accuracy.
Advantages of the Invention
[0023] As described above, according to the present invention, it is possible to accurately evaluate the similarity of movement between two moving subjects or between two measurement targets.
Brief Description of the Drawings
[0024] [Figure 1] Block diagram showing the configuration of the motion evaluation system S1 according to the first embodiment of the present invention [Figure 2] (A) is an explanatory diagram showing the usage mode of the motion evaluation system S1, and (B) to (D) are explanatory diagrams for explaining the preprocessing in the motion evaluation system S1 [Figure 3] Explanatory diagram of the pose recognition model 40 [Figure 4] (A) is an explanatory diagram showing an example of an image of the subject 1, and (B) is an explanatory diagram showing the key points 41 and the center of gravity CGa of the whole body 1AL in the image of the subject 1 [Figure 5] (A) is a graph representing the movement of the first subject 1a in the x direction, (B) is a graph representing the movement of the second subject 1b in the x direction, and (C) is a graph obtained by normalizing the movement of the first subject 1a in the x direction within a range of ±1 [Figure 6] (A) and (B) are explanatory diagrams showing the timing of acquiring the movement information of the subject 1 [Figure 7] (A) and (B) are explanatory diagrams for explaining the process of deriving the similarity in the second embodiment of the present invention [Figure 8]This diagram illustrates a method for visualizing the movement of the subject 1 in a third embodiment of the present invention. [Figure 9] Block diagram showing the configuration of the motion evaluation system S1 according to the fourth embodiment of the present invention. [Figure 10] Block diagram showing the configuration of the motion detection unit 3. [Figure 11] Block diagram showing the configuration of the motion evaluation system S1 according to the fifth embodiment of the present invention. [Figure 12] Block diagram showing the configuration of the motion evaluation system S1 according to the sixth embodiment of the present invention. [Modes for carrying out the invention]
[0025] (First Embodiment) Hereinafter, a first embodiment of the present invention will be described with reference to the drawings. Figure 1 is a block diagram showing the configuration of a motion evaluation system S1 according to the first embodiment of the present invention. The motion evaluation system S1 consists of a control unit 10, a display unit 15, an imaging unit 13, and a sound information output unit 16. A sound information acquisition unit 17 is provided as needed, as will be described later.
[0026] The control unit 10 consists of an arithmetic unit 10a, a storage unit 10b, and a communication unit 10c. The arithmetic unit 10a consists of a CPU (Central Processing Unit), etc. The storage unit 10b consists of ROM (Read Only Memory), RAM (Random Access Memory), etc., and the arithmetic unit 10a operates according to a control program stored in the storage unit 10b. The storage unit 10b includes non-volatile memory (EEPROM (Electrically Erasable Programmable Read-Only Memory), etc.). Music files used for generating sound information are stored in the non-volatile memory. The storage unit 10b may also include so-called storage (mass storage device) such as an SSD (Solid State Drive) or HDD (Hard Disk Drive), and the control program and music files may be stored in the storage device. The arithmetic unit 10a is connected to the other components by a bus 20, etc., and the arithmetic unit 10a controls the other components via the bus 20, etc.
[0027] The control unit 10 may be configured as, for example, a PC (Personal Computer) or a server. The display unit 15 may be separate from the control unit 10, or it may be integrated with the control unit 10, such as a tablet terminal or a notebook PC. The communication unit 10c includes a communication module (not shown) that conforms to wireless communication standards such as LTE, LTE-M, 4G, or 5G. Furthermore, the communication unit 10c may include a communication module (not shown) that conforms to a short-range wireless communication standard such as BLE (Bluetooth® Low Energy). The communication unit 10c establishes communication with the imaging unit 13, the sound information output unit 16, and the sound information acquisition unit 17, enabling the control unit 10 to send and receive information to and from these units. Of course, these may be connected by wires.
[0028] The imaging unit 13 includes an image sensor composed of a CMOS (Complementary Metal Oxide Semiconductor) or a CCD (Charge Coupled Device). The imaging unit 13 may also be a camera provided in an information terminal such as a smartphone. In this case, the information terminal may include a communication module (not shown) compliant with a predetermined wireless communication standard, and transmit image data to the control unit 10 via the network 50. The imaging unit 13 may also be configured integrally with the control unit 10; in this case, the imaging unit 13 transmits image data to the processing unit 10a via the bus 20.
[0029] The sound information output unit 16 includes an amplifier, speaker, etc. (not shown). The control unit 10 generates an analog audio signal based on the music file stored in the storage unit 10b and outputs it to the sound information output unit 16. The sound information output unit 16 plays sound information (music / song) based on the analog sound signal. The sound information may be played back in the room where the subject 1 is dancing, etc., via a speaker, or it may be played back on a device such as wireless headphones worn by the subject 1. Here, sound information refers to either the digital data that constitutes the music file, the analog data obtained by decoding the digital data, or the sound output from the sound information output unit 16 based on the analog data.
[0030] The subject 1 may be fitted with a sound information acquisition unit 17. The sound information acquisition unit 17 includes a microphone, an AD converter, and a communication module compliant with a short-range wireless communication standard (none of which are shown). In this embodiment, the sound information reproduced by the sound information output unit 16 is acquired and digitized by the sound information acquisition unit 17 and transmitted to the control unit 10 by short-range wireless communication. The control unit 10 stores the received sound information as a music file in the storage unit 10b and may use this for motion detection and similarity derivation, which will be described later. As a result, even if the subject 1 and the sound information output unit 16 are far apart, fitting the sound information acquisition unit 17 to the subject 1 eliminates the effect of the time delay until the sound information reaches the subject 1.
[0031] Here, subject 1 is, for example, a human being. Subject 1 performs a dance according to a predetermined choreography in accordance with the sound information (music) output from the sound information output unit 16. The content of the sound information can be selected arbitrarily, and it is particularly preferable to select music with a clear beat, rhythm, or tempo. The dance choreography can also be selected arbitrarily, and it is particularly preferable to select choreography that increases the movement (displacement) of subject 1 at the timing when the beat occurs (on-beat). Note that the sound information is not limited to music or songs; for example, a sound source that emits a periodic sound such as a metronome may also be used.
[0032] The imaging unit 13 photographs the subject 1 performing a dance. The control unit 10 temporarily stores the image data received from the imaging unit 13 in the storage unit 10b. If the imaging unit 13 has a function to store image data in a portable storage medium, the control unit 10 may read the image data stored in the portable storage medium and store it in the storage unit 10b. The control unit 10 extracts motion information of the subject 1 using the stored image data. Here, subject 1 includes a first subject 1a and a second subject 1b. Based on the motion information, the control unit 10 derives the degree of similarity of movement between the first subject 1a and the second subject 1b.
[0033] Figure 2(A) is an explanatory diagram showing how the motion evaluation system S1 is used, and Figures 2(B) to (D) are explanatory diagrams explaining the preprocessing in the motion evaluation system S1. In the following explanation, the side of the subject 1 visible in Figure 2(B) may be referred to as the front, the opposite direction as the back, the direction of the right arm as the right, the direction of the left arm as the left, the direction of the head 1HD as the top, and the opposite direction as the bottom. In Figure 2(A), the sound information output unit 16 includes a speaker (not shown), and the sound information output unit 16 and the subject 1 are spatially separated by a predetermined distance. Of course, the position of the sound information output unit 16 can be determined arbitrarily.
[0034] Furthermore, regarding the positional relationship between the subject 1 and the imaging unit 13, it is preferable that the angle θ formed by the line extended forward from the front of the subject 1 and the optical axis AxL of the imaging unit 13 be in the range of, for example, 10°≦θ≦45°[deg]. This reduces occlusion and other issues during shooting, and allows many of the pose landmarks described later to be acquired with high reliability. The distance L (or field of view) between the subject 1 and the imaging unit 13 can be arbitrarily determined, but it is preferable to determine it while considering that the entire body 1AL (see Figure 3) of the subject 1 is captured when the subject 1 has both arms 1A raised (Figure 2(B)) and when the arms 1A are spread to the left and right (Figure 2(C)), and further considering the range of movement of the subject 1 during the dance.
[0035] From here on, the explanation will continue with reference to Figure 1. In the preprocessing of the motion evaluation system S1, after adjusting the field of view of the imaging unit 13, the subject 1 is photographed in the postures shown in Figures 2(B) to (D) before the subject 1 begins to move in accordance with the dance. Based on the image data received from the imaging unit 13, the control unit 10 measures the distance between the hands and feet (first hand-foot distance H1a) in the posture shown in Figure 2(B), where the subject 1 (here, the first subject 1a) has both arms 1A raised (the so-called "hands-up" posture). It also measures the distance between the hands (first hand-hand distance Wa) in the state shown in Figure 2(C), where the first subject 1a has its arms 1A spread to the left and right. Furthermore, it measures the distance between the top of the head 1HD and the feet (first head-foot distance H2a) in the posture shown in Figure 2(D), where the first subject 1a is standing upright (the so-called "attention" posture). Pose landmarks, which will be described later, can be used to measure this distance information.
[0036] For the second subject 1b, the distance between the second limbs H1b (Figure 2(B)), the distance between the second arms Wb (Figure 2(C)), and the distance between the second head and feet H2b (Figure 2(D)) are measured in the same way as for the first subject 1a. This distance information may be obtained by so-called 3D distance measurement, or it may be measured manually with a measuring tape or the like, and the values may be input to the control unit 10 via an input unit not shown. Based on the obtained distance information, the control unit 10 calculates the following. (i) Width direction (x direction) correction coefficient SFx = Distance between the first hands Wa / Distance between the second hands Wb (ii) Height correction coefficient (y-direction) SFy = Distance between first limbs H1a / Distance between second limbs H1b (Alternatively, SFy2 = distance between the first cephalic feet H2a / distance between the second cephalic feet H2b) In evaluating the motion, the control unit 10 corrects the positional information of the first subject 1a or the second subject 1b using these correction coefficients (details will be described later).
[0037] Figure 3 is an explanatory diagram of the pose recognition model 40. In the first embodiment, the open-source library MediaPipe Pose is used as the pose recognition model 40. As shown in Figure 3, for an image (still image or video) of a person, MediaPipe Pose recognizes key points 41 (pose landmarks) from 0.nose to 32.right_foot_index (33 locations in total) and outputs the positional information (coordinate values (x,y coordinates)) of the recognized key points 41.
[0038] The control unit 10 uses the obtained position information to calculate the coordinates of the center of gravity that represent the movement of the subject 1 for each of the following parts: the whole body 1AL, the head 1HD, the torso 1BD, and the legs 1L. Specifically, the following center of gravity is calculated using the position information of each key point 41. • Average values of each coordinate of the center of gravity CGa:0.nose~32.right_foot_index of the whole body 1AL • Average values of the coordinates of the center of gravity CGh:0.nose~12.right_shoulder of the head 1HD • Average values of the coordinates of the center of gravity CGb: 11.left_shoulder~24.right_hip of the torso 1BD. • Average values of the coordinates of the center of gravity of leg 1L CGl:23.left_hip~32.right_foot_index
[0039] In calculating the center of gravity CGh of the head 1HD and the center of gravity CGb of the torso 1BD, 11.left_shouldeu and 12.right_shoulde are both referenced, and in calculating the center of gravity CGb of the torso 1BD and the center of gravity CGl of the legs 1L, 23.right_hip and 24.left_hip are both referenced. However, it is preferable to exclude the coordinate values of keypoints 41 that are not obtained due to occlusion or other reasons, or that have low reliability, from the calculation of each center of gravity. In the following description, the position information of the center of gravity may be referred to as "representative position information". Representative position information is information that represents the movement of each subject 1. Thus, in the first embodiment, the position information of the center of gravity described above is used as representative position information, but the representative position information may also be calculated based on keypoints 41 that more strongly reflect the movement of the subject 1 according to the choreography of the dance. Alternatively, instead of (or in conjunction with) the representative position information, the first distance between both hands Wa or the distance between both feet (the distance between 31. left foot index and 32. right foot index) described above may be used.
[0040] Figure 4(A) is an explanatory diagram showing an example of an image taken of subject 1, and Figure 4(B) is an explanatory diagram showing key point 41 and the center of gravity CGa of the whole body 1AL in the image taken of subject 1. The process from taking a picture of subject 1 to obtaining representative position information will be explained below. The first subject 1a (for example, an instructor) and the second subject 1b (for example, a student) perform a dance with the same predetermined choreography to the same music output from the sound information output unit 16. First, as shown in Figure 4(A), the subject 1 (first subject 1a or second subject 1b) performing the dance is photographed by the imaging unit 13.
[0041] Here, the first subject 1a and the second subject 1b may be photographed using different imaging units 13, at different locations and at different times, or they may be photographed simultaneously using the same imaging unit 13. The imaging unit 13 photographs subject 1 in a time series at a predetermined frame rate (for example, 60 fps (frames per seconds)) and transmits the image data to the control unit 10.
[0042] The control unit 10 generates a video file based on the received image data and stores it in the storage unit 10b. Subsequently, the control unit 10 accesses the storage unit 10b to retrieve the image file, obtains keypoints 41 (and their coordinate values) from each frame image (evaluation image) that makes up the image file, and calculates representative position information (in this case, the coordinate value of the centroid CGa of the whole body 1AL). Specifically, the control unit 10 processes the evaluation image using the MediaPipe Pose API (Application Programming Interface) described above. As a result, as shown in Figure 4(B), multiple keypoints 41 are recognized for the subject 1 (in this case, the first subject 1a), and the x,y coordinates and representative position information of each keypoint 41 are calculated. The display unit 15 then displays the subject 1, keypoints 41, and representative position information (centroid CGa) superimposed, and further shows the skeleton connecting the main keypoints 41 and the outer edge encompassing the group of keypoints 41 as line segments. The same processing is performed for evaluation images taken of the second subject 1b.
[0043] Figure 5(A) is a graph showing the movement of the first subject 1a in the x-direction, Figure 5(B) is a graph showing the movement of the second subject 1b in the x-direction, and Figure 5(C) is a graph normalizing the movement of the first subject 1a in the x-direction to a range of ±1. Here, the vertical axis (x-direction) in Figures 5(A) to (C) corresponds to the left-right direction of subject 1 shown in Figures 4(A) and (B), and the horizontal axis is the time axis t. As mentioned above, since each evaluation image is obtained discretely (periodically) in a time series, the representative position information generated based on the evaluation images is also obtained discretely. However, in the graphs of Figures 5(A) to (C), the intervals between each representative position information are interpolated and drawn as curves (Figure 6, which will be described later, is similar). The representative position information uses the x-coordinate value of the centroid CGa of the whole body 1AL as described above.
[0044] The process for acquiring motion information of subject 1 is described below. First, when photographing subject 1, the control unit 10 synchronizes the timing of outputting sound information from the sound information output unit 16 with the timing of starting shooting with the imaging unit 13. The music file used for playing the sound information may be prepared in advance. Of course, the sound information output from the sound information output unit 16 may also be acquired (recorded) via the sound information acquisition unit 17 and used as a music file. In this case, the control unit 10 controls the timing so that the timing of starting shooting and the timing of starting recording are the same.
[0045] In this way, music and video files are obtained in which the start of music playback and the start of recording are synchronized. If the distance from the sound information output unit 16 to the subject 1 is large and affects the evaluation of motion (for example, if the time it takes for sound information to reach the subject 1 from the sound information output unit 16 exceeds half the period of the music beat), it is preferable to adjust the origin of the time axis t by adjusting the timestamp of the music file or image file. In Figures 5(A) to (C), the point in time when the subject 1 starts moving is set as the origin (0) of the time axis t, and the timestamp of the music file at this point is adjusted to 0.
[0046] The positional relationship between the imaging unit 13 and the first subject 1a (second subject 1b) is usually different each time an image is taken. Therefore, the control unit 10 performs a process to match the initial positions (in this case, the x-coordinate value of the representative position information) of the first subject 1a and the second subject 1b. Specifically, when the initial position of the representative position information of the first subject 1a is 540 (Figure 5(A)), an offset is added to the representative position information of the second subject 1b to match its initial value to 540 (Figure 5(B) shows the graph after the initial positions have been matched). In this way, the control unit 10 performs calibration to match the initial values of the representative position information of the first subject 1a and the second subject 1b. The y-coordinate value of the representative position information is also calibrated in the same way.
[0047] Hereinafter, the time-series change in the representative position information of the first subject 1a (i.e., Figure 5(A)) may be referred to as CGax(x,t), and the time-series change in the representative position information of the second subject 1b (i.e., Figure 5(B)) may be referred to as CGax'(x',t'). Here, the x-coordinate value of the centroid CGa of the whole body 1AL is used as an example of the representative position information.
[0048] The control unit 10 calculates the average value of the representative position information in the time series for the first subject 1a based on CGax(x,t). Then, as shown in Figure 5(C), it normalizes the data so that the average value is set to 0 and each representative position information falls within the range of ±1 (hereinafter, the normalized representative position information may be referred to as "normalized representative position information"). Hereafter, the normalized CGax(x,t) may be referred to as CGax_fin(x,τ). Furthermore, the control unit 10 normalizes CGax'(x',t') for the second subject 1b in the same way. Hereafter, the normalized CGax'(x',t') may be referred to as CGax_fin'(x',τ') (see Figure 6(B)).
[0049] Furthermore, when performing normalization, the normalized representative position information may be corrected using the correction coefficient (in this case, SFx) described above. Specifically, if SFx = first inter-arm distance Wa / second inter-arm distance Wb = 0.9, then for example, the normalized representative position information of the second subject 1b will be multiplied by 0.9. Of course, the normalized representative position information of the first subject 1a may also be multiplied by 1 / 0.9. This eliminates the influence based on differences in body shape, build, etc., of each subject 1. When evaluating the vertical movement (i.e., y-direction) of subject 1, the normalized representative position information can be corrected using SFy or SFy2.
[0050] Figures 6(A) and 6(B) are explanatory diagrams showing the timing for acquiring motion information for subject 1. Here, Figure 6(A) shows the timing extracted from sound information (τ1~τ14), normalized representative position information adopted as motion information (x1~x14), and the timing at which motion information was acquired (τ1a, τ2a, etc.) added to CGax_fin(x,τ) (see Figure 5(C)). Figure 6(B) shows the vertical axis of CGax'(x',t') (see Figure 5(B)) normalized to ±1, with the timing extracted from sound information (τ'1~τ'14), normalized representative position information adopted as motion information (x'1~x'14), and the timing at which motion information was acquired (τ'1a~τ'14a) added to it.
[0051] From here on, we will continue the explanation using Figure 1. The control unit 10 opens a music file in a digital audio format (WAV, MP3, etc.) stored in the memory unit 10b and performs decoding. Through decoding, the music data is converted into time-series sound pressure data, which is obtained by sampling the sound pressure at a fixed period. The control unit 10 detects the regular and irregular beats that make up the music from the sound pressure data. Here, the regular beats are related to the rhythm of the music, and from this perspective, it can be said that the control unit 10 also detects the tempo (BPM (Beats Per Minute)) based on the rhythm of the music. For BPM detection, methods such as FFT (Fast Fourier Transform) can be used.
[0052] The control unit 10 determines, for example, that a beat has occurred when the change in sound pressure data exceeds a predetermined threshold. Alternatively, it may determine that a beat has occurred when the sound pressure data exceeds a predetermined value. Furthermore, it may determine that a beat has occurred when the sound pressure data exceeds a predetermined value AND the change in sound pressure data over time exceeds a predetermined threshold. In other words, the control unit 10 extracts a predetermined timing when the sound pressure of the sound information exceeds a predetermined value, or when the change in the sound pressure of the sound information exceeds a predetermined value. This makes it possible to easily obtain the timing for acquiring motion information.
[0053] Regarding beat detection, it is also possible to distinguish between regularly occurring beats and irregular beats. For detecting regular beats, for example, an algorithm based on Sound Energy Variation can be used (https: / / mziccard.me / 2015 / 05 / 28 / beats-detection-algorithms-1 / ). This algorithm analyzes the energy for each measure of music and extracts regular beat patterns from these energy peaks. On the other hand, for detecting irregular beats, for example, an algorithm based on multipath search and cluster analysis can be used (Hindawi Complexity Volume 2021, "Music Rhythm Detection Algorithm Based on Multipath Search and Cluster Analysis"). This algorithm transforms sample data into the frequency domain using Short-Time Fourier Transform (STFT), extracts amplitude peak and phase information, and extracts PCM (Pulse Code Modulation) feature values from this information.
[0054] Thus, in the motion evaluation system S1 of the first embodiment, the sound information is music, and the control unit 10 determines a predetermined timing (the timing for acquiring motion information) based on the change in sound pressure of the sound information. This makes it possible to unify the timing for acquiring motion information for the first subject 1a and the second subject 1b, who are performing a dance in time with the music.
[0055] The control unit 10 acquires motion information based on the timing at which the beat is detected. In Figure 6(A), τ1 to τ14 and in Figure 6(B), τ'1 to τ'14 correspond to the timing at which the beat is detected. In this situation, when the first subject 1a and the second subject 1b dance to the same song, τ1 and τ'1, τ2 and τ'2...τ14 and τ'14 are at the same timing.
[0056] As described above, evaluation images are captured at predetermined intervals. The control unit 10 extracts multiple evaluation images captured in close proximity in time, centered around the timing when the beat is detected, and acquires motion information for each subject 1 based on the evaluation images. At this time, the normalized representative position information described above is referenced. The control unit 10 acquires normalized representative position information for each evaluation image captured within a predetermined period (for example, within ±1 / 3 of the beat period) centered around the timing when the beat is detected (for example, τ1 shown in Figure 6(A)), and adopts the normalized representative position information that satisfies predetermined criteria as the motion information for subject 1. Then, by combining this motion information with the time information (τ1a, etc.) at the time it was obtained, the first x-direction dataset: (x1,τ1a), (x2,τ2a)...(x14,τ14a) is obtained from CGax_fin(x,τ).
[0057] The following are some examples of criteria for extracting motion information of subject 1 from normalized representative position information. (C1) If multiple peaks of normalized representative position information are detected before and after the timing τ in which the beat was detected: The normalized representative position information with the largest absolute value is adopted as the motion information.
[0058] The following describes an example of applying the criterion (C1). In Figure 6(B), there are multiple peaks before and after τ'1, τ'6, and τ'8. By processing according to (C1), the shown P1, P2, and P3 are not adopted as motion information, and as a result, for the second subject 1b, the second x-direction dataset: (x'1,τ'1a), (x'2,τ'2a)...(x'14,τ'14a) is obtained from CGax_fin'(x',τ').
[0059] As described above, the motion evaluation system S1 of the first embodiment comprises a sound information output unit 16 that outputs sound information, an imaging unit 13 that photographs an object 1 that displaces based on the sound information, and a control unit 10. The control unit 10 acquires motion information of the object 1 from the output of the imaging unit 13 based on predetermined timings extracted from the sound information. In other words, the control unit 10 acquires motion information in substantially synchronization with the beat or rhythm contained in the sound information. This makes it possible to accurately acquire motion information of the object 1 that displaces based on the sound information.
[0060] Furthermore, the control unit 10 detects the beat or rhythm contained in the sound information and extracts the peak value of representative position information that is temporally ahead or behind the timing based on the detected beat or rhythm as motion information. It is known that skilled dancers enhance their expressiveness by intentionally delaying the timing of large body movements from the moment the beat is struck. Conversely, beginners may not be able to keep up with the rhythm of the music, and their body movements may lag behind the timing of the beat. With the present invention, it is possible to accurately acquire motion information even for movements that are intentionally (or due to lack of skill, etc.) out of sync with the beat (or are otherwise out of sync with the beat).
[0061] The application of the above-mentioned criterion (C1) is optional. For example, if multiple peaks of normalized representative position information exist within a predetermined period before and after the timing τ at which the beat was detected, all of these may be used as motion information. In other words, multiple motion information may be obtained for a single beat. Even if dances are performed to the same music, the number of peaks detected for the first subject 1a and the second subject 1b may differ, and this difference in the number of peaks may be reflected in the derivation of the similarity score.
[0062] The control unit 10 uses the x-direction first dataset, which is motion information of the first subject 1a, to determine the peak-to-peak average time Tpp_P_axn in the positive region and the peak-to-peak average time Tpp_N_axn in the negative region. Specifically, these are calculated as follows. In the following equations, kaxp represents the number of peaks in the positive region, and kaxn represents the number of peaks in the negative direction. ·Tpp_P_axn ={(τ3a-τ1a)+(τ6a-τ3a)+(τ8a-τ6a)+...+(τ14a-τ12a)} / kaxp =(τ12a-τ1a) / kaxp ·Tpp_N_axn ={(τ4a-τ2a)+(τ5a-τ4a)+(τ7a-τ5a)+...+(τ13a-τ11a)} / kaxn =(τ11a-τ2a) / kaxn
[0063] Similarly, using the second x-direction dataset, which contains motion information for the second subject 1b, we calculate the peak-to-peak average time in the positive region (Tpp_P_axn') and the peak-to-peak average time in the negative region (Tpp_N_axn'). These are calculated specifically as follows. In the following formulas, jaxp is the number of peaks in the positive region, and jaxn is the number of peaks in the negative region. ·Tpp_P_axn' =(τ'14a-τ'3a) / jaxp ·Tpp_N_axn' =(τ'13a-τ'1a) / jaxn
[0064] The average value between peaks in the positive / negative regions and the number of peaks for the first subject 1a and the second subject 1b are closely related to the beat. If there is a difference in these, it can be determined, for example, that the second subject 1b (student) made a mistake in the dance movements. The control unit 10 calculates F1 as an evaluation function, for example, as follows. Note that α and β in F1 are weighting coefficients and may be determined as appropriate. F1 =α{|(Tpp_P_axn)-(Tpp_P_axn')|+|(Tpp_N_axn)-(Tpp_N_axn')}+β(|kaxp-jaxp|+|kaxn-jaxn|)
[0065] Furthermore, the control unit 10 may calculate the difference between the sum of the absolute values of all elements of CGax_fin(x,τ) for the first subject 1a and the sum of the absolute values of all elements of CGax_fin'(x',τ') for the second subject 1b (evaluation function F2). Note that δ is a weighting coefficient and may be determined as appropriate. ·F2 =δ(Σ|CGax_fin(x,τ)|-Σ|CGax_fin'(x',τ')|)
[0066] F1 and F2 can be used as similarity indicators. These indicators approach zero as the difference in motion information between the first subject 1a and the second subject 1b decreases, i.e., as the similarity of motion increases. Of course, we may also use F1 and F2 to define the following evaluation function F3. F3 = F1 + F2 F3 can also be used as a similarity measure. Like F3, F3 approaches zero as the similarity between the movements of the two objects increases.
[0067] Thus, in the motion evaluation system S1 of the first embodiment, the subject 1 includes a first subject 1a and a second subject 1b, and the control unit 10 derives the degree of similarity of the movements of the first subject 1a and the second subject 1b based on motion information of the first subject 1a and the second subject 1b detected based on predetermined timings extracted from sound information. This makes it possible to accurately evaluate the degree of similarity of the movements of the first subject 1a and the second subject 1b that are displaced based on sound information.
[0068] Furthermore, in the motion evaluation system S1 of the first embodiment, the control unit 10 calculates at least one representative position information that represents the position information of the first subject 1a and the second subject 1b, and derives the similarity score based on the time-series changes in motion information extracted from the representative position information. This makes it possible to calculate the similarity score with high accuracy and speed without processing a large amount of motion information.
[0069] Furthermore, the motion evaluation system S1 of the first embodiment includes a sound information output unit 16 that outputs sound information, an imaging unit 13 that photographs a first subject 1a and a second subject 1b that are displaced based on the sound information, and a control unit 10. The control unit 10 calculates at least one representative position information representing the respective position information of the first subject 1a and the second subject 1b from the output of the imaging unit 13 based on predetermined timings extracted from the sound information, performs calibration to match the initial values of the representative position information of the first subject 1a and the second subject 1b, and derives the degree of similarity of movement between the first subject 1a and the second subject 1b based on the time-series changes in motion information extracted from the representative position information of the first subject 1a and the second subject 1b. By performing calibration, it becomes possible to evaluate motion with high accuracy even if the positions where the dance (action, movement) starts are different for the first subject 1a and the second subject 1b, or even if their positions in the image are different.
[0070] Furthermore, in the motion evaluation system S1 of the first embodiment, the control unit 10 pre-calculates a correction coefficient to match the sizes of the first subject 1a and the second subject 1b based on their height or width dimensions when the first subject 1a and the second subject 1b are in the same posture, and corrects the representative position information of the first subject 1a or the second subject 1b using the correction coefficient. This makes it possible to evaluate motion with high accuracy even if the sizes of the first subject 1a and the second subject 1b are different.
[0071] The above example demonstrates how to derive similarity between the first subject 1a and the second subject 1b based on the average value between peaks in the positive / negative regions and the number of peaks in the x-direction of the motion information of the whole body 1AL. Of course, similarly, similarity may also be derived based on the motion information of both subjects 1 in the y-direction (height direction). Furthermore, similarity may also be derived based on the average value between peaks in the positive / negative regions and the number of peaks in the x and y directions of the head 1HD, torso 1BD, and legs 1L, respectively, or similarity may be derived by integrating this motion information. Note that if the imaging unit 13 is configured as a stereo camera, the displacement of subject 1 in the front-to-back direction (see Figure 2(A)) can be measured. Similarity may also be derived by extracting motion information from this displacement amount in the front-to-back direction.
[0072] (Second Embodiment) Figures 7(A) and 7(B) are explanatory diagrams illustrating the process of deriving similarity in a second embodiment of the present invention. Here, Figure 7(A) shows the x,y distribution of representative position information of the first subject 1a in a time series, and is an image generated based on the movement of the first subject 1a in the x direction (CGax(x,t) shown in Figure 5(A)) and the movement in the y direction (not shown). Here, if the evaluation image is captured for 24 seconds at 60fps, for example, 60[fps] × 24[s] = 1440 representative position information points (x,y coordinate values) are acquired. These representative position information points are plotted as pixels in the x,y coordinates (i.e., a coordinate space that does not include the time axis). In Figure 7(A), the plotted area is shown as Tra. The range of the x,y coordinates is normalized to, for example, the range of 0 to 511. The representative position information is plotted as 8-bit monochrome image data, and for example, the pixel value is set to 255. Hereinafter, the image shown in Figure 7(A) will be referred to as the "first motion information image".
[0073] Furthermore, Figure 7(B) shows the x and y distribution of representative position information for the second subject 1b over time, and is generated in the same way as Figure 7(A) based on the x-direction movement of the second subject 1b (CGax'(x',t') shown in Figure 5(B)) and the y-direction movement (not shown). In Figure 7(B), the plotted region is shown as Trb. Hereafter, the image shown in Figure 7(B) will be referred to as the "second motion information image".
[0074] The configuration of the motion evaluation system S1 in the second embodiment is the same as in the first embodiment. The explanation will continue below with reference to Figure 1. The control unit 10 generates a first motion information image and a second motion information image, and substitutes the elements constituting each image into [Equation 1] below to obtain the structural similarity index (SSIM: Structural Similarity Index Measure).
number
[0075] SSIM provides an evaluation index (image quality evaluation index) that takes into account the characteristics of the human visual system based on three elements: brightness, contrast, and structure of an image. In [Equation 1], x and y are vectors representing each pixel within the window (here, 512 × 512) in the first motion information image and the second motion information image, respectively. μ is the average pixel value within the window, σ x ,σ y σ is the standard deviation of the pixel values within the same window. xy This is the covariance between x and y. C1 and C2 are constants that prevent the evaluation value from becoming unstable when the denominator becomes very small. Here, C1 = (K1L) 2 C2=(K2L) 2 L is the dynamic range of the pixel value (8 bits: 255 in this case). Also, K1 and K2 are constants, for example, K1 = 0.01 and K2 = 0.03.
[0076] As described above, in the second embodiment, the SSIM is calculated for the first subject 1a and the second subject 1b using representative position information in the x and y directions of the whole body 1AL (see Figure 3) (the centroid CGa of the whole body 1AL). When the first motion information image and the second motion information image match perfectly, SSIM(x,y)=1, and as the similarity decreases, the value of SSIM approaches 0. The control unit 10 displays the calculated SSIM as a similarity score on the display unit 15. Of course, the SSIM may also be calculated using representative position information of the head 1HD, torso 1BD, and legs 1L, and the values of these individual SSIMs may be combined as appropriate to serve as an index of similarity.
[0077] In calculating SSIM(x,y), the first motion information image and the second motion information image may each be divided into small regions, an SSIM may be calculated for each small region, and these may be averaged to obtain MSSIM (Mean SSIM). Furthermore, when deriving the similarity between images, instead of SSIM and MSSIM, or in conjunction with SSIM, for example, SNR (Signal to Noise Ratio) and PSNR (Peak Signal to Noise Ratio) may be used. Thus, in this second embodiment, both the first motion information image and the second motion information image are treated as image data, and the similarity is derived by comparing the two image data.
[0078] Furthermore, in the example described above, the first motion information image and the second motion information image were explained as two-dimensional images with pixels plotted in the x,y coordinate space. However, for example, the imaging unit 13 may be configured as a stereo camera to obtain depth information for the first subject 1a and the second subject 1b, and the similarity may be derived based on the three-dimensional information to which this depth information has been added. Here, depth information refers to motion information in the direction (z axis) that is orthogonal to both the x and y axes as shown in Figure 4. Of course, two-dimensional image data corresponding to the xy plane, yz plane, and zx plane may be obtained from the obtained three-dimensional information, and the similarity of motion between the first subject 1a and the second subject 1b may be derived based on each image data. Moreover, the first motion information image and the second motion information image may be one-dimensional images plotted in the x coordinate (or y coordinate, z coordinate).
[0079] As described above, the motion evaluation system S1 of the second embodiment comprises a sound information output unit 16 that outputs sound information, an imaging unit 13 that photographs the first subject 1a and the second subject 1b, and a control unit 10. The first subject 1a and the second subject 1b are displaced based on the sound information output by the sound information output unit 16. The control unit 10 calculates at least one representative position information representing the respective position information of the first subject 1a and the second subject 1b in a time series, and further generates a first motion information image and a second motion information image (image data) by plotting each representative position information as pixels in a coordinate space that does not include the time axis (in this case, a two-dimensional space), and derives a similarity based on the first motion information image and the second motion information image. This makes it possible to represent the movement (trajectory) of a position representing the subject 1 as a two-dimensional image and derive a similarity based on the difference between the images (by comparing the image data).
[0080] Furthermore, in the motion evaluation system S1 of the second embodiment, the control unit 10 calculates a structural similarity (SSIM) index based on the first motion information image and the second motion information image. This makes it possible to replace the movement of subject 1 with an image and derive a similarity score while taking into account the characteristics of the human visual system.
[0081] The following describes a modified version of the second embodiment. The control unit 10 generates a first motion information image with respect to the first subject 1a based on a first x-direction dataset: (x1,τ1a), (x2,τ2a)...(x14,τ14a) acquired based on CGax_fin(x,τ) shown in Figure 6(A), and a first y-direction dataset: (y1,τ1a), (y2,τ2a)...(y14,τ14a) acquired in the same manner as the first x-direction dataset.
[0082] Furthermore, with respect to the second subject 1b, a second motion information image is generated based on the second x-direction dataset: (x'1,τ'1a), (x'2,τ'2a)...(x'14,τ'14a) obtained based on CGax_fin'(x',τ') shown in Figure 6(B), and the second y-direction dataset: (y'1,τ'1a), (y'2,τ'2a)...(y'14,τ'14a) obtained in the same manner as the second x-direction dataset. That is, in the modified example, the first x-direction dataset, the first y-direction dataset, the second x-direction dataset, and the second y-direction dataset all use motion information of subject 1 obtained in synchronization with sound information.
[0083] In the modified example, the control unit 10 derives the similarity based on the first motion information image and the second motion information image. Thus, in the modified motion evaluation system S1, the control unit 10 generates a first motion information image and a second motion information image representing the time-series change of motion information for the first subject 1a and the second subject 1b, respectively, and derives the similarity based on the first motion information image and the second motion information image. In this case, SSIM may be used as in the second embodiment. This makes it possible to represent the movement of a position representative of subject 1 as a two-dimensional image and derive the similarity based on the difference between the images.
[0084] However, in the modified example, the number of points (pixels) constituting the motion information image is very small, only 14 in the example above (shown in Figure 6(A) as (x1,τ1a) to (x14,τ14a)). When an image is composed of a small number of pixels (dots), the structure of the first motion information image and the second motion information image may differ significantly, potentially leading to a very small SSIM being calculated and the similarity not being properly evaluated. Therefore, in the modified example, the pixels constituting the first and second motion information images are replaced with objects having an area larger than one pixel. Specifically, for example, one pixel is replaced with a circle with a predetermined radius r (e.g., r=5 pixels) centered on the x,y coordinates of that pixel. In this case, the interior of the circle may be filled with a predetermined value (e.g., 255), or a gradient may be provided in which the pixel value decreases radially from the center of the circle. By providing a gradient, sensitivity to edge structures can be reduced. Furthermore, in areas where multiple circles overlap, the average value of each gradient may be used for replacement, thereby suppressing the edges of objects and intentionally reducing features related to the image structure.
[0085] In this way, the modified motion evaluation system S1 renders representative position information (representative motion information described later) as an object of a predetermined size exceeding the size of one pixel in the first motion information image and the second motion information image. This makes it possible to properly acquire SSIM even when the number of representative position information items is small.
[0086] As mentioned above, the representative position information for the first subject 1a and the second subject 1b is acquired in a time series, for example, at a period of 60 fps. By differentiating the representative position information with respect to time, it is possible to obtain velocity information (representative velocity information), and further by differentiating the representative velocity information with respect to time, it is possible to calculate acceleration information (representative acceleration information). The representative position information, representative velocity information, and representative acceleration information (hereinafter, these may be collectively referred to as "representative motion information") are all representative information of the movement of each subject 1. That is, the representative motion information may be any of the position information, velocity information, or acceleration information. Instead of representative position information, the representative velocity information and representative acceleration information may be plotted as pixels to generate the first motion information image and the second motion information image, and the similarity may be derived by calculating SSIM, MSSIM, SNR, PSNR, etc. between these images. This makes it possible to evaluate the similarity of the movement between the first subject 1a and the second subject 1b in more detail using not only position information but also velocity information and acceleration information.
[0087] (Third embodiment) Figure 8 is an explanatory diagram illustrating a method for visualizing the movement of subject 1 in a third embodiment of the present invention. Figure 8 shows CGax(x,t) shown in Figure 5(A) and CGax'(x',t') shown in Figure 5(B) superimposed on a radar chart. The movement of the first subject 1a is shown by a solid line (hereinafter referred to as the "first graph"), and the movement of the second subject 1b is shown by a dashed line (hereinafter referred to as the "second graph"). The "movement" values here represent representative position information. The radial direction of the radar chart represents the movement (displacement) of subject 1 in the x direction, and the circumferential direction represents the passage of time (here, one revolution is 24 seconds). Subject 1 starts dancing at 0° and ends dancing at 360°. By representing the movements of the first subject 1a and the second subject 1b as a radar chart in this way, the similarity between the two can be easily evaluated visually.
[0088] In the radar chart, the representative position information of the first subject 1a and the second subject 1b is the same at 0° and 360°, and both the first and second graphs are drawn as closed curves. In the radar chart, the radial direction corresponds to the magnitude of the movement of subject 1, so the larger the movement, the more easily it is reflected in the increase of the area enclosed by the closed curve. Thus, in the third embodiment, it is possible to evaluate the dynamism of the movement of subject 1 using the area of the area enclosed by the first graph and the area of the area enclosed by the second graph. Of course, for example, the ratio of the area of the first graph to the area of the second graph may be used as the similarity.
[0089] (Fourth Embodiment) Figure 9 is a block diagram showing the configuration of the motion evaluation system S1 according to the fourth embodiment of the present invention. In the first embodiment, motion information of the subject 1 is extracted based on an evaluation image captured by the imaging unit 13 (see Figure 1), but in the fourth embodiment, motion information of the measurement target 2 is extracted using the motion detection unit 3. In the motion evaluation system S1 of the fourth embodiment, the imaging unit 13 shown in Figure 1 is replaced with the motion detection unit 3, and the subject 1 is replaced with the measurement target 2. That is, the measurement target 2 is, for example, a human being, and includes a first measurement target 2a (corresponding to the first subject 1a in the first embodiment) and a second measurement target 2b (corresponding to the second subject 1b in the first embodiment). Similar to the first embodiment, the first measurement target 2a and the second measurement target 2b are displaced in accordance with the sound information output from the sound information output unit 16.
[0090] A box-shaped motion detection unit 3 is attached to the left and right wrists and left and right ankles of the measurement target 2, respectively, using a wristband or the like. In order to detect the movement of the arm 1A and leg 1L (see Figure 2) with high accuracy, it is preferable that the motion detection unit 3 be attached to a part where the displacement is large. From this viewpoint, it is preferable that the motion detection unit 3 that detects the movement of the arm 1A is attached to the wrist or grasped in the palm. Also, it is preferable that the motion detection unit 3 corresponding to the leg 1L is attached to the ankle. The motion detection unit 3 may also be installed on the head 1HD or torso 1BD (see Figure 2) of the measurement target 2.
[0091] Here, it is preferable that the correspondence between each motion detection unit 3 and the part to which it is attached is predetermined. For example, each motion detection unit 3 is clearly marked with the part to which it should be attached, such as "for arm (right)," and the measurement target 2 attaches the motion detection unit 3 to the marked part.
[0092] Figure 10 is a block diagram showing the configuration of the motion detection unit 3. As shown in Figure 10, the motion detection unit 3 consists of a second control unit 3a, a second storage unit 3b, a second communication unit 3c, and an inertial sensor 3d. The second control unit 3a is composed of a CPU, etc., and operates according to a control program stored in the second storage unit 3b, which is composed of ROM, RAM, etc. The second control unit 3a and the other components are connected by a bus, etc., and the control unit 10 controls the other components via the bus, etc. The second storage unit 3b also stores identifiers (IDs) that represent each individual motion detection unit 3. The second communication unit 3c includes a communication module (not shown) that conforms to a short-range wireless communication standard, such as BLE. The second control unit 3a acquires the IDs stored in the second storage unit 3b and the output of the inertial sensor 3d, and transmits this information to the control unit 10 (see Figure 9) via the second communication unit 3c at a predetermined period (for example, a 10ms period).
[0093] The inertial sensor 3d is composed of, for example, a three-axis accelerometer and / or a gyroscope. Here, the three-axis accelerometer outputs the acceleration (in acceleration) of each part of the object being measured 2 in the direction and by how much. The gyroscope outputs the angular velocity (in angular velocity) of each part of the object being measured 2 in the direction and by how much speed. In this way, the inertial sensor 3d detects the movement of the arm 1A and leg 1L of the object being measured 2 and outputs three-axis acceleration information and / or three-axis angular velocity information (hereinafter sometimes referred to as "three-axis acceleration information, etc.") based on this to the control unit 10. At this time, the aforementioned ID is also output.
[0094] The control unit 10, having received the three-axis acceleration information and ID, calculates representative position information based on the three-axis acceleration information. By referring to the ID, the control unit 10 determines which motion detection unit 3 the output of the inertial sensor 3d originated from. The control unit 10 integrates the output (three-axis acceleration information) of each inertial sensor 3d to obtain velocity information, and further integrates this to obtain position information. Then, it averages the position information of each motion detection unit 3 to calculate the representative position information of the measurement target 2 in time series. Alternatively, by referring to the ID, representative position information may be obtained based on the output of a specific inertial sensor 3d. Furthermore, similar to the first embodiment, the control unit 10 acquires motion information of the measurement target 2 based on the beat and rhythm of the sound information. Based on this motion information, the control unit 10 derives the degree of similarity of motion between the first measurement target 2a and the second measurement target 2b.
[0095] As described above, the motion evaluation system S1 of the fourth embodiment comprises a sound information output unit 16 that outputs sound information, a motion detection unit 3 that detects the movement of the measurement target 2 that is displaced based on the sound information, and a control unit 10. The control unit 10 acquires motion information of the measurement target 2 based on predetermined timings extracted from the sound information. This makes it possible to accurately acquire motion information of the measurement target 2 that is displaced based on the sound information.
[0096] Furthermore, in the motion evaluation system S1 of the fourth embodiment, the measurement target 2 includes a first measurement target 2a and a second measurement target 2b, and the control unit 10 derives the similarity of the movements of the first measurement target 2a and the second measurement target 2b based on motion information of the first measurement target 2a and the second measurement target 2b detected based on predetermined timings extracted from sound information. This makes it possible to accurately evaluate the similarity of the movements of the first measurement target 2a and the second measurement target 2b that are displaced based on sound information.
[0097] Of course, the fourth embodiment and the second embodiment may be combined. That is, based on the representative position information (or representative velocity information, representative acceleration information) of the first measurement target 2a and the second measurement target 2b, a first motion information image and a second motion information image may be generated, respectively, and the similarity may be derived by applying image quality evaluation indicators such as SSIM, SNR, and PSNR to these image data.
[0098] (Fifth embodiment) Figure 11 is a block diagram showing the configuration of a motion evaluation system S1 according to a fifth embodiment of the present invention. The motion evaluation system S1 consists of a control unit 10, a display unit 15, and an imaging unit 13. The control unit 10, the display unit 15, and the imaging unit 13 have the same configuration as those described in the first embodiment, and their description is omitted here. However, in the fifth embodiment, the components of the motion evaluation system S1 do not need to include a sound information output unit 16 and a sound information acquisition unit 17 (see Figure 1). Therefore, music files and the like used to generate sound information do not need to be stored in the non-volatile memory of the storage unit 10b.
[0099] In the fifth embodiment, the subject 1 is, for example, a human. The subject 1 may also be an animal, a moving object, or an item whose form changes. The imaging unit 13 photographs the subject 1 performing a predetermined action (examples of predetermined actions will be described later). The control unit 10 extracts motion information (representative position information) of the subject 1 using the captured image data. In this case, motion information may be extracted by recognizing pose landmarks, as in the first embodiment. Alternatively, key points may be extracted from the subject 1 using, for example, SIFT (Scale-Invariant Feature Transform), and motion information may be extracted for multiple specific key points.
[0100] Here, subject 1 includes a first subject 1a and a second subject 1b. The control unit 10 calculates representative position information representing the movement of each subject 1 in a time series, for example, with a period of 60 fps, similar to the first embodiment. Here, the control unit 10 may determine a width direction (x direction) correction coefficient (SFx) and a height direction (y direction) correction coefficient (SFy) for the first subject 1a and the second subject 1b, and correct the representative position information for each subject 1, or it may perform calibration to match the initial values of the representative position information for the first subject 1a and the second subject 1b.
[0101] The control unit 10 derives the degree of similarity in movement between the first subject 1a and the second subject 1b based on the calculated representative position information. In deriving the degree of similarity, as in the second embodiment, a first motion information image and a second motion information image are generated by plotting the representative position information of the first subject 1a and the second subject 1b as pixels in a coordinate space that does not include the time axis. The control unit 10 then uses the first motion information image and the second motion information image to calculate evaluation values such as SSIM, MSSIM, SNR, PSNR, etc., i.e., derive the degree of similarity. Of course, the degree of similarity may also be derived based on the first motion information image and the second motion information image, which are plotted not only on the representative position information but also on the representative motion information described above (i.e., any of the representative position information, representative velocity information, or representative acceleration information).
[0102] As described above, the motion evaluation system S1 of the fifth embodiment comprises an imaging unit 13 that photographs a first subject 1a and a second subject 1b, and a control unit 10. Based on the output of the imaging unit 13, the control unit 10 calculates at least one representative motion information that represents the respective movements of the first subject 1a and the second subject 1b in a time series. Furthermore, it generates a first motion information image and a second motion information image by plotting each representative motion information as pixels in a coordinate space that does not include the time axis, and derives the degree of similarity of the movements of the first subject 1a and the second subject 1b based on the first motion information image and the second motion information image. This makes it possible to represent the representative motion information that represents the movement of each subject 1 as an image and to easily derive the degree of similarity based on the differences between the images.
[0103] Now, in the fifth embodiment, the following are examples of predetermined movements for which the similarity of movements between subjects 1 can be evaluated. • Choreography for singing Movements when performing a dance Movements when playing a musical instrument • Movements when operating vehicles, ships, aircraft, rockets, etc. • Movements when cooking • Movements during surgery (finger movements, hand movements, relative movements to specific organs, etc.) Movements performed by doctors, osteopaths, acupuncturists, physical therapists, etc. • Actions performed while working (cash register operation, smartphone operation, keyboard input, customer service posture, etc.) • Movements when using tools • Movements involved in sports (baseball and golf swings, table tennis, badminton, tennis, fencing, kendo, judo, wrestling, boxing, figure skating, skiing, skateboarding, rugby, soccer, swimming, gymnastics, bowling, etc.) • Movement and behavior of animals and insects, including pets. • Movement of moving objects in factories, etc., and movement of manufacturing equipment and production machinery (moving parts in factories) • Robot movements, including industrial robots
[0104] Furthermore, it is preferable to place a reference marker, such as a predetermined color marker or a marker configured in a predetermined shape, on the floor where each subject 1 stands, or near the target part of the body (e.g., hands or fingers) where the movement of each subject 1 is to be evaluated, and for each subject 1 to stand on the reference marker or touch the reference marker with a specific part of their body, and to start moving with these states as the initial position. This allows for effective calibration of the initial positions among the subjects 1.
[0105] In the fifth embodiment, the movement of subject 1 from which the similarity is derived does not have to be linked or synchronized with "sound" or "music" (of course, it may be linked or synchronized). Furthermore, the images of each subject 1 may include images of the same task or sports technique being performed repeatedly. There is also no particular restriction on the period during which each subject 1 is photographed. When evaluating the similarity of hand and finger movements such as surgery or cashiering, tracking software such as MediaPipe's Multi Hand Tracking mentioned above can be used. In addition, in combat sports such as wrestling and team sports such as soccer, it is preferable to trace the subject (player) whose movement is to be evaluated using known image recognition technology and perform pre-processing such as cropping out everything except the player. In this way, if the image captured by the imaging unit 13 contains multiple subjects 1, a specific subject 1 is extracted by pre-processing. Of course, each subject 1 may be included in the same image, in which case at least two subjects 1 are extracted from one image, and the similarity of movement between each subject 1 is derived.
[0106] (Sixth Embodiment) Figure 12 is a block diagram showing the configuration of a motion evaluation system S1 according to the sixth embodiment of the present invention. The motion evaluation system S1 consists of a control unit 10, a display unit 15, and a motion detection unit 3. The control unit 10, the display unit 15, and the motion detection unit 3 have the same configuration as those described in the fourth embodiment (Figures 9 and 10), and their description is omitted here. However, in the sixth embodiment, the components of the motion evaluation system S1 do not need to include a sound information output unit 16 and a sound information acquisition unit 17 (see Figure 9). Therefore, the non-volatile memory of the storage unit 10b does not need to store music files or the like used to generate sound information.
[0107] The following explanation will continue with reference to Figure 10. In the sixth embodiment, the control unit 10 integrates the output (i.e., acceleration information) of each inertial sensor 3d (here, a three-axis acceleration sensor) included in the motion detection unit 3 to obtain velocity information, and further integrates this to obtain position information. Then, the position information of each motion detection unit 3 is averaged to obtain representative position information that represents the movement of the measurement target 2. Of course, representative acceleration information may be obtained by averaging the acceleration information, and representative velocity information may be obtained by averaging the velocity information. Thus, in the sixth embodiment, representative motion information including representative acceleration information and representative velocity information is obtained in the process of obtaining representative position information. Of course, the number of inertial sensors 3d is arbitrary. If necessary, for example, an inertial sensor 3d may be placed on each of multiple fingers, and averaged representative motion information may be calculated based on their outputs. Alternatively, representative motion information may be obtained for each inertial sensor 3d (i.e., without averaging, for example, on a per-finger basis). This point is also the same for the fifth embodiment described above, where representative motion information may be derived based on the coordinate values of key points corresponding to the hand or fingers, for example.
[0108] The control unit 10 derives the degree of similarity between the movements of the first measurement target 2a and the second measurement target 2b based on the calculated representative position information (representative motion information). In deriving the degree of similarity, as in the second embodiment, a first motion information image and a second motion information image are generated by plotting the representative motion information of the first measurement target 2a and the second measurement target 2b as pixels in a coordinate space that does not include the time axis. The control unit 10 then uses the first motion information image and the second motion information image to calculate evaluation values such as SSIM, MSSIM, SNR, and PSNR. In the sixth embodiment as well, the movement of the measurement target 2 from which the degree of similarity is derived does not have to be linked or synchronized with "sound" or "music" (of course, it may be linked or synchronized with "sound" or "music").
[0109] As described above, the motion evaluation system S1 of the sixth embodiment comprises a motion detection unit 3 that detects the motion of a first measurement target 2a and a second measurement target 2b, and a control unit 10. Based on the output of the motion detection unit 3, the control unit 10 calculates at least one representative motion information that represents the motion of the first measurement target 2a and the second measurement target 2b in a time series, and further generates a first motion information image and a second motion information image by plotting each representative motion information as pixels in a coordinate space that does not include the time axis, and derives the similarity of the motion of the first measurement target 2a and the second measurement target 2b based on the first motion information image and the second motion information image. This makes it possible to represent the representative motion information that represents the motion of each measurement target 2 as an image and to easily derive the similarity based on the differences between the images.
[0110] The motion evaluation system S1 and motion evaluation method according to the present invention have been described in detail above based on specific embodiments. However, these embodiments are merely illustrative, and the present invention is not limited to these embodiments. For example, the first subject 1a and the second subject 1b (or the first measurement target 2a and the second measurement target 2b) may be the same person. By acquiring video files of the same subject 1 performing a dance at different points in time and deriving their similarity, it becomes possible to numerically represent the training results of the same person.
[0111] Furthermore, Subject 1 or Measurement Target 2 does not have to be a human. Specifically, for example, either the first subject 1a or the second subject 1b may be a robot. In this case, the robot is programmed to displace in accordance with sound information. Based on the similarity, the smoothness of the robot's movement, response speed, and displacement amount can then be evaluated. Of course, both the first subject 1a and the second subject 1b may be robots.
[0112] Furthermore, in the first to third embodiments, representative position information is calculated based on the coordinate values of the key points 41 of the pose recognition model 40 (see Figure 3). However, so-called motion capture technology may be used when calculating the representative position information. Specifically, multiple reflective markers attached to the subject 1 are photographed by the imaging unit 13, and representative position information is obtained based on the coordinates of the detected reflective markers.
[0113] Furthermore, in the first embodiment, the peak value of representative position information that is temporally before or after the timing based on the detected beat or rhythm is extracted as motion information. However, it is also possible to select the frame image (evaluation image) that is closest in time to the timing when the beat was detected, and to use the representative position information obtained from this evaluation image as motion information. Alternatively, the average value of multiple representative position information obtained within a predetermined period before and after the timing when the beat was detected may be used as motion information.
[0114] Furthermore, in each embodiment, even when performing the same choreography to the same song, differences in skill and proficiency will occur in the movements of the two subjects 1 (measurement target 2). Therefore, the first subject 1a was mainly described as a dance instructor and the second subject 1b as her student. On the other hand, the present invention may also be applied to rehabilitation conducted between doctors and the elderly, and may also be applied to the guidance and support of children with developmental disabilities, etc. [Industrial applicability]
[0115] The movement evaluation system S1 and movement evaluation method according to the present invention can improve dance performance by evaluating the similarity of movements between instructors and students and reflecting this in instruction. Furthermore, it is possible to further improve motor function and attention in children with developmental disabilities and the elderly. Therefore, it can be widely used in dance studios, support settings for children with developmental disabilities, nursing homes, and home care settings. Moreover, since the movement evaluation system S1 according to the present invention can easily derive the similarity of movements between subjects or measurement targets performing actions such as work, equipment operation, and sports, it can be widely used in improving efficiency in manufacturing sites based on motion analysis and work analysis, skills training, certification and transmission, technical guidance, sports instruction, pet toilet training, and abnormality detection in factory operating parts. [Explanation of Symbols]
[0116] 1. Subject 1a First subject 1b 2nd subject 2. Measurement target 3. Motion detection unit 10 Control Unit 13 Imaging Unit 16. Audio Information Output Section 40 Pose Recognition Models 41 Key Points 50 Networks S1 Motion Evaluation System
Claims
1. An imaging unit that photographs the first subject and the second subject, Control unit and Equipped with, The control unit, Based on the output of the imaging unit, at least one representative motion information item representing the movement of the first subject and the second subject is calculated in time series. Furthermore, a first motion information image and a second motion information image are generated by plotting each of the aforementioned representative motion information as pixels in a coordinate space that does not include the time axis. A motion evaluation system characterized by deriving an image quality evaluation index between images based on the pixel values of a plurality of pixels constituting the first motion information image and the second motion information image, as a degree of similarity in motion between the first subject and the second subject.
2. A motion detection unit that detects the movement of the first measurement target and the second measurement target, Control unit and Equipped with, The control unit, Based on the output of the motion detection unit, at least one representative motion information item representing the motion of the first measurement target and the second measurement target is calculated in time series. Furthermore, a first motion information image and a second motion information image are generated by plotting each of the aforementioned representative motion information as pixels in a coordinate space that does not include the time axis. A motion evaluation system characterized by deriving an image quality evaluation index between images based on the pixel values of a plurality of pixels constituting the first motion information image and the second motion information image, as a degree of similarity between the motion of the first measurement target and the second measurement target.
3. The motion evaluation system according to claim 1 or 2, characterized in that the representative motion information is any one of position information, velocity information, or acceleration information.
4. The motion evaluation system according to claim 1 or 2, characterized in that the representative motion information is drawn as an object of a predetermined size exceeding the size of one pixel in the first motion information image and the second motion information image.
5. The control unit, The motion evaluation system according to claim 1 or 2, characterized in that a structural similarity (SSIM) index is calculated as the similarity based on the first motion information image and the second motion information image.
6. Take a picture of the first subject and the second subject. For the first subject and the second subject, at least one representative motion information that represents their respective movements is calculated in time series. Furthermore, a first motion information image and a second motion information image are generated by plotting each of the aforementioned representative motion information as pixels in a coordinate space that does not include the time axis. A motion evaluation method characterized by deriving an image quality evaluation index between images based on the pixel values of a plurality of pixels constituting the first motion information image and the second motion information image, as a degree of similarity in motion between the first subject and the second subject.
7. The movement of the first measurement target and the second measurement target is detected. For the first measurement target and the second measurement target, at least one representative motion information that represents their respective movements is calculated in time series. Furthermore, a first motion information image and a second motion information image are generated by plotting each of the aforementioned representative motion information as pixels in a coordinate space that does not include the time axis. A motion evaluation method characterized by deriving an image quality evaluation index between images based on the pixel values of a plurality of pixels constituting the first motion information image and the second motion information image, as a degree of similarity in motion between the first measurement target and the second measurement target.
Citation Information
Patent Citations
Physical action evaluation device, karaoke system, and program
JP2014217627A
Motion form determination device, determination method, determination program, and determination system
JP2017055913A
Personal identification system and method
JP2021531539A
Information processing device, information processing method, and program
WO2019124069A1