Video editing method, video editing device, and display system
By estimating the subject skeleton model in the video and marking key operation frames, the problem of difficult shooting of user action videos at different time points is solved, and the unified playback time after video editing is achieved, and the accuracy of action training and evaluation is improved.
Patent Information
- Application Number
- JP2024501009
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-02-18
- Filing Date
- 2022-12-28
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-12-28
AI Technical Summary
The prior art is difficult to effectively compare user action videos taken at different time points, especially when shooting time and operation synchronization are poor.
By estimating the subject skeleton model in the video and identifying and marking key operation frames, we ensure that videos at different time points have a unified playback time after editing, making it easier to compare.
It is achieved easier to visually compare key operations in user action videos at different time points, improving the accuracy of action training and evaluation.
Smart Images

Figure 0007678540000001 
Figure 0007678540000002 
Figure 0007678540000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a video editing method, a video editing device, and a display system. [Background technology]
[0002] Traditionally, nursing homes and other facilities provide services that train users, such as the elderly, to live independently (so-called rehabilitation). In facilities with experts such as physical therapists, users' physical abilities are examined through interviews and medical examinations by the experts, and optimal functional training tailored to each user can be provided. On the other hand, in facilities without experts or when users train at their own homes, it is difficult to receive guidance from experts.
[0003] Therefore, for example, Patent Document 1 discloses a training evaluation device that displays a model video of a training performed by a user on the display of a user terminal together with a video of the user's training.
[0004] This allows the user to easily compare the model video of the training he or she desires with the user's actual training video on the display, making it easy to understand the appropriate posture for training, etc. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2020-195573 A Summary of the Invention [Problem to be solved by the invention]
[0006] Here, for example, a video of a user's past actions such as training or rehabilitation and a video of the user's current actions are simultaneously played back and displayed on a display. In such a case, if the shooting time of the subject to generate the two videos, that is, the video playback time of the videos, and / or the timing at which the user starts to move after shooting starts, are different, it is difficult to visually compare the actions in the two videos.
[0007] The present invention provides a video editing method, a video editing device, and a display system that make it easy to visually compare user actions included in each of a plurality of videos. [Means for solving the problem]
[0008] A video editing method according to one aspect of the present invention is a computer-executed video editing method, comprising: an acquisition step of acquiring a first video including a subject who performed a specific action in a first period of time, and a second video including the subject who performed the specific action in a second period of time different from the first period of time; an estimation step of estimating a skeletal model of the subject in each of the first video and the second video; and a step of identifying a first image in which the subject performs a key action of the specific action in the first video based on the skeletal model of the subject in the first video, and identifying a first image in which the subject performs the key action of the specific action in the second video based on the skeletal model of the subject in the second video. the step of: identifying a second image performing an operation; an editing step of editing at least one of the first video and the second video based on the first image and the second image to generate a third video based on the first video and a fourth video based on the second video so that their video playback times are aligned; and an output step of outputting the third video and the fourth video, wherein in the editing step, (i) the third video is edited so that the first image is included in the third video and the second image is included in the fourth video, and (ii) the number of images included in the at least one video is reduced or the number of images included in the at least one video is increased.
[0009] A video editing device according to an aspect of the present invention includes an acquisition unit that acquires a first video including a subject who has performed a specific action in a first period and a second video including the subject who has performed the specific action in a second period different from the first period, and an estimation unit that estimates a skeletal model of the subject in each of the first video and the second video, and an estimation unit that identifies a first image in which the subject performs a key action of the specific action in the first video based on the skeletal model of the subject in the first video, and estimates an image in which the subject performs the key action in the second video based on the skeletal model of the subject in the second video. an editing unit that generates a third video based on the first video and a fourth video based on the second video by editing at least one of the first video and the second video based on the first image and the second image so that their video playback times are aligned; and an output unit that outputs the third video and the fourth video, wherein the editing unit (i) edits at least one of the first video so that the third video includes the first image and the fourth video so that the second image is included in the fourth video, and (ii) edits at least one of the first video so that the first image is included in the third video and the second image is included in the fourth video, and
[0010] Moreover, a display system according to one aspect of the present invention includes the video editing device described above and an information terminal, wherein the video editing device further includes a first communication unit that communicates with the information terminal, and the information terminal includes a second communication unit that communicates with the video editing device, an instruction unit that instructs the subject to perform the specific action, a camera that generates the first moving image and the second moving image by photographing the subject performing the specific action, a control unit that outputs the first moving image and the second moving image to the video editing device via the second communication unit and acquires the third moving image and the fourth moving image from the video editing device via the second communication unit, and a display unit that displays the third moving image and the fourth moving image. Effect of the Invention
[0011] According to the present invention, a video editing method, a video editing device, and a display system are realized that can easily visually compare user actions included in each of a plurality of videos. [Brief description of the drawings]
[0012] [Figure 1] FIG. 1 is a block diagram showing a functional configuration of a display system according to an embodiment. [Diagram 2] FIG. 2 is a diagram for explaining a skeletal model of a subject estimated by an estimation unit according to the embodiment. [Diagram 3] FIG. 3 is a diagram for explaining an example of a time change in a user's movement according to the embodiment. [Figure 4] FIG. 4 is a diagram showing a specific example of the relationship between a specific action and a key action according to the embodiment. [Figure 5A] FIG. 5A is a diagram showing a first example of a relationship between positions of skeleton points corresponding to a specific action and images in a moving image including, as a subject, a user who has performed the specific action, in accordance with an embodiment. [Figure 5B] FIG. 5B is a diagram showing a second example of the relationship between the positions of skeleton points corresponding to a specific action and images in a moving image that includes, as a subject, a user who has performed the specific action in accordance with the embodiment. [Figure 5C] FIG. 5C is a diagram showing a third example of the relationship between the positions of skeleton points corresponding to a specific action and images in a moving image including, as a subject, a user who has performed the specific action in accordance with the embodiment. [Figure 5D] FIG. 5D is a diagram showing a fourth example of the relationship between the positions of skeleton points corresponding to a specific action and images in a moving image including, as a subject, a user who has performed the specific action in accordance with the embodiment. [Figure 5E] FIG. 5E is a diagram showing a fifth example of the relationship between the positions of skeleton points corresponding to a specific action and images in a moving image that includes, as a subject, a user who has performed the specific action in accordance with the embodiment. [Figure 5F]FIG. 5F is a diagram showing a sixth example of the relationship between the positions of skeleton points corresponding to a specific action and images in a moving image that includes, as a subject, a user who has performed the specific action in accordance with the embodiment. [Figure 6] FIG. 6 is a diagram for explaining an example of a video editing process executed by the editing unit according to the embodiment. [Figure 7] FIG. 7 is a diagram showing a specific example of two moving images displayed by the display unit according to the embodiment. [Figure 8] FIG. 8 is a flowchart illustrating a processing procedure of the video editing device according to the embodiment. [Figure 9] FIG. 9 is a sequence diagram showing a processing procedure of the display system according to the embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component arrangement and connection forms, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the present invention. In addition, among the components in the following embodiments, components that are not described in the independent claims will be described as optional components.
[0014] In addition, each drawing is a schematic diagram and is not necessarily a precise illustration. In addition, in each drawing, the same reference numerals are used for substantially the same configuration, and duplicated explanations may be omitted or simplified.
[0015] (Embodiment) [composition] First, the configuration of a display system according to the embodiment will be described.
[0016] FIG. 1 is a block diagram showing a functional configuration of a display system 10 according to an embodiment.
[0017] The display system 10 is a system that determines the level of activities of daily living (ADL) that a subject can perform based on moving images that include the subject performing a specific movement as a subject (i.e., moving images that show the subject).
[0018] The display system 10 includes an information terminal 30 and a video editing device 40.
[0019] A user photographs a subject at different time periods by, for example, operating the information terminal 30. A plurality of videos thus generated are transmitted to the video editing device 40. The video editing device 40 edits the received plurality of videos and transmits the edited plurality of videos to the information terminal 30. The information terminal 30 displays the received plurality of edited videos.
[0020] Here, the subject is a person whose degree of executable daily living activities is being determined, for example, a person whose physical function, i.e., the ability to move the body, has been reduced due to illness, injury, aging, or disability.
[0021] Furthermore, the user may be, for example, a physical therapist, an occupational therapist, a nurse, or a rehabilitation specialist.
[0022] Activities of daily living are the minimum daily activities necessary to lead a normal life, such as getting up and down, transferring, moving, eating, changing clothes (e.g. putting on and taking off shoes and putting on clothes), excretion, bathing (e.g. washing hair), and grooming.
[0023] The specific movement is a movement related to the activities of daily living. For example, the specific movement is a movement that is common to or similar to at least a part of the movements included in the activities of daily living. Specific examples of the specific movement will be described later.
[0024] The information terminal 30 is a computer that instructs the subject to perform a specific action, acquires a moving image (moving image data) including the subject as a subject, which is generated by photographing the subject with the camera 20, and transmits the acquired moving image to the video editing device 40. Specifically, the information terminal 30 generates a moving image made up of a plurality of images by photographing the subject, and transmits the generated moving image to the video editing device 40.
[0025] The information terminal 30 is, for example, a portable computer device such as a smartphone or a tablet terminal used by a user. The information terminal 30 may be a stationary computer device such as a personal computer.
[0026] The information terminal 30 includes a camera 20, a communication unit 31, a control unit 32, a storage unit 33, a reception unit , a display unit 35, and an instruction unit .
[0027] Camera 20 is a camera that captures an image of a subject performing a specific action, thereby generating a video including the subject performing the specific action as a subject. In this embodiment, camera 20 captures an image of a subject performing a specific action, thereby generating a plurality of videos, such as a first video and a second video.
[0028] The multiple moving images are, for example, moving images including as a subject a subject who performed a specific action during mutually different time periods. For example, the first moving image is a moving image including as a subject a subject who performed a specific action during a first time period, and the second moving image is a moving image including as a subject a subject who performed a specific action during a second time period different from the first time period. The first time period and the second time period may be different time periods.
[0029] The camera 20 may be a camera using a complementary metal oxide semiconductor (CMOS) image sensor, or may be a camera using a charge coupled device (CCD) image sensor.
[0030] The camera 20 may be an external camera attached to the information terminal 30. In this case, the information terminal 30 does not need to include the camera 20, and only needs to include a communication interface for connecting to the camera 20 so as to be able to communicate with it.
[0031] The communication unit 31 is a communication interface that communicates with the video editing device 40. Specifically, the communication unit 31 enables the information terminal 30 to communicate with the video editing device 40 via a network 5 such as the Internet. The communication unit 31 is an example of a second communication unit. The communication unit 31 is realized by, for example, a wireless communication circuit for wirelessly communicating with the video editing device 40.
[0032] The communication standard for communication performed by the communication unit 31 is not particularly limited.
[0033] Furthermore, the communication unit 31 may be connected to the video editing device 40 so as to be capable of wireless communication, or may be connected to the video editing device 40 so as to be capable of wired communication. For example, when the communication unit 31 is connected to the video editing device 40 so as to be capable of wired communication, the communication unit 31 is realized by a connector connected to a communication line or the like.
[0034] The control unit 32 is a processing unit that performs various information processes in the information terminal 30. The control unit 32 outputs the moving image generated by the camera 20 to the video editing device 40 via the communication unit 31, for example.
[0035] Specifically, the control unit 32 outputs the first and second videos to the video editing device 40 via the communication unit 31, and also acquires the third and fourth videos from the video editing device 40 via the communication unit 31.
[0036] The third video is a video obtained by editing the first video by video editing device 40. The fourth video is a video obtained by editing the second video by video editing device 40.
[0037] It is sufficient that video editing device 40 edits at least one of the first video and the second video. In other words, the first video and the third video may be the same video. Also, the second video and the fourth video may be the same video.
[0038] For example, when outputting a moving image, the control unit 32 outputs the moving image by linking the time information at which each image was generated with a plurality of images constituting the moving image. Also, for example, the control unit 32 acquires a plurality of moving images edited by the moving image editing device 40 (more specifically, the editing unit 42d) from the moving image editing device 40 via the communication unit 31. For example, the control unit 32 causes the display unit 35 to simultaneously display the acquired plurality of moving images. Also, for example, the control unit 32 acquires an evaluation result (evaluation result information) by the moving image editing device 40 (more specifically, the evaluation unit 42e) from the moving image editing device 40 via the communication unit 31. For example, the control unit 32 causes the display unit 35 to display information indicating the acquired evaluation result together with the plurality of moving images. Also, for example, the control unit 32 performs various processes based on an operation input accepted by the acceptance unit 34. The control unit 32 is realized by, for example, a microcomputer. Alternatively, the control unit 32 may be realized by a processor. The functions of the control unit 32 are realized, for example, by a microcomputer or processor constituting the control unit 32 executing a dedicated application program stored in the storage unit 33.
[0039] The storage unit 33 is a storage device that stores dedicated application programs and the like to be executed by the control unit 32. The storage unit 33 is realized by, for example, a semiconductor memory or a hard disk drive (HDD).
[0040] The reception unit 34 is an input interface that receives operation input by a user of the information terminal 30 (e.g., a subject or a rehabilitation specialist). For example, the reception unit 34 receives user input operations such as an instruction to shoot a moving image or an instruction to transmit the moving image to the video editing device 40. The reception unit 34 is realized by, for example, a touch panel display. For example, when the reception unit 34 is realized by a touch panel display, the touch panel display functions as the display unit 35 and the reception unit 34.
[0041] The reception unit 34 is not limited to a touch panel display, and may be, for example, a keyboard, a pointing device such as a touch pen or a mouse, or a hardware button. The reception unit 34 may be a microphone when receiving voice input. The reception unit 34 may be a camera when receiving gesture input. In this case, the reception unit 34 may be realized by the camera 20, or may be realized by a camera different from the camera 20.
[0042] Display unit 35 is a display device that displays the video edited by video editing device 40 and the evaluation results by video editing device 40. Display unit 35 displays, for example, the third video and the fourth video. In the present embodiment, display unit 35 simultaneously displays the third video and the fourth video so that the videos are played at the same time.
[0043] The display unit 35 may be realized, for example, by a display panel such as a liquid crystal panel or an organic EL (Electro Luminescence) panel, or by an audio device such as a speaker or earphones, or by a display panel and an audio device.
[0044] The instruction unit 36 is an instruction device that instructs the subject to perform a specific action. The instruction unit 36 instructs the subject to perform a specific action, for example, by video and audio. That is, the instruction unit 36 may instruct the user by video or audio.
[0045] The instruction unit 36 issues instructions by video and / or audio, such as "Please raise your arms up (raise your arms from a lowered position)", "Touch your back and maintain that position", "Touch the back of your head and maintain that position", "Touch your toes and maintain that position", etc. The instruction unit 36 may be realized by a display panel such as a liquid crystal panel or an organic EL panel, or may be realized by an audio device such as a speaker or earphones, or may be realized by a display panel and an audio device.
[0046] The instruction unit 36 and the display unit 35 may be realized by the same device such as a display panel and / or an audio device.
[0047] Furthermore, the video information and / or audio information used by the instruction unit 36 to instruct the subject to perform a specific action may be stored in the storage unit 33 in advance.
[0048] The video editing device 40 is a computer that receives a plurality of video images transmitted from the information terminal 30 and edits at least one of the received video images.
[0049] The video editing device 40 includes a communication unit 41 , an information processing unit 42 , and a storage unit 43 .
[0050] The communication unit 41 is a communication interface that communicates with the information terminal 30. Specifically, the communication unit 41 allows the video editing device 40 to communicate with the information terminal 30 via a network 5 such as the Internet. The communication unit 41 is an example of a first communication unit. The communication unit 41 is realized by, for example, a wireless communication circuit for wirelessly communicating with the information terminal 30.
[0051] The communication standard for communication performed by communication unit 41 is not particularly limited.
[0052] Furthermore, the communication unit 41 may be connected to the information terminal 30 so as to be capable of wireless communication, or may be connected to the information terminal 30 so as to be capable of wired communication. For example, when the communication unit 41 is connected to the information terminal 30 so as to be capable of wired communication, the communication unit 41 is realized by a connector connected to a communication line or the like.
[0053] The information processing unit 42 is a processing unit that performs various information processing in the video editing device 40. The information processing unit 42 is realized, for example, by a microcomputer. Alternatively, the information processing unit 42 may be realized by a processor. The functions of the information processing unit 42 are realized, for example, by the microcomputer or processor constituting the information processing unit 42 executing a computer program stored in the storage unit 43.
[0054] The information processing unit 42 includes an acquisition unit 42a, an estimation unit 42b, a specification unit 42c, an editing unit 42d, an evaluation unit 42e, and an output unit 42f.
[0055] The acquiring unit 42a is a processing unit that acquires moving images from the information terminal 30 via the communication unit 41. Specifically, the acquiring unit 42a acquires moving images output (transmitted) from the information terminal 30 via the communication unit 41. In the present embodiment, a first moving image including, as a subject, a subject who performed a specific action in a first period, and a second moving image including, as a subject, a subject who performed a specific action in a second period different from the first period are acquired.
[0056] The estimation unit 42b is a processing unit that estimates (calculates) a skeletal model of a subject in a moving image that includes a subject performing a specific action as a subject. Specifically, the estimation unit 42b estimates a skeletal model of the subject in the moving image based on the moving image acquired by the acquisition unit 42a. More specifically, the estimation unit 42b estimates a skeletal model in each of a plurality of images that constitute the moving image based on the moving image.
[0057] Fig. 2 is a diagram for explaining a skeletal model of the subject 1 estimated by the estimation unit 42b according to the embodiment. Specifically, Fig. 2 is a diagram that diagrammatically illustrates an image including the subject 1 as a subject, in which the skeletal model of the subject 1 estimated by the estimation unit 42b is superimposed on the subject 1.
[0058] The skeletal model is a model generated by connecting a plurality of skeletal points, which are specific positions such as joints of the subject 1 in an image, with links (lines). Specifically, the skeletal model is coordinate data of a plurality of skeletal points. For example, the estimation unit 42b estimates the positions (more specifically, coordinates) of a plurality of predetermined skeletal points of the subject 1 in an image, including a neck skeletal point, an elbow skeletal point, and a wrist skeletal point, by performing image analysis or the like. Furthermore, the estimation unit 42b connects, for example, predetermined skeletal points such as an elbow skeletal point and a wrist skeletal point among the estimated plurality of skeletal points with lines. In this way, the estimation unit 42b estimates the skeletal model of the subject 1.
[0059] In addition, the skeletal model may be estimated by any method, as long as an existing posture and skeleton estimation algorithm is used, for example.
[0060] The estimation unit 42b may estimate a two-dimensional skeletal model of the subject, or may estimate a three-dimensional skeletal model of the subject. That is, the estimation unit 42b may estimate two-dimensional coordinates of skeletal points of the subject in an image, or may estimate three-dimensional coordinates of skeletal points of the subject. For example, the estimation unit 42b estimates a two-dimensional skeletal model of the subject (i.e., coordinates of each skeletal point in a two-dimensional orthogonal coordinate system) based on the moving image acquired by the acquisition unit 42a, and estimates a three-dimensional skeletal model of the subject (i.e., coordinates of each skeletal point in a three-dimensional orthogonal coordinate system) based on the estimated two-dimensional skeletal model using a trained model, which is a trained machine learning model.
[0061] The trained model is a classifier constructed in advance by machine learning using a two-dimensional skeletal model in which the three-dimensional coordinate data of each joint is known as training data and the three-dimensional coordinate data as teacher data. The trained model receives the two-dimensional skeletal model as input and outputs three-dimensional coordinate data corresponding to the two-dimensional skeletal model, that is, a three-dimensional skeletal model. The trained model is stored in advance in the storage unit 43, for example.
[0062] In this manner, the estimation unit 42b may estimate a three-dimensional skeletal model of the subject in the moving image acquired by the acquisition unit 42a.
[0063] In the present embodiment, the estimation unit 42b estimates a skeletal model of the subject in each of the first video sequence and the second video sequence. Specifically, the estimation unit 42b estimates a skeletal model of the subject in the first video sequence based on the first video sequence. Also, the estimation unit 42b estimates a skeletal model of the subject in the second video sequence based on the second video sequence.
[0064] The identification unit 42c is a processing unit that identifies an image (also called a key image) in which the subject performs a key movement in a specific movement in the video, based on a skeletal model of the subject in the video estimated by the estimation unit 42b.
[0065] A key motion is a motion or posture included in a specific motion, and is a motion or posture that is particularly important when evaluating a subject who performs the specific motion. For example, one or more key motions are arbitrarily determined in advance for each specific motion.
[0066] Fig. 3 is a diagram for explaining an example of a time change in a user's motion according to an embodiment. Specifically, Fig. 3 is a graph showing the position of the subject's wrist height against the elapsed time when the subject performs a "banzai" motion as a specific motion.
[0067] In the example shown in FIG. 3, the subject first performs a specific motion of "banzai" (five-handed gesture) from a state in which the hands are lowered, thereby raising the hands, and then performs a motion of lowering the hands. In this specific motion of "banzai" (five-handed gesture), the motion (posture) at the highest position (time B) of the hands is the key motion. The identification unit 42c identifies an image in which the key motion is being performed from among a plurality of images included in the video in accordance with the specific motion. Specifically, the identification unit 42c identifies a first image in which the subject is performing a key motion in the specific motion in the first video, based on a skeletal model of the subject in the first video. Furthermore, the identification unit 42c identifies a second image in which the subject is performing a key motion in the second video, based on a skeletal model of the subject in the second video. The first video and the second video are each an example of a key image.
[0068] For example, when the specific motion is "banzai", the identification unit 42c identifies an image in which the position of the wrist skeleton point in the skeleton model is the highest as an image in which the key motion is being performed. Alternatively, when the specific motion is "banzai", the identification unit 42c may identify an image in which the amount of change in the position of the wrist skeleton point in the skeleton model is the smallest after the motion is started, or an image including a change point where the position changes from rising to falling, as a key image. In this way, for example, the identification unit 42c identifies a key image based on at least one of the position of one of a plurality of joints in the skeleton model of the subject corresponding to the key motion and the amount of change in the position.
[0069] Information indicating the positions and / or amounts of change of skeleton points for identifying a key image is, for example, arbitrarily set in advance and stored in the storage unit 43.
[0070] For example, the identification unit 42c may set (calculate) multiple three-dimensional areas around the skeletal model based on the positions of multiple skeletal points in the skeletal model estimated by the estimation unit 42b, and further, the identification unit 42c may identify a key image based on a three-dimensional area (target three-dimensional area) in which the skeletal point of the subject's wrist is located during a specific movement (i.e., while performing a specific movement) among the multiple three-dimensional areas set by the identification unit 42c.
[0071] For example, the identification unit 42c identifies in which of the multiple three-dimensional regions the coordinates of the subject's wrist skeleton point are located (in other words, included) based on the three-dimensional coordinate data (i.e., three-dimensional skeletal model) of the subject in the image. Also, for example, the identification unit 42c identifies one or more three-dimensional regions in which the subject's wrist skeleton point is located while the subject is performing the specific movement based on the three-dimensional skeletal model (i.e., three-dimensional coordinate data) of the subject in a moving image including the subject performing the specific movement as a subject. The identification unit 42c may identify a key image based on the identified one or more three-dimensional regions.
[0072] Further, for example, the identification unit 42c may identify, based on a skeletal model of the subject, a start image in the moving image in which the subject is performing a start movement, which is a movement (posture) that starts a specific movement, and an end image in which the subject is performing an end movement, which is a movement (posture) that ends the specific movement.
[0073] In the example shown in Fig. 3, in the specific action of "banzai," an image (image at time A) including the action (start action) of starting to raise a lowered hand is the start image. Also, an image (image at time C) including the action (end action) of completely lowering a raised hand via a key action is the end image. The start action and end action are, for example, arbitrarily determined in advance for each specific action.
[0074] For example, the specification unit 42c specifies the start image and the end image based on at least one of the position of any one of a plurality of joints in a skeletal model of the subject corresponding to the start motion and the end motion, and the amount of change in the position.
[0075] Information indicating the positions and / or amounts of change of skeleton points for identifying the start image and the end image is, for example, arbitrarily set in advance and stored in the storage unit 43.
[0076] FIG. 4 is a diagram showing a specific example of the relationship between a specific action and a key action according to the embodiment.
[0077] For example, when the specific motion is "standing up", two motions, a sitting position (e.g., sitting on a chair) and a standing position (standing up from a chair), are set as the motion to be determined (a motion including at least one of a start motion, a key motion, and an end motion). For example, the sitting position and the standing position are determined based on the position of the waist and the position of the knee. Specifically, the motion in which the position of the waist skeleton point and the position of the knee skeleton point in the skeletal model are closest in the height direction is the sitting position, and is set as the start motion and the end motion. The specification unit 42c specifies an image in the moving image in which the motion of the subject is in the sitting position as the start image or the end image. For example, when two images in which the motion of the subject is in the sitting position are specified, the specification unit 42c specifies the image that is played first in the moving image as the start image, and specifies the image that is played later in the moving image as the end image. Also, for example, the motion in which the position of the waist skeleton point and the position of the knee skeleton point in the skeletal model are farthest in the height direction is the standing position, and is set as the key motion. The identification unit 42c identifies an image in which the subject is in a standing position among the moving images as a key image.
[0078] Also, for example, when the specific motion is "stomping", two motions, one with the lowest ankle position and the other with the highest ankle position, are set as the motion to be determined. Specifically, the motion with the lowest ankle skeletal point position in the skeletal model is set as the start motion and the end motion. The identification unit 42c identifies an image in the moving image including the motion with the lowest ankle skeletal point position in the skeletal model as the start image or the end image. For example, when two images including the motion are identified, the identification unit 42c identifies the image that is played first in the moving image as the start image, and identifies the image that is played later in the moving image as the end image. Also, for example, the motion with the highest ankle skeletal point position in the skeletal model is set as the key motion. The identification unit 42c identifies an image in the moving image including the motion with the highest ankle skeletal point position in the skeletal model as the key image.
[0079] Also, for example, when the specific motion is "back touch", the motion with the highest wrist position is set as the key motion. Specifically, the motion with the highest wrist position in the skeletal model is set as the key motion. The specifying unit 42c specifies an image in the moving image including the motion with the highest wrist position in the skeletal model as the key image. Alternatively, for example, when the specific motion is "back touch", the motion with the most bent elbow angle is set as the key motion. For example, this motion is determined based on the positional relationship between the elbow skeletal point position, the shoulder skeletal point position, and the wrist skeletal point position in the skeletal model. Specifically, the motion with the smallest angle between the line connecting the wrist skeletal point and the elbow skeletal point and the line connecting the elbow skeletal point and the shoulder skeletal point in the skeletal model is set as the key motion. The specifying unit 42c specifies an image in the moving image including the motion with the smallest angle between the line connecting the wrist skeletal point and the elbow skeletal point and the line connecting the elbow skeletal point and the shoulder skeletal point.
[0080] Also, for example, when the specific motion is "touching the toes", the motion with the lowest wrist position is set as the key motion. Specifically, the motion with the lowest wrist skeletal point position in the skeletal model is set as the key motion. The specification unit 42c specifies, from among the moving images, an image including the motion with the lowest wrist skeletal point position in the skeletal model as the key image.
[0081] Also, for example, when the specific motion is "banzai," the motion in which the wrist position is the highest is set as the key motion. Specifically, the motion in which the wrist skeleton point position in the skeleton model is the highest is set as the key motion. The specification unit 42c specifies, from among the moving images, an image including the motion in which the wrist skeleton point position in the skeleton model is the highest as the key image.
[0082] Also, for example, when the specific motion is "touching the back of the head", the motion with the highest wrist position is set as the key motion. Specifically, the motion with the highest wrist skeletal point position in the skeletal model is set as the key motion. The specification unit 42c specifies, from among the moving images, an image including the motion with the highest wrist skeletal point position in the skeletal model as the key image.
[0083] In addition, the positions and positional relationships of the skeleton points used to specify the above-mentioned judgment target motion may be determined arbitrarily and are not limited to the above. For example, if the specific motion is "toe touch", the motion in which the wrist position and the ankle position are closest may be set as the key motion.
[0084] 5A to 5F are diagrams showing examples of the relationship between the positions of skeleton points corresponding to a specific action according to the embodiment and images in a video including a user who performed the specific action as a subject. The horizontal axis of the graphs shown in Figs. 5A to 5F indicates image numbers, which are natural numbers uniquely assigned to each image in order starting from 1, from the image that is played first among the multiple images included in the video, and the vertical axis indicates positional relationships such as the positions of the skeleton points of the subject or the distances between multiple skeleton points included in the image with the corresponding image number. In other words, the graphs shown in Figs. 5A to 5F are graphs showing the time changes in feature amounts such as the positions of the skeleton points of the subject or the distances between multiple skeleton points against the video playback time.
[0085] The first example shown in FIG. 5A is a graph showing the change over time in the distance in the height direction between the hip skeleton point and the knee skeleton point when the specific movement is "standing up".
[0086] The identification unit 42c, for example, identifies image number A1, which is an image before the key image in the video and includes an action in which the distance changes significantly in the direction of increasing, as the start image, identifies image number B1, which is the largest distance, as the key image, and identifies image number C1, which is an image after the key image in the video and includes an action in which the distance becomes smaller and little change is seen, as the end image.
[0087] In addition, when there are multiple key images in a moving image, such as the two peaks shown in FIG. 5A, the identification unit 42c, for example, arbitrarily identifies one of the key images and identifies the start image and the end image corresponding to that one of the key images.
[0088] The second example shown in FIG. 5B is a graph showing the change over time in height of the ankle skeleton point when the specific motion is "stepping".
[0089] For example, the specification unit 42c specifies the image with image number B2, which includes the movement in which the height of the ankle skeleton point is the highest, as the key image.
[0090] A third example shown in FIG. 5C is a graph showing the change over time in height of the elbow skeleton point when the specific motion is "back touch."
[0091] The identification unit 42c, for example, identifies the image with image number A3, which is an image before the key image in the video and includes an action in which the height changes significantly, as the start image, identifies the image with image number B3, which has the highest height, as the key image, and identifies the image with image number C3, which is an image after the key image in the video and includes an action in which the height becomes lower and little change is seen, as the end image.
[0092] A fourth example shown in FIG. 5D is a graph showing the change over time in the distance between the ankle skeletal point and the wrist skeletal point when the specific motion is a "toe touch."
[0093] The identification unit 42c, for example, identifies image number A4, which is an image before the key image in the video and includes an action that changes significantly in the direction of reducing the distance, as the start image, identifies image number B41, which includes an action that reduces the distance and little change is seen, as the key image (first key image), identifies image number B42, which is an image after the first key image and includes an action that changes significantly in the direction of increasing the distance, as the key image (second key image), and identifies image number C4, which is an image after the key image in the video and includes an action that increases the distance and little change is seen, as the end image.
[0094] In this manner, more than one key image may be identified.
[0095] A fifth example shown in FIG. 5E is a graph showing the change over time in height of the wrist skeleton point when the specific motion is "banzai (five-handed)."
[0096] The identification unit 42c, for example, identifies the image with image number A5, which is an image before the key image in the video and includes the action with the shortest height, as the start image, and identifies the image with image number B5, which has the highest height, as the key image.
[0097] The sixth example shown in FIG. 5F is a graph showing the change over time in height of the wrist skeleton point when the specific motion is "touch behind the head."
[0098] For example, the specification unit 42c specifies the image with image number B6 including the action with the highest height as the key image, and specifies the image with image number C6 which follows the key image and has the lowest height as the end image.
[0099] The thresholds for the amount of change in the feature amount and the like used by the specification unit 42c to specify the start image, the key image, and the end image may be determined arbitrarily and are not particularly limited.
[0100] For example, when the specific motion is "standing up," the determination unit 42c determines whether or not the distance in the height direction between the hip skeletal point and the knee skeletal point in an image played before the key image in the moving image changes by a predetermined amount or more in a direction away from the distance in the predetermined image to the distance in the image played after the predetermined image. When the determination unit 42c determines that the distance changes by a predetermined amount or more, the determination unit 42c determines the predetermined image as the start image. When multiple start images are determined, the determination unit 42c may determine any one image as the start image, or may determine one image as the start image based on a method arbitrarily determined in advance. The process of determining the key image and the end image is also similar to the process for determining the start image described above.
[0101] Such threshold information is stored in advance in the storage unit 43, for example.
[0102] The editing unit 42d is a processing unit that edits the video. Specifically, the editing unit 42d edits one or more of the multiple videos so that the video playback times of the multiple videos are aligned. For example, the editing unit 42d edits at least one of the first video and the second video based on the first image and the second image to generate a third video based on the first video and a fourth video based on the second video so that their video playback times (total video playback times from start to end) are aligned. Furthermore, the editing unit 42d (i) edits at least one of the first video and the second video so that the first image is included in the third video and the second image is included in the fourth video. Furthermore, the editing unit 42d (ii) edits at least one of the first video and the second video so as to reduce images included in at least one of the first video and the second video, or to increase images included in at least one of the first video and the second video. For example, the editing unit 42d generates the third video and the fourth video so that their video playback times are aligned by editing at least one of the first video and the second video so that the number of images contained in each of the first video and the second video are the same.
[0103] When adding an image to a video, an image that is the same as either the image before or after the position where the image is to be added may be added, or an image that is estimated from the images before and after the position may be generated and added.
[0104] In addition, aligning the video playback times does not only mean that the video playback times match completely, but also that they are substantially the same, for example, the video playback times differ by a few seconds, or the number of images included in the video differs by a few.
[0105] FIG. 6 is a diagram for explaining an example of a video editing process executed by editing unit 42d according to the embodiment. Specifically, (a) of FIG. 6 is a diagram that typically shows a first video. (b) of FIG. 6 is a diagram that typically shows a third video obtained by editing the first video. (c) of FIG. 6 is a diagram that typically shows a second video. (d) of FIG. 6 is a diagram that typically shows a fourth video obtained by editing the second video. The widths of the first to fourth videos shown in FIG. 6 typically indicate the video playback times of the respective videos. In the example shown in FIG. 6, the first and second videos have different video playback times. Therefore, editing unit 42d edits the first and second videos, respectively, to generate a third video based on the first video and a fourth video based on the second video, which have the same video playback times.
[0106] For example, the specification unit 42c specifies a start image, a key image, and an end image for each of the first video and the second video.
[0107] The editing unit 42d deletes images before the start images of the first and second videos, for example, based on the start images of the first and second videos. Next, the editing unit 42d deletes images after the end images of the first and second videos, for example, based on the end images of the first and second videos. Next, the editing unit 42d deletes one or more images between the start and end images of the first and second videos, for example, based on the start, key, and end images of the first and second videos, for example, by a predetermined number, so that the start, key, and end images remain and the start, key, and end images are played at the same timing. This generates a third video and a fourth video with the same video playback time (i.e., the same video playback time).
[0108] In this manner, for example, the editing unit 42d edits at least one of the first video and the second video based on the start image and the end image. Also, for example, the editing unit 42d edits at least one of the first video and the second video by deleting at least one of the video after the first image in the first video and the video after the second image in the second video. Also, for example, the editing unit 42d edits at least one of the first video and the second video by deleting a predetermined number of images included in at least one of the first video and the second video.
[0109] The predetermined number may be determined arbitrarily in advance and is not particularly limited. Information indicating the predetermined number is stored in the storage unit 43 in advance, for example.
[0110] Also, the timing of the key image playback in the third video and the timing of the key image playback in the fourth video may be different, but it is preferable that they are synchronized, which makes it easier to visually compare the key operations when the third video and the fourth video are played back simultaneously, as described below.
[0111] The start image and the end image may or may not be included in each of the edited videos (for example, the third video and the fourth video).
[0112] The evaluation unit 42e is a processing unit that evaluates a specific motion performed by the subject included in the video based on the specific motion. For example, the evaluation unit 42e generates difference information indicating a difference between an evaluation point in the first video and an evaluation point in the second video, the evaluation point being at least one of (iii) a position of any one of a plurality of skeletal points in a skeletal model of the subject, (iv) a positional relationship between two or more of the plurality of skeletal points, and (v) a moving speed of any one of the plurality of skeletal points, which corresponds to a key motion.
[0113] For example, when the specific action is "banzai," the evaluation unit 42e calculates the difference between the height of the wrist skeletal point in a key image (first image) included in the first video and the height of the wrist skeletal point in a key image (second image) included in the second video, using the height of the wrist skeletal point as the evaluation point, and generates difference information indicating the calculation result.
[0114] The evaluation score corresponding to a specific action may be determined arbitrarily in advance and is not particularly limited.
[0115] Furthermore, the evaluation unit 42e may determine (calculate) the level of the daily living activities corresponding to the specific movement that the subject can perform based on the above-mentioned evaluation points. For example, the evaluation unit 42e may determine the level of the daily living activities that the subject can perform. This determination result may be included in the evaluation information. Furthermore, for example, the evaluation unit 42e may generate evaluation information indicating an evaluation of the subject for the specific movement based on the calculated difference.
[0116] For example, if the specific motion is "banzai," it is considered that the higher the height of the wrist skeletal point, which is the evaluation point, the more appropriately the specific motion is executed. Therefore, for example, the evaluation unit 42e subtracts the evaluation point of the first or second video, which is generated later, from the evaluation point of the video, and if the subtraction is positive, it determines that the specific motion of the subject has improved, and if the subtraction is negative, it determines that the specific motion of the subject has deteriorated.
[0117] Furthermore, for example, the evaluation unit 42e may generate individual evaluation information indicating an evaluation of the subject with respect to a specific action in a moving image, based on the evaluation score in the moving image.
[0118] The evaluation unit 42e performs an evaluation of the subject with respect to a specific action based on a database stored in the storage unit 43, for example.
[0119] The database is data in which, for example, a specific motion, an evaluation score for the specific motion such as the height of the wrist, an evaluation method according to the evaluation score, an evaluation method corresponding to the difference, and an activity of daily living (ADL) corresponding to the specific motion are stored in association with each other. In the database, for example, when the specific motion is "raise your arms up," the evaluation score is the height of the wrist skeleton point, and an evaluation method corresponding to the evaluation score is set so that the higher the evaluation score, the more appropriately the specific motion is executed, and an evaluation method corresponding to the difference is set so that the higher the evaluation score in the later generated video image, the better the performance, and the ADL is set to be "eating."
[0120] The output unit 42f is a processing unit that outputs a plurality of images edited by the editing unit 42d to the information terminal 30 via the communication unit 41. Specifically, the output unit 42f outputs the third moving image and the fourth moving image. The control unit 32 acquires the outputted plurality of moving images and causes the display unit 35 to display the plurality of moving images. Also, for example, the output unit 42f outputs difference information and an evaluation result by the evaluation unit 42e together with the moving images via the communication unit 41. The control unit 32 acquires the outputted difference information and evaluation result and causes the display unit 35 to display the difference information and evaluation result together with the moving images.
[0121] FIG. 7 is a diagram showing a specific example of two moving images displayed by the display unit 35 according to the embodiment.
[0122] The control unit 32 causes the display unit 35 to display, for example, two videos 50, 51 based on videos generated in different time periods. For example, the video 50 is a video generated by capturing a subject in the past, such as one week before the videos 50, 51 are displayed on the display unit 35, and the video 51 is a video generated by capturing a subject more recently, such as on the day the videos 50, 51 are displayed on the display unit 35, after the video 50. The control unit 32 controls the display unit 35 so that the videos 50 and 51 are played side by side at the same time.
[0123] Also, for example, the control unit 32 causes the display unit 35 to display evaluation information such as "It's better than last time."
[0124] Furthermore, for example, the control unit 32 causes the display unit 35 to display difference information such as "degree of improvement: +□□mm."
[0125] In addition, the control unit 32 may display on the display unit 35, together with the moving images 50 and 51, information indicating a specific action, such as "action: cheers", information indicating an evaluation point for generating differential information, such as "evaluation point: wrist height", and information indicating the evaluation points for each of the moving images 50 and 51 for generating differential information, such as "last time: 〇〇 mm" and "this time: △△ mm".
[0126] This allows a user such as a subject using the information terminal 30 or a caregiver who cares for the subject to easily grasp changes in the subject performing a specific action.
[0127] The output unit 42f may output information such as the moving image 50, the moving image 51, the difference information, and the evaluation information to the information terminal 30 via the communication unit 41 at the same timing or at different timings.
[0128] The output unit 42f may output information indicating an ADL associated with a specific action. The control unit 32 may cause the display unit 35 to display the information indicating the ADL together with the moving images 50 and 51.
[0129] Furthermore, the moving images 50, 51 displayed on the display unit 35 may or may not include a skeletal model of the subject. For example, the output unit 42f may or may not output information indicating a skeletal model corresponding to the subject included in the moving images output to the information terminal 30 via the communication unit 41 to the information terminal 30 via the communication unit 41.
[0130] The storage unit 43 is a storage device that stores information such as the moving image (moving image data) acquired by the acquisition unit 42a, the control program executed by the information processing unit 42, the trained model, the database, and the various thresholds. The storage unit 43 is realized by, for example, a semiconductor memory or a HDD.
[0131] [Processing Procedure] Next, the processing procedure of the display system 10 will be described.
[0132] 8 is a flowchart showing a processing procedure of video editing device 40 according to an embodiment. In the following description, a case will be described in which video editing device 40 edits at least one of a first video including, as a subject, a subject who performed a specific action in a first period of time, and a second video including, as a subject, a subject who performed a specific action in a second period of time different from the first period of time, to generate a third video based on the first video and a fourth video based on the second video.
[0133] First, the acquisition unit 42a acquires the first video and the second video (S100). For example, the acquisition unit 42a acquires the first video and the second video from the information terminal 30 via the communication unit 41. When the first video and the second video are stored in the storage unit 43, the acquisition unit 42a acquires the first video and the second video from the storage unit 43.
[0134] Next, the estimation unit 42b estimates a skeletal model of the subject in each of the first video and the second video (S101). Specifically, the estimation unit 42b estimates a skeletal model of the subject in the first video (also referred to as a first skeletal model) based on the first video, and estimates a skeletal model of the subject in the second video (also referred to as a second skeletal model) based on the second video.
[0135] Next, the identification unit 42c identifies an image (key image) of a person performing a key action based on the estimated skeletal model of the subject in each of the first and second video sequences (S102). Specifically, the identification unit 42c identifies a first image that is a key image in the first video sequence based on the first skeletal model, and identifies a second image that is a key image in the second video sequence based on the second skeletal model.
[0136] Next, the editing unit 42d edits at least one of the first and second videos based on the first and second images to generate a third video based on the first video and a fourth video based on the second video so that their video playback times are aligned (S103). The editing unit 42d also edits at least one of the first and second videos so that (i) the third video includes the first image and the fourth video includes the second image, and (ii) the number of images included in at least one of the first and second videos is reduced or the number of images included in at least one of the first and second videos is increased.
[0137] Next, the evaluation unit 42e performs evaluation of the subject's specific motion based on the first and second moving images (S104). For example, the evaluation unit 42e generates an evaluation result including the difference information and evaluation information described above based on the first and second moving images.
[0138] The evaluator 42e may evaluate the subject's performance of a particular movement based on the third and fourth moving images.
[0139] Next, the output unit 42f outputs the third and fourth moving images and the evaluation result (S105).
[0140] FIG. 9 is a sequence diagram showing a processing procedure of the display system 10 according to the embodiment.
[0141] First, the instruction unit 36 instructs the subject to perform a specific movement (S201). For example, when the reception unit 34 receives an instruction from the user to perform a process of determining the degree of the daily living activities that the subject can perform (i.e., an instruction to have the subject start a specific movement), the instruction unit 36 gives an instruction such as "Please raise your arms in the air."
[0142] When the receiving unit 34 receives an instruction, the control unit 32 may acquire a video captured by the camera 20 and identify the subject in the acquired video. For example, a known image analysis technique such as pattern matching is used to identify the subject in the video.
[0143] Next, the camera 20 captures an image of a subject performing a specific action as a subject, thereby generating a moving image including the subject performing the specific action as a subject (S202).
[0144] Next, the control unit 32 outputs the moving image generated by the camera 20 to the video editing device 40 via the communication unit 31 (S203). At this time, the control unit 32 may anonymize the image and transmit it to the video editing device 40. This protects the privacy data of the subject.
[0145] Information terminal 30 executes steps S201 to S203 during mutually different time periods, thereby transmitting to video editing device 40 the first and second videos that have been generated by capturing images during mutually different time periods.
[0146] The first and second videos may be stored once in the storage unit 33 and then transmitted to the video editing device 40 at the same time.
[0147] Through the above-described processing of information terminal 30, acquisition unit 42a acquires the first video sequence and the second video sequence (S100).
[0148] Next, the estimation unit 42b estimates a skeletal model of the subject in each of the first and second video sequences (S101).
[0149] Next, the identification unit 42c identifies a key image in each of the first video sequence and the second video sequence based on the estimated skeletal model of the subject (S102).
[0150] Next, the editing unit 42d edits at least one of the first and second videos based on the first and second images to generate a third video based on the first video and a fourth video based on the second video so that their video playback times are aligned (S103). The editing unit 42d also edits at least one of the first and second videos so that (i) the third video includes the first image and the fourth video includes the second image, and (ii) the number of images included in at least one of the first and second videos is reduced or the number of images included in at least one of the first and second videos is increased.
[0151] Next, the evaluation unit 42e performs an evaluation of the subject's performance of the specific movement based on the first video sequence and the second video sequence (S104).
[0152] Next, output unit 42f outputs the third and fourth moving images and the evaluation result (S105). For example, output unit 42f outputs the third and fourth moving images and the evaluation result to information terminal 30 via communication unit 41.
[0153] Next, control unit 32 acquires, via communication unit 31, the third and fourth moving images and the evaluation result that have been output by output unit 42f via communication unit 41 (S204).
[0154] Next, display unit 35 displays the third and fourth moving images and the evaluation result acquired by control unit 32 (S205). Specifically, control unit 32 causes display unit 35 to simultaneously play back the acquired third and fourth moving images.
[0155] The information terminal 30 may perform the processes of steps S201 to S203 as one loop process, executing them each time the subject performs each of the multiple specific movements, or may perform the processes of steps S201 and S202 for each of the multiple specific movements, and execute step S203 after the subject has completed all of the specific movements.
[0156] In addition, when the subject performs a plurality of specific actions, the video images and the evaluation results associated with each of the specific actions may be displayed, or only the actions with the poorest evaluation results may be displayed. These judgment results may be displayed in order of poorest results.
[0157] Furthermore, the information terminal 30 may select a specific motion to be performed by the subject according to the subject's physical function before instructing the subject to perform the specific motion. For example, before step S201, the information terminal 30 may instruct the subject to perform a motion to stand up from a sitting position. At this time, the information terminal 30 may determine whether the subject can perform the motion to stand up based on an image of the subject captured by the camera 20. Alternatively, the specific motion to be performed by the subject may be selected based on a user's instruction received by the receiving unit 34.
[0158] This allows specific movements to be selected according to the subject's physical functions, making it possible to efficiently and accurately assess the state of the subject's daily living movements.
[0159] [Effects, etc.] As described above, the video editing method according to the embodiment is a computer-executed video editing method, and includes an acquisition step (S100) of acquiring a first video including a subject who performed a specific action in a first period and a second video including a subject who performed the specific action in a second period different from the first period, an estimation step (S101) of estimating a skeletal model of the subject in each of the first video and the second video, and an estimation step (S102) of estimating a key of the specific action of the subject in the first video based on the skeletal model of the subject in the first video. The method includes a step (S102) of identifying a first image of the subject performing an action and identifying a second image of the subject performing a key action in the second video based on a skeletal model of the subject in the second video, an editing step (S103) of generating a third video based on the first video and a fourth video based on the second video by editing at least one of the first video and the second video based on the first image and the second image so that their video playback times are aligned, and an output step (S105) of outputting the third video and the fourth video. In the editing step, at least one of the video is edited so that (i) the third video includes the first image and the fourth video includes the second image, and (ii) the at least one of the video is edited so that the image included in the at least one of the video is reduced or the image included in the at least one of the video is increased.
[0160] According to this, the video playback times of the first video and the second video are aligned so that key actions considered to be particularly important among specific actions for evaluating the subject's actions are included in each video. Therefore, it is possible to easily visually compare the user's actions included in each of the multiple videos. In particular, it is possible to easily visually compare the user's key actions included in each of the multiple videos.
[0161] Furthermore, for example, in the identification step, a start image in which the subject is performing a start movement to start a specific movement and an end image in which the subject is performing an end movement to end the specific movement are identified in at least one of the images based on a skeletal model of the subject, and in the editing step, at least one of the images is edited based on the start image and the end image.
[0162] This makes it possible to appropriately delete, for example, video during which a particular action included in the video is not being performed.
[0163] Also, for example, in the editing step, at least one of a video following the first image in the first video and a video following the second image in the second video is edited by deleting the video.
[0164] This makes it possible to easily delete video images during times when other than key operations are being performed.
[0165] Also, for example, in the editing step, a predetermined number of images included in the at least one of the images are deleted to edit the at least one of the images.
[0166] This makes it possible to reduce the discomfort felt by the user due to a sudden change in the subject's posture when the edited video is played back, while partially deleting the video to shorten the video playback time of the video.
[0167] Also, for example, in the identification step, the first image and the second image are identified based on at least one of the position of one of a plurality of skeletal points in a skeletal model of the subject corresponding to the key action and the amount of change in the position.
[0168] This allows the first image and the second image to be appropriately identified using the skeletal model.
[0169] Also, for example, the video editing method according to the embodiment further includes an evaluation step (S104) of generating difference information indicating a difference between an evaluation point in the first video and an evaluation point in the second video for an evaluation point which corresponds to a key action and is at least one of (iii) the position of any one of a plurality of skeletal points in a skeletal model of a subject, (iv) the positional relationship between two or more of the plurality of skeletal points, and (v) the movement speed of any one of the plurality of skeletal points, and further includes outputting the difference information in the output step.
[0170] This makes it possible to easily calculate changes in evaluation scores using the skeleton model.
[0171] Also, for example, in the evaluation step, evaluation information indicating an evaluation of the subject with respect to the specific motion is generated based on the difference, and in the output step, the evaluation information is further output.
[0172] This allows the user to easily understand whether the subject is properly performing a particular action or not.
[0173] Furthermore, video editing device 40 according to the embodiment includes an acquisition unit 42a that acquires a first video including a subject who performed a specific action in a first period and a second video including a subject who performed the specific action in a second period different from the first period, an estimation unit 42b that estimates a skeletal model of the subject in each of the first video and the second video, and an estimation unit 42c that identifies a first image in which the subject performs a key action of the specific action in the first video, based on the skeletal model of the subject in the first video, and identifies a second image in which the subject performs the key action in the second video, based on the skeletal model of the subject in the second video. the third video and the fourth video are generated based on the second video by editing at least one of the first video and the second video based on the first image and the second image so that their video playback times are aligned; and an output unit 42f that outputs the third video and the fourth video, wherein the editing unit 42d edits at least one of the first video and the second video (i) so that the third video includes the first image and the fourth video includes the second image, and (ii) so that the number of images included in at least one of the first video and the fourth video are reduced or increased.
[0174] Furthermore, the display system 10 according to the embodiment includes a video editing device 40 and an information terminal 30. The video editing device 40 further includes a first communication unit (communication unit 41) that communicates with the information terminal 30. The information terminal 30 includes a second communication unit (communication unit 31) that communicates with the video editing device 40, an instruction unit 36 that instructs a subject to perform a specific action, a camera 20 that generates a first moving image and a second moving image by photographing the subject performing the specific action, a control unit 32 that outputs the first moving image and the second moving image to the video editing device 40 via the second communication unit and acquires a third moving image and a fourth moving image from the video editing device 40 via the second communication unit, and a display unit 35 that displays the third moving image and the fourth moving image.
[0175] According to these, the same effects as those of the video editing method according to the above-mentioned embodiment can be achieved.
[0176] (Other embodiments) Although the embodiment has been described above, the present invention is not limited to the above embodiment.
[0177] For example, an instruction to cause a target person to perform a specific action may be given by a user, in which case the information terminal does not need to have an instruction unit.
[0178] Furthermore, the information terminal 30 may transmit to the video editing device 40 a video of the subject performing the specific action together with information indicating the specific action.
[0179] Also, for example, the information terminal 30 may transmit to the video editing device 40 an instruction to edit a moving image based on an instruction from a user received by the receiving unit 34. In this case, for example, the video editing device 40 may transmit to the information terminal 30 information indicating an instruction to have the subject perform a specific action. Also, in this case, the information terminal 30 may cause the subject to perform a specific action by the instructing unit 36 based on the received information, and may also capture an image of the subject by the camera 20.
[0180] The evaluation unit 42e may also calculate a feature amount indicating the feature of the movement of the subject in a specific movement based on the skeletal model estimated by the estimation unit 42b, and determine the physical function, which is the ability of the subject to perform a physical movement, based on the calculated feature amount. For example, the evaluation unit 42e calculates the angle (joint angle) between two links connected to a specific skeletal point of the subject as a feature amount based on the skeletal model estimated by the estimation unit 42b. Alternatively, for example, the evaluation unit 42e calculates the distance between a specific skeletal point and an extremity part in a specific movement, and the fluctuation range of the position of the specific skeletal point in a specific movement, as a feature amount. For example, the evaluation unit 42e determines the physical function of the subject based on whether each calculated value is equal to or greater than a predetermined threshold value, or whether each calculated value is within a predetermined range.
[0181] According to this, for example, for a subject who has no problem with daily living activities, a training plan necessary for maintaining or improving a physical function based on a physical function such as muscle strength can be provided. The predetermined skeletal points, the predetermined threshold, and the predetermined range may be arbitrarily determined. These pieces of information may be stored in the storage unit 43 in advance.
[0182] The evaluation unit 42e may further determine the degree of daily living activities that the subject can perform based on the determination result of whether or not the subject can perform an action involving finger movement (for example, opening and closing the hand (opening the hand) or opposing the fingers (OK sign)). For example, when the reception unit 34 of the information terminal 30 receives an instruction to determine whether or not the subject can perform an action involving finger movement, the control unit 32 instructs the instruction unit 36 to perform an action involving finger movement. When the information terminal 30 obtains an image including a subject performing an action involving finger movement captured by the camera 20, the information terminal 30 transmits the instruction received by the reception unit 34 and the image captured by the camera 20 to the video editing device 40. The evaluation unit 42e of the video editing device 40 determines whether or not the subject can perform an action involving opening and closing the hand, for example, by using another trained model (not shown) different from the trained model. The evaluation unit 42e may also use another trained model to determine whether the tip of the index finger and the tip of the thumb are touching in the image, and to identify the shape and size of the space between the index finger and the thumb to determine whether the finger opposition motion is possible. The other trained model may be stored in the storage unit 43 in advance.
[0183] This makes it possible to determine whether the subject can grasp an object, thereby making it possible to more accurately determine the degree of daily living activities that the subject can perform.
[0184] Furthermore, the information regarding the subject's physical functions may be stored in advance in the storage unit 43, or the information may be received from the user by the reception unit 34 and acquired by the acquisition unit 42a from the information terminal 30.
[0185] The evaluation unit 42e may also generate a rehabilitation training plan based on the assessment result. At this time, for example, the evaluation unit 42e may generate a rehabilitation training plan based on the subject's physical function in addition to the assessment result.
[0186] In the embodiment described above, two videos are edited by editing unit 42d so that their video playback times are aligned, and are simultaneously displayed on display unit 35. For example, three or more videos may be edited by editing unit 42d so that their video playback times are aligned, and are simultaneously displayed on display unit 35.
[0187] Furthermore, a plurality of moving images including, as subjects, a subject performing specific actions different from one another may be simultaneously displayed by display unit 35.
[0188] Furthermore, for example, the control unit 32 may link information indicating what specific action the subject is performing included in the moving image to the moving image and transmit the information to the video editing device 40 via the communication unit 31. For example, the control unit 32 may obtain the information from the user via the reception unit 34, or may identify what the specific action is based on the content of an instruction given by the instruction unit 36 and obtain the identified result as the information.
[0189] It is to be noted that step S104 does not have to be executed, in which case the video editing device 40 does not have to include the editing unit 42d.
[0190] Furthermore, in step S105, the evaluation result does not have to be output.
[0191] Furthermore, the video editing device 40 may repeatedly execute the processes of steps S100 to S104 as one loop process for each particular action performed by a subject person included in a video.
[0192] In addition, for example, in the above embodiment, the process executed by a specific processing unit may be executed by another processing unit. Furthermore, the order of multiple processes may be changed, or multiple processes may be executed in parallel.
[0193] Also, for example, in the above embodiment, each component of a processing unit such as the information processing unit 42 may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.
[0194] Furthermore, each component may be realized by hardware. Each component may be a circuit (or an integrated circuit). These circuits may form a single circuit as a whole, or each may be a separate circuit. Furthermore, each of these circuits may be a general-purpose circuit, or a dedicated circuit.
[0195] In addition, the general or specific aspects of the present invention may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.
[0196] Also, for example, the present invention may be realized as a video editing method, or as a program for causing a computer to execute the video editing method, or as a non-transitory recording medium readable by a computer on which such a program is recorded.
[0197] In the above embodiment, the display system 10 includes the information terminal 30 and the video editing device 40, but the display system according to the present invention may be realized as a single device such as an information terminal, or may be realized by multiple devices. For example, the display system may be realized as a client-server system. When the display system is realized by multiple devices, the components of the display system described in the above embodiment may be distributed among the multiple devices in any manner.
[0198] In addition, the present invention also includes forms obtained by applying various modifications to each embodiment that a person skilled in the art may think of, or forms realized by arbitrarily combining the components and functions of each embodiment within the scope that does not deviate from the spirit of the present invention. [Explanation of symbols]
[0199] 1. Target Audience 10 Display System 20 Camera 30 Information terminal 31, 41 Communications Department 32 Control section 35 Display section 36 Instruction section 40 Video editing equipment 42a Acquisition Department 42b Estimation part 42c Specific part 42d Editorial Department 42e Evaluation section 42f Output section 50, 51 Video
Claims
1. 1. A computer-implemented method for editing video, comprising the steps of: acquiring a first moving image including, as a subject, a subject who performed a specific action in a first period of time, and a second moving image including, as a subject, the subject who performed the specific action in a second period of time different from the first period; an estimation step of estimating a skeletal model of the subject in each of the first video sequence and the second video sequence; identifying a first image in the first video as a key image in which the subject performs a key action of the specific action, based on a skeletal model of the subject in the first video, and identifying a second image in the second video as the key image in which the subject performs the key action, based on a skeletal model of the subject in the second video; an editing step of editing at least one of the first video and the second video based on the first image and the second image to generate a third video based on the first video and a fourth video based on the second video so that their video playback times are aligned; an output step of outputting the third moving image and the fourth moving image, In the editing step, (i) the third video is edited so that the first image is included in the third video and the second image is included in the fourth video, and (ii) the at least one video is edited so that an image included in the at least one video is reduced or an image included in the at least one video is increased; The identifying step further includes identifying, based on a skeletal model of the subject, a start image in which the subject is performing a start motion to start the specific motion and an end image in which the subject is performing an end motion to end the specific motion, in at least one of the subjects; In the editing step, (i) editing the first video and the second video by deleting images before the start image and images after the end image from each of the first video and the second video, so as to leave the start image, the key image, and the end image; and (ii) editing at least one of the first video and the second video by deleting images before the start image and images after the end image, and then deleting one or more images between the start image and the end image in units of a predetermined number, so that the timings at which the start image, the key image, and the end image are played are aligned. How to edit videos.
2. In the editing step, at least one of a video following the first image in the first video and a video following the second image in the second video is deleted to edit the at least one of the video. The video editing method according to claim 1 .
3. In the editing step, the at least one of the images is edited by deleting a predetermined number of images included in the at least one of the images. The video editing method according to claim 1 or 2.
4. In the identifying step, the first image and the second image are identified based on at least one of a position of a plurality of skeletal points in a skeletal model of the subject, which corresponds to the key action, and an amount of change in the position. The video editing method according to claim 1 or 2.
5. further comprising an evaluation step of generating difference information indicating a difference between an evaluation point in the first video and an evaluation point in the second video, the evaluation point being at least any of: (iii) a position of any one of a plurality of skeletal points in a skeletal model of the subject, (iv) a positional relationship between two or more of the plurality of skeletal points, and (v) a moving speed of any one of the plurality of skeletal points, the evaluation point corresponding to the key action; The output step further includes outputting the difference information. The video editing method according to claim 1 or 2.
6. The evaluating step further includes generating evaluation information indicating an evaluation of the specific motion of the subject based on the difference; The output step further includes outputting the evaluation information. The video editing method according to claim 5.
7. an acquisition unit that acquires a first moving image including a subject who has performed a specific action in a first period of time and a second moving image including the subject who has performed the specific action in a second period of time different from the first period of time; an estimation unit that estimates a skeletal model of the subject in each of the first video and the second video; an identification unit that identifies a first image in the first video as a key image in which the subject performs a key action of the specific action, based on a skeletal model of the subject in the first video, and identifies a second image in the second video as the key image in which the subject performs the key action, based on a skeletal model of the subject in the second video; an editing unit that edits at least one of the first video and the second video based on the first image and the second image to generate a third video based on the first video and a fourth video based on the second video so that their video playback times are aligned; an output unit that outputs the third moving image and the fourth moving image, the editing unit (i) edits the at least one of the video streams so that the third video stream includes the first image and the fourth video stream includes the second image, and (ii) edits the at least one of the video streams so that an image included in the at least one of the video streams is reduced or an image included in the at least one of the video streams is increased; The identification unit further identifies, based on a skeletal model of the subject, a start image in which the subject is performing a start motion to start the specific motion and an end image in which the subject is performing an end motion to end the specific motion, in the at least one of the images; The editorial department: (i) editing the first video and the second video by deleting images before the start image and images after the end image from each of the first video and the second video, so as to leave the start image, the key image, and the end image; and (ii) editing at least one of the first video and the second video by deleting images before the start image and images after the end image, and then deleting one or more images between the start image and the end image in units of a predetermined number, so that the timings at which the start image, the key image, and the end image are played are aligned. Video editing equipment.
8. The video editing device according to claim 7 ; An information terminal, The video editing device further includes a first communication unit that communicates with the information terminal, The information terminal includes: A second communication unit that communicates with the video editing device; An instruction unit that instructs the subject to perform the specific action; a camera that captures an image of the subject performing the specific action to generate the first moving image and the second moving image; a control unit that outputs the first moving image and the second moving image to the video editing device via the second communication unit and acquires the third moving image and the fourth moving image from the video editing device via the second communication unit; a display unit that displays the third moving image and the fourth moving image. Display system.
Citation Information
Patent Citations
Apparatus and method for evaluation of achievement of target
JP2001000420A
Medical image display apparatus and program
JP2014138661A
Movement information processor and program
JP2014155693A
Dynamic image comparison device, method and program thereof, and dynamic image comparison system
JP2017229081A
Training evaluation device, method, and program
JP2020195573A