Video generation device and its program

The video generation device addresses the challenge of character rotation by calculating rotation angles and generating both front and rear views, enhancing realism through three-dimensional representation.

JP2026061007APending Publication Date: 2026-04-09NIPPON HOSO KYOKAI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional methods for generating videos with hand-drawn characters struggle to reproduce character rotation, as they handle characters in 2D and cannot effectively depict changes in orientation, such as facing forward and then turning backward.

Method used

A video generation device that includes a character image processing unit, an action image processing unit, a motion data storage unit, a joint point assignment unit, a frame division unit, a character rotation unit, and a back-side image generation unit to generate a video that reproduces character rotation by calculating rotation angles and creating both front and rear views.

Benefits of technology

The device effectively reproduces character rotation and expresses three-dimensionality, enhancing the realism of the video by depicting both the front and back of the character, even when facing diagonally, using trapezoidal deformation and rear-view images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026061007000001_ABST
    Figure 2026061007000001_ABST
Patent Text Reader

Abstract

We provide a video generation device that can reproduce character rotation. [Solution] The video generation device 1 comprises a character image processing unit 16 that detects the joint points of a character, an action image processing unit 11 that generates motion data, a joint point assignment unit 19 that assigns the joint points of the character to the joint points of a person, a first video generation unit 20 that generates a first video in which the character is performing an action, a character rotation unit 22 that calculates the rotation angle of the person and generates a rotation image, a back-side image generation unit 23 that generates a back-side image when the character is facing backward, and a second video generation unit 24 that generates a second video composed of the rotation image and the back-side image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a video generation device and a program therefor. [Background technology]

[0002] Methods for generating videos in which hand-drawn characters come to life have been proposed before (for example, Non-Patent Document 1). The method described in Non-Patent Document 1 involves uploading an image of a hand-drawn character to a server, identifying the character in the image, and generating a video in which the character performs a predetermined action. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] “Animated Drawings”, [Searched August 1, 2024], Internet<URL:https: / / sketch.metademolab.com / > [Overview of the project] [Problems that the invention aims to solve]

[0004] In the conventional technology described above, hand-drawn characters are handled in 2D, and the video is generated with the character always facing forward. Therefore, the conventional technology has the problem that it is difficult to reproduce character rotation, such as when a character that is facing forward turns its back and then turns back to face forward.

[0005] The present invention aims to provide a video generation device and program capable of reproducing the rotation of a character. [Means for solving the problem]

[0006] To solve the aforementioned problems, the present invention provides a video generation device that generates a video of a character performing a predetermined action from a character image in which a character is depicted, and comprises a character image processing unit, an action image processing unit, a motion data storage unit, a joint point assignment unit, a first video generation unit, a frame division unit, a character rotation unit, a back side image generation unit, and a second video generation unit.

[0007] The character image processing unit detects the joint points of the character from the character image. The motion image processing unit detects the joint points of a person from motion images of a person performing an action and generates motion data. The motion data storage unit stores motion data for each movement. The joint point assignment unit assigns the joint points of the character to the joint points of the human. The first video generation unit generates a first video in which a character, to which joint points of a person have been assigned, performs an action based on motion data of a pre-selected action.

[0008] The frame division section divides the first video into individual frames. The character rotation unit calculates the rotation angle of the character based on the pre-defined positional relationships between the character's joint points for each frame, and generates a rotated image of the character rotated at the calculated rotation angle. The back-view image generation unit determines, frame by frame, whether the character is facing backward based on its rotation angle. If the character is facing backward, it generates a back-view image representing the character's silhouette. The second video generation unit generates a second video consisting of a rotated image and a back-side image.

[0009] With this configuration, the video generation device generates a second video consisting of a rotated image representing the front of the character and a rear-view image representing the back of the character, according to a pre-selected operation, thereby reproducing the rotation of the character.

[0010] Note that the present invention can also be implemented by a program for causing a computer to function as the above-described video generation device.

Advantages of the Invention

[0011] According to the present invention, the rotation of a character can be reproduced.

Brief Description of the Drawings

[0012] [Figure 1] It is a block diagram showing the configuration of the video generation device according to the embodiment. [Figure 2] (a) to (c) are diagrams for explaining the detection of human joint points in the embodiment. [Figure 3] (a) to (c) are diagrams for explaining the detection of character joint points in the embodiment. [Figure 4] It is a diagram for explaining the calculation of the rotation angle of a human in the embodiment. [Figure 5] (a) to (d) are diagrams for explaining the generation of a rotated image and a rear image in the embodiment. [Figure 6] It is a diagram for explaining the generation of a rear image in the embodiment. [Figure 7] It is a flowchart showing the operation of the video generation device when generating motion data in the embodiment. [Figure 8] It is a flowchart showing the operation of the video generation device when generating a second video in the embodiment.

Embodiments for Carrying Out the Invention

[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, each of the embodiments described below is for embodying the technical idea of the present invention, and the present invention is not limited to the following unless specifically described. Also, the same means may be denoted by the same reference numerals, and the description may be omitted.

[0014] [Configuration of Video Generation Device] Referring to Figure 1, the configuration of the video generation device 1 according to this embodiment will be described. The video generation device 1 generates a video of a character performing a predetermined action from a character image in which the character is depicted. In other words, when a character image is input to the video generation device 1, it generates a video in which the character contained in that character image is acting like a human. For example, the character image is an image of a hand-drawn character taken with a digital camera or smartphone.

[0015] As shown in Figure 1, the video generation device 1 comprises a motion image input unit 10, a motion image processing unit 11, a motion data storage unit 14, a character image input unit 15, a character image processing unit 16, a joint point assignment unit 19, a first video generation unit 20, a frame division unit 21, a character rotation unit 22, a back side image generation unit 23, and a second video generation unit 24.

[0016] The motion image input unit 10 receives motion images in order to generate the motion data described later. In this embodiment, the motion image input unit 10 receives motion images representing various actions. The motion images are moving images of a person performing some kind of action (for example, dancing, walking around, doing gymnastics). The motion image input unit 10 outputs the motion image to the motion image processing unit 11 (joint point 2D coordinate detection unit 12).

[0017] The motion image processing unit 11 detects the joint points of a person from motion images of a person performing an action and generates motion data. As shown in Figure 1, the motion image processing unit 11 includes a joint point 2D coordinate detection unit 12 and a joint point 3D coordinate transformation unit 13.

[0018] The joint point 2D coordinate detection unit 12 detects the joint points of a person included in the motion image input from the motion image input unit 10 in a 2D coordinate system. For example, the joint point 2D coordinate detection unit 12 detects the joint points of a person using the method described in references 1 and 2. As shown in Figure 2(a), a motion image D in which person H raises both arms will be explained as an example. In this case, the joint point 2D coordinate detection unit 12 detects predetermined joint points P from the motion image D in Figure 2(a), such as the joint point P1 of person H's right shoulder and the joint point P2 of person H's left shoulder, as shown in Figure 2(b). In Figure 2(b), each joint point P of person H is shown as a black circle. The joint point 2D coordinate detection unit 12 outputs the joint points of the person detected in the 2D coordinate system to the joint point 3D coordinate transformation unit 13.

[0019] Reference 1: “ROKOKO”, [Accessed August 8, 2024], Internet<URL:https: / / www.rokoko.com / ja> Reference 2: “YOLOv3: An Incremental Improvement”, Joseph Redmon, Ali Farhadi, [Retrieved August 8, 2020], Internet<URL:https: / / arxiv.org / pdf / 1804.02767>

[0020] The joint point 3D coordinate transformation unit 13 transforms the joint points of a person detected in a 2D coordinate system into a 3D coordinate system. For example, using the methods described in References 1 and 3, the joint point 3D coordinate transformation unit 13 first estimates the posture of person H in 3D, as shown in Figure 2(c), and then transforms the 2D coordinates of the joint points into 3D coordinates. The joint point 3D coordinate transformation unit 13 arranges the joint points of the person, which have been transformed into a 3D coordinate system, in chronological order and writes them as motion data to the motion data storage unit 14.

[0021] Reference 3: “Exploiting Temporal Contexts with Strided Transformer for 3D Human Pose Estimation”, [Retrieved September 9, 2024], Internet<URL:https: / / github.com / Vegetebird / StridedTransformer-Pose3D>

[0022] Here, as shown in Figure 2(a), we have explained a portion of the action performed by person H (one frame of motion image D). In reality, the motion image processing unit 11 performs the same processing for a series of actions (all frames of motion image D) from the start to the end of person H's action. Furthermore, the motion image processing unit 11 can generate motion data for various actions such as dancing, walking around, and gymnastics, and write it to the motion data storage unit 14 in advance.

[0023] The motion data storage unit 14 stores motion data for each action. For example, the motion data storage unit 14 is an HDD (Hard Disk Drive), SSD (Solid State Drive), or memory that stores motion data for various actions such as dancing, walking around, and gymnastics.

[0024] The character image input unit 15 receives a character image in which a character is depicted. The character image input unit 15 also outputs the character image to the character image processing unit 16 (character area detection unit 17).

[0025] The character image processing unit 16 detects the joint points of a character from a character image. As shown in Figure 1, the character image processing unit 16 includes a character region detection unit 17 and a character joint point detection unit 18.

[0026] The character region detection unit 17 detects character regions included in the character image input from the character image input unit 15. A character region is a region (mask region) that represents a character included in the character image. The image representing the character region (mask region) included in the character image is called the mask image. In this embodiment, the character region detection unit 17 uses the method described in Non-Patent Literature 1 to detect the mask region from the character image and generate a mask image. The character region detection unit 17 then outputs the mask region of the generated mask image as a character region to the character joint point detection unit 18.

[0027] As shown in Figure 3(a), consider the case where character T represents a hand-drawn bear in character image C. In this case, as shown in Figure 3(c), the character region detection unit 17 detects the white region representing the bear as a mask region and generates a mask image M. The character region detection unit 17 then outputs the white region of the mask image M as the character region to the character joint point detection unit 18.

[0028] The character joint point detection unit 18 detects the joint points of a character in a two-dimensional coordinate system from the character region. The character joint point detection unit 18 can detect the joint points of a character from the character region detected by the character region detection unit 17 using any method (for example, the method described in Non-Patent Document 1). For example, the character joint point detection unit 18 detects predetermined joint points R in a two-dimensional coordinate system from the character region in Figure 3(a), such as the joint point R1 of the bear's right shoulder and the joint point R2 of its left shoulder, as shown in Figure 3(b). In Figure 3(b), each joint point R of the bear is shown as a black circle. The character joint point detection unit 18 outputs each joint point of the character detected in a two-dimensional coordinate system to the joint point assignment unit 19.

[0029] The joint point assignment unit 19 assigns the joint points of a character to the joint points of a person. The joint point assignment unit 19 uses an arbitrary method (for example, deep learning) to associate each joint point of a person included in the motion data of the motion data storage unit 14 with each joint point of a character detected by the character joint point detection unit 18. In this embodiment, the joint point assignment unit 19 assigns the joint point R of the bear in Figure 3(b) to the joint point P of the person H in Figure 2(b). The joint point assignment unit 19 outputs the joint point assignment result to the first video generation unit 20.

[0030] Furthermore, because the skeletal structure of characters, such as fish or automobiles, differs significantly from that of humans, it may not be possible to correctly assign joint points. In this case, the joint point assignment unit 19 may manually assign the joint points.

[0031] The first video generation unit 20 generates a first video in which a character, to which joint points of a person are assigned, performs an action based on motion data of a pre-selected action. For example, a user of the video generation device 1 selects a desired action from among the actions corresponding to the motion data in the motion data storage unit 14. Then, the first video generation unit 20 generates a first video in which the character performs the selected action. For example, the first video generation unit 20 can generate the first video using the method described in Non-Patent Literature 1. It is assumed that the first video is pre-assigned information representing the order of each frame (for example, a time code). The first video generation unit 20 outputs the generated first video to the frame division unit 21.

[0032] Furthermore, in the first video, the character is always depicted facing forward. This means that even if the person who is the source of the motion data turns their back in the first video, the character will still appear to be facing forward. Also, the first video does not express the three-dimensionality of the character. The three-dimensionality (perspective) of a character is expressed so that the part of the character closer to the viewer is larger and the part further away is smaller.

[0033] The frame division unit 21 divides the first video into individual frames. In other words, the frame division unit 21 divides the first video input from the first video generation unit 20 into frame units. The frame division unit 21 then outputs each frame constituting the first video to the character rotation unit 22.

[0034] The character rotation unit 22 calculates the rotation angle of the character for each frame based on the pre-set positional relationship between the character's joint points, and generates a rotated image of the character rotated at the calculated rotation angle. In this embodiment, the character rotation unit 22 reads motion data corresponding to the character's movement from the motion data storage unit 14 and calculates the rotation angle (depth angle) of the character for each frame input from the frame division unit 21.

[0035] <Calculation of the rotation angle of person H> Referring to Figure 4, we will now explain in detail how to calculate the rotation angle of person H. In Figure 4, the X-axis represents the horizontal direction, the Y-axis represents the vertical direction, and the Z-axis represents the depth direction. Also, the line segment passing through the joint point R1(X1,Z1) of the right shoulder and the joint point R2(X2,Z2) of the left shoulder of person H is N. 12 Let X1 represent the horizontal coordinate of joint point R1, and Z1 represent the depth coordinate of joint point R1. Also, X2 represents the horizontal coordinate of joint point R2, and Z2 represents the depth coordinate of joint point R2. Furthermore, let N be a line segment parallel to the X axis. x Let the line segment N 12 and line segment N x Let Q0(0,0) be the intersection point with [the other line].

[0036] Line segment N 12 and line segment N x Let the angle between the line segment N be the rotation angle θ of person H. When person H is facing forward, that is, line segment N 12 is line segment N x When parallel to the given plane, the rotation angle of person H is set to θ = 0°.

[0037] The character rotation unit 22 calculates the rotation angle θ of the person H from the horizontal coordinates and depth coordinates of two joint points by using trigonometric functions. In the present embodiment, the character rotation unit 22 uses the joint points R1(X1, Z1) , R2(X2, Z2) of both shoulders of the person H to calculate the rotation angle θ of the person H.

[0038] First, the character rotation unit 22 obtains a line segment N 12 that passes through the joint point R1(X1, Z1) and the joint point R2(X2, Z2). Next, the character rotation unit 22 obtains an intersection point Q2(X2, 0) of a line segment extended in the Z-axis direction from the joint point R2(X2, Z2) and the line segment N x . Next, the character rotation unit 22 obtains a length L X in the X-axis direction and a length L Z in the Z-axis direction from the intersection point Q0(0, 0) to the joint point R2(X2, Z2).

[0039] That is, the length L X represents the length from the intersection point Q0(0, 0) to the intersection point Q2(X2, 0). Also, the length L Zは、 represents the length from the intersection point Q2(X2, 0) to the joint point R2(X2, Z2). Therefore, the character rotation unit 22 calculates the rotation angle θ of the person H from the lengths L X , L Y by using a calculation formula represented by the inverse tangent function.

[0040] Note that, in order to calculate the rotation angle θ of the person H, the joint points R1 and R2 of both shoulders have been described as being set in advance, but it is not limited to this. For example, in order to calculate the rotation angle θ of the person H, the joint points R of both arms or both legs may be set in advance.

[0041] <Generation of Rotated Image> Referring to FIG. 5, the generation of the rotated image will be specifically described. In the upper part of FIG. 5, a person H performing an action, which is the source of the motion data, is illustrated. Also, in the lower part of FIG. 5, a rotated image or a back-side image in which the character T is rotating in accordance with the action of the person H is illustrated. Hereinafter, it will be described that the rotation angles θ of the person H and the character T are equal.

[0042] The character rotation unit 22 generates a rotated image F by deforming the frame into a trapezoidal shape so that the front side of character T is larger and the back side of character T is smaller, according to the rotation angle θ of character H.

[0043] As shown in Figure 5(a), consider the case where person H is facing forward and the rotation angle θ of character T is 0°. In this case, the character rotation unit 22 does not need to deform the frame of the first animation into a trapezoidal shape, so it generates a rectangular rotated image F1. In the rotated image F1, the length of the left side is L. L and the length L of the right-hand side R They become equal, and the width does not change from the original frame.

[0044] As shown in Figure 5(b), consider the case where person H is facing diagonally forward and the rotation angle of character T is θ = 45°. In this case, the character rotation unit 22 generates a rotated image F2 by deforming the frame of the first video into a trapezoid. In the rotated image F2, the length L of the right side, which is the far side, is... R The length L of the left side, which is the front side. L It becomes shorter than [original]. Also, in the rotated image F2, the width W becomes shorter than the original frame. Note that the length L on the right side depends on the rotation angle θ of the character T. R The length L of the left side L The ratio for shortening the frame, and the ratio for reducing the frame width W, can be set manually.

[0045] Subsequently, the character rotation unit 22 outputs the generated rotated image F to the second video generation unit 24. The character rotation unit 22 also outputs the rotation angle θ of the rotated image F and character T to the back-side image generation unit 23.

[0046] Returning to Figure 1, we will continue the explanation of the video generation device 1. The rear-view image generation unit 23 determines, frame by frame, whether the character is facing backward based on its rotation angle, and if the character is facing backward, it generates a rear-view image representing the character's silhouette.

[0047] <Generating the back side image> The generation of the back side image will be explained in detail with reference to Figures 5 and 6. First, the back-view image generation unit 23 determines whether character T is facing backward for each frame of the first video. Then, if character T is facing backward (90°≦θ<270°), the back-view image generation unit 23 generates a back-view image B from the rotated image F.

[0048] Consider the case where character T is facing the back, as shown in Figure 5(c) or Figure 5(d). In this case, the back-side image generation unit 23 calculates rotated images F1 and F2 in which character T is symmetrical, using a rotation angle θ=90 as a reference. Then, the back-side image generation unit 23 flips the calculated rotated images F1 and F2 horizontally, extracts the outline of character T from the horizontally flipped rotated images F3 and F4, and fills the inside of the outline with a predetermined color to generate back-side images B3 and B4.

[0049] As shown in Figure 5(c), when the rotation angle θ of character T is 135°, the rotated image F2 with a rotation angle θ = 45° is symmetrical, as shown in Figure 6(a). Therefore, as shown in Figure 6(b), the character rotation unit 22 flips the rotated image F2 horizontally to generate the rotated image F3. Next, as shown in Figure 6(c), the character rotation unit 22 extracts the contour of character T from the rotated image F3. At this time, it is preferable for the character rotation unit 22 to perform a closing process to reduce noise contained in the contour line (see enlarged view of Figure 6(c)). Next, as shown in Figure 6(d), the character rotation unit 22 fills the inside of the contour of character T with a predetermined color (for example, gray) to generate the back side image B3. Note that the color used to fill character T in the back side image is not limited to gray.

[0050] As shown in Figure 5(d), when the rotation angle θ of character T is 180°, the rotated image F1 with a rotation angle θ = 0° is symmetrical, as shown in Figure 5(a). Therefore, the character rotation unit 22 generates the back side image B4 from the rotated image F1 using the same procedure as for the back side image B3.

[0051] In other words, as shown in Figure 5(c), the back image B3 is an image created by horizontally flipping the rotated image F2 and filling in character T with gray. Also, as shown in Figure 5(d), the back image B4 is an image created by horizontally flipping the rotated image F1 and filling in character T with gray. In this way, the back of character T is filled with gray, which gives a sense of three-dimensionality and makes it easier for the user to understand that character T is rotating.

[0052] Subsequently, the rear image generation unit 23 outputs the generated rear image to the second video generation unit 24.

[0053] Returning to Figure 1, we will continue the explanation of the video generation device 1. The second video generation unit 24 generates a second video consisting of a rotated image and a back-side image. In other words, the second video generation unit 24 generates the second video by arranging the rotated image input from the character rotation unit 22 and the back-side image input from the second video generation unit 24 in chronological order. In the example in Figure 5, the second video generation unit 24 generates the rotated images F1, ..., F 2,…, A second video is generated, consisting of rear view images B3, ..., and rear view image B4 arranged in sequence.

[0054] Thus, the second video depicts not only the front but also the back of the character. Furthermore, the second video expresses the character's three-dimensionality (perspective) when the character T is facing diagonally forward or diagonally backward.

[0055] [Video generation device operation: Motion data generation] Referring to Figure 7, the operation of the video generation device 1 when generating motion data will be explained. As shown in Figure 7, in step S1, an action image is input to the action image input unit 10. In step S2, the joint point 2D coordinate detection unit 12 detects the joint points of the person included in the motion image input in step S1 in a 2D coordinate system.

[0056] In step S3, the joint point 3D coordinate transformation unit 13 transforms the joint points of the person detected in the 2D coordinate system into a 3D coordinate system. In step S4, the joint point 3D coordinate transformation unit 13 writes the joint points of the person, which have been transformed into a 3D coordinate system, as motion data to the motion data storage unit 14.

[0057] [Video generation device operation: Generation of second video] Referring to Figure 8, the operation of the video generation device 1 when it generates the second video will be explained. As shown in Figure 8, in step S10, a character image is input to the character image input unit 15. In step S11, the character region detection unit 17 detects the character region included in the character image input in step S10.

[0058] In step S12, the character joint point detection unit 18 detects the joint points of the character in a two-dimensional coordinate system from the character area. In step S13, the joint point assignment unit 19 assigns the character's joint points to the human's joint points.

[0059] In step S14, the first video generation unit 20 generates a first video in which a character to which joint points of a person have been assigned is performing an action, based on motion data of a pre-selected action. In step S15, the frame division unit 21 divides the first video into individual frames. In step S16, the character rotation unit 22 calculates the rotation angle of the character from the positional relationship between the joint points of the character that have been set in advance, and generates a rotated image in which the character has rotated at the calculated rotation angle.

[0060] In step S17, the back-view image generation unit 23 determines whether the character is facing backward based on the rotation angle of the person, and if the character is facing backward, it generates a back-view image representing the silhouette of the character. In step S18, the second video generation unit 24 generates a second video consisting of a rotated image and a back-side image.

[0061] [effect] As described above, the video generation device 1 generates a second video consisting of a rotated image representing the front of the character and a rear-view image representing the back of the character, according to a pre-selected operation, thereby reproducing the rotation of the character. Furthermore, the video generation device 1 can also express a sense of three-dimensionality by deforming the frame into a trapezoidal shape so that the character in front is larger and the character in the background is smaller. Thus, the video generation device 1 can reproduce the rotation of characters and express a sense of three-dimensionality, thereby improving the sense of realism.

[0062] The second video generated by the video generation device 1 can be combined with other videos using video editing software. For example, the background of the second video can be changed, multiple characters can be drawn, or it can be combined with images of real people in motion. Furthermore, if motion data representing local dances (for example, Awa Odori) is generated, the second video can be used as content that conveys the charm of a region, such as having characters move in accordance with the local dance.

[0063] Although embodiments have been described in detail above, the present invention is not limited to the embodiments described above, and includes design changes and the like that do not depart from the spirit of the present invention.

[0064] In the embodiments described above, the video generation device was described as an independent piece of hardware, but the present invention is not limited thereto. For example, the present invention can also be realized by a program that causes hardware resources such as the CPU, memory, and hard disk of a computer to function as the video generation device described above. This program may be distributed via a communication line, or it may be written to a recording medium such as a CD-ROM or flash memory and distributed. [Explanation of Symbols]

[0065] 1. Video generation device 10 Motion Image Input Section 11 Motion Image Processing Unit 12. Joint point 2D coordinate detection unit 13. Joint Point 3D Coordinate Transformation Unit 14 Motion data storage unit 15 Character Image Input Section 16 Character Image Processing Unit 17 Character Area Detection Unit 18 Character joint point detection unit 19. Joint point assignment area 20. First Video Generation Unit 21 Frame division section 22 Character rotation section 23 Rear side image generation unit 24 Second Video Generation Unit

Claims

1. A video generation device that generates a video of a character performing a predetermined action from a character image in which the character is depicted, A character image processing unit that detects the joint points of the character from the character image, A motion image processing unit that detects the joint points of a person performing the aforementioned action from an image of the person performing the aforementioned action and generates motion data, A motion data storage unit that stores the motion data for each of the aforementioned operations, A joint point assignment unit that assigns the joint points of the character to the joint points of the person, A first video generation unit generates a first video in which a character to which the joint points of the person are assigned performs the action, based on motion data of the action that has been selected in advance. A frame division unit that divides the aforementioned first video into frames, A character rotation unit calculates the rotation angle of the character from the positional relationship between the joint points of the character, which is predetermined for each frame, and generates a rotated image of the character rotated by the calculated rotation angle. For each frame, a back-side image generation unit determines from the rotation angle whether the character is facing backward, and if the character is facing backward, generates a back-side image representing the silhouette of the character. A second video generation unit generates a second video composed of the aforementioned rotated image and the aforementioned rear view image, A video generation device characterized by comprising the following features.

2. The character rotation unit is characterized in that it calculates the rotation angle from the horizontal and depth coordinates of the two joint points using trigonometric functions, as described in claim 1.

3. The video generation apparatus according to claim 1, characterized in that the character rotation unit generates the rotated image by deforming the frame into a trapezoidal shape such that the front side of the character is larger and the back side of the character is smaller, according to the rotation angle.

4. The video generation apparatus according to claim 1, characterized in that the rear image generation unit generates the rear image by horizontally flipping the rotated image, extracting the outline of the character from the horizontally flipped rotated image, and filling the inside of the outline with a predetermined color.

5. The aforementioned character image processing unit is A character region detection unit for detecting character regions included in the character image, A character joint point detection unit detects the joint points of the character in a two-dimensional coordinate system from the character region, The video generation device according to claim 1, characterized by comprising the following:

6. The aforementioned motion image processing unit is A joint point two-dimensional coordinate detection unit that detects the joint points of a person included in the aforementioned motion image in a two-dimensional coordinate system, A joint point 3D coordinate transformation unit that transforms the joint points of the person detected in a 2D coordinate system into a 3D coordinate system, The video generation device according to claim 1, characterized by comprising the following:

7. A program for causing a computer to function as a video generation device according to any one of claims 1 to 6.