Information processing device, information processing method, and computer-readable non-transitory storage medium

The information processing device and method address the challenge of connecting multiple takes in motion capture by generating and superimposing planned motions, ensuring smooth transitions and reducing post-processing, thus improving video quality.

WO2025177733A1PCT designated stage Publication Date: 2025-08-28SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/001103
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-19
Filing Date
2025-01-16
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing motion capture technologies struggle to smoothly connect multiple takes in video production, particularly in scenarios where discrepancies in poses and movements between takes require complex post-processing to ensure natural-looking footage.

Method used

An information processing device and method that generates a second-take opening motion to smoothly connect first and second-half take motions, using a motion generation unit to infer and a presentation image generation unit to superimpose planned motions, supported by a motion capture system and display, facilitating real-time motion checking and adjustment.

Benefits of technology

This configuration reduces the need for extensive post-processing by ensuring seamless transitions between takes, maintaining natural motion continuity and allowing performers to easily align their movements with ideal poses, thereby enhancing the quality of the final video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025001103_28082025_PF_FP_ABST
    Figure JP2025001103_28082025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device includes a motion generating unit and a presentation video generating unit. The motion generating unit infers a motion in a beginning part of a second-half take, for smoothly connecting a first-half take motion adopted in a first-half take and a second-half take reference motion planned for the second-half take, as a second-half take beginning motion. The presentation video generating unit generates a presentation video by superimposing a video indicating the second-half take beginning motion onto a video indicating the motion of a performer.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and computer-readable non-transitory storage medium

[0001] The present invention relates to an information processing device, an information processing method, and a computer-readable non-transitory storage medium.

[0002] Motion capture is a well-known method of video production. Motion capture filming is performed in a dedicated studio, but depending on the size of the studio and the content of the filming, it may be necessary to film a series of movements in multiple takes. In this case, if the motions at the beginning and end of each segment are not smoothly connected, the footage will look unnatural. Hereinafter, the filming at the beginning of each segment will be referred to as the first take, and the filming at the end of each segment will be referred to as the second take.

[0003] Patent No. 7007330

[0004] It is known to use bamite to align the first and second takes. Bamite refers to tape stuck on the floor to indicate the performer's position. Bamite can specify the position of the feet, but cannot specify poses such as arms. If there is a discrepancy in the pose between takes, the motion data must be corrected after filming is complete.

[0005] Therefore, the present disclosure proposes an information processing device, an information processing method, and a computer-readable non-transitory storage medium that are capable of realizing smooth connections between takes.

[0006] According to the present disclosure, there is provided an information processing device having a motion generation unit that infers, as a second-take opening motion, a motion at the beginning of a second-take that smoothly connects a first-half take motion adopted in the first-half take and a second-half take reference motion planned for the second-half take, and a presentation image generation unit that generates a presentation image by superimposing an image showing the second-half take opening motion on an image showing the motion of the performer. Also, according to the present disclosure, there is provided an information processing method in which the information processing of the information processing device is executed by a computer, and a computer-readable non-transitory storage medium that stores a program that causes a computer to realize the information processing of the information processing device.

[0007] 1 is a diagram illustrating an example of division of motion. FIG. 1 is a diagram illustrating an example of division of motion. FIG. 2 is a diagram illustrating an example of the configuration of a motion assistance system. FIG. 3 is a diagram illustrating an example of display in superimposition mode. FIG. 4 is a diagram illustrating an example of display in rotation mode. FIG. 5 is a diagram illustrating an example of a skeleton superimposition method. FIG. 6 is a diagram illustrating variations of second half take reference motion. FIG. 7 is a diagram illustrating variations of second half take reference motion. FIG. 8 is a diagram illustrating an example of a processing flow of motion assistance. FIG. 9 is a diagram illustrating an example of a processing flow of motion assistance. FIG. 10 is a diagram illustrating an example of the hardware configuration of an information processing device.

[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0009] The explanation will be given in the following order: [1. Example of motion division] [2. Configuration of motion support system] [3. Display mode of presented image] [4. Variations of reference motion for second half take] [5. Processing flow] [6. Example of hardware configuration] [7. Effects]

[0010] 1. Examples of Motion Division FIGS. 1 and 2 are diagrams for explaining examples of motion division.

[0011] Motion capture records the movement of joints and other parts of the body as three-dimensional data. Filming takes place in a dedicated studio, but depending on the size and equipment of the studio and the content you want to film, it may be necessary to film a single continuous motion multiple times (splitting the motion).

[0012] FIG. 1 shows a scene of climbing stairs. When filming the motion of climbing stairs, boxes or the like are stacked in the studio to create multiple steps ST. The performer AC acts by imitating the stacked steps ST as stairs. In this case, taking into consideration the number of steps ST that can be prepared and safety aspects, it may be possible to prepare fewer steps than are necessary for the motion. In this case, a take TK from the start to the middle of climbing the stairs (first half take TK) is taken. 1 ), Take TK from the middle to the end of the climb (TK in the second half) 2 ), so it is necessary to take multiple shots.

[0013] The same thing happens in scenes where motions branch off depending on the scenario. Figure 2 shows a rock-paper-scissors scene. In rock-paper-scissors motions, three types of motions, "rock," "scissors," and "paper," branch off from a set pose (for example, a motion of "rock first" in which the hand AM for playing rock-paper-scissors is held out in front). To prevent players from predicting which hand will be played, it is desirable that the set pose be the same for each of "rock," "scissors," and "paper." In this case, the motion from which the branch originates (first half take motion M) 1real ) and the branching motion (later take motion M 2real ) must be photographed separately.

[0014] On set, motion adjustments between takes are made using Bamite. However, the scope of Bamite's use is limited to adjusting things like standing position, and there are many elements that are difficult to Bamite, such as the pose of the upper body other than the foot FT (for example, the angle of the elbow) and the speed of body movement. Any parts that cannot be adjusted using Bamite must be addressed by editing the motion data.

[0015] However, editing motion data can require complex work that takes into account the relationships between joints. For example, editing a shoulder can also move an elbow or hand, so editing one part can require adjustments to other parts that depend on it. The greater the discrepancy in motion between takes, the more time is required for adjustments, so it is best to keep the discrepancy as small as possible.

[0016] First half take TK1 The end of the song and the second half of the take TK 2 When using a motion that has a gap between the beginning and end of the first half, game engines such as Unity have a function to blend the motions together and connect them smoothly. By setting the blend rate change method (linear, etc.) and the time required for blending, the first half of the take motion M 1real Gradually take motion M 2real It can be played back so as to transition to

[0017] However, normal blending tends to result in simple movements, and the continuity between takes TK is not fully guaranteed. For example, 1 And the second half of the take TK 2 If the difference is large and the blend time is short, the motion will not transition smoothly. When blending cyclical movements, if the cycles do not match, the motion will look unnatural. For example, in a walking motion, if the timing of putting out the feet between takes TK is not matched, the motion will look like the feet are slipping. Therefore, even when using the blend function to transition between motions, it is necessary to adjust the blend time, etc.

[0018] The present disclosure has been made in consideration of the above-mentioned problems. 1 Filmed motion (first half take motion M 1real ) and the second half of the take TK 2 Information on the planned filming content (second half take reference motion M 2ref ) based on the first half of the take TK 1 The poses and movements are smoothly connected in the second half of the TK. 2 The motion at the beginning of the second half of the take (the motion at the beginning of the second half of the take) 2synth ) is generated. The generated second half take opening motion M 2synth The second half of the take TK 2 This will be displayed on the screen when shooting the second half of the movie. 2synth Acting in accordance with this is encouraged.

[0019] This configuration supports the performance of a performance in which the motions are smoothly connected between takes TK. Since there is less deviation in the motions between takes TK, the amount of post-processing required to correct the motion data is reduced. 2 The performance, including the opening part, is performed by the performer AC himself / herself. Therefore, motion data that reflects the individuality of the performer AC is obtained. Below, a motion support system for carrying out the processing of the present disclosure will be specifically described.

[0020] 2. Configuration of Motion Assistance System FIG. 3 is a diagram showing an example of the configuration of the motion assistance system 1. As shown in FIG.

[0021] The motion support system 1 has an information processing device 60, a display 30, a motion capture system 40, and a camera 50. The information processing device 60 generates support information for realizing smooth motion connections between takes TK based on the motion information of each take TK. The information processing device 60 presents the support information to the performer AC via the display 30, and 2 For example, the information processing device 60 includes a motion generation unit 10 and a presentation image generation unit 20.

[0022] The motion generation unit 10 generates the first half take motion M 1real And the second half take reference motion M 2ref First half take motion M 1real Actually, the first half of the take TK 1 This refers to the motion of the performer AC that was adopted in the second half of the take. 2ref The second half of the take TK 2 This means a sample of the motion that is planned. 2ref are used as samples of motion, and may be performed by a person other than the performer AC, or may be created using CG.

[0023] The motion generation unit 10 generates the first half take motion M 1real And the second half take reference motion M 2refThe second half of the take TK smoothly connects 2 The motion of the beginning of the second half of the take 2synth In order to make the transition between the takes TK smooth, the motion generator 10 infers that the first half take TK 1 And the second half of the take TK 2 The pose and the change of the pose at the boundary of the first half of the take motion M 1real The motion that matches the motion at the beginning of the second half of the take is M. 2synth It can be inferred as:

[0024] There are several possible ways to implement the motion generation unit 10. For example, a rule-based implementation that generates a motion that gradually transitions from the first half motion to the second half motion at a predetermined ratio, like Unity's blend motion creation, is possible.

[0025] Another possible implementation is to use a machine learning model that has been trained on similar past cases. An example of a method similar to the latter is described in the following document. This document describes a technology that uses a recurrent transition network (RTN) based on a long short term memory (LSTM) to output a movement motion that complements two movement motions.

[0026] [References] Recurrent Transition Networks for Character Locomotion / Felix G. Harvey and Christopher Pal: [Retrieved February 9, 2024], Internet, <https: / / arxiv.org / abs / 1810.02363>

[0027] Second half take reference motion M 2ref The first half of the take motion M 1real This means the following sample motion. In this disclosure, the first half take motion M 1real And the second half take reference motion M 2ref There is no time gap between them. 2synth Reference motion M2ref It is assumed that the beginning of the movie will be replaced. However, the first half of the take motion M 1real And the second half take reference motion M 2ref There may be a time gap between the first half take motion M 1real Immediately after that, the second half of the opening motion M 2synth Continued, the second half of the take opening motion M 2synth Immediately after that, the second half of the take reference motion M 2ref The following structure may also be used.

[0028] For example, the opening motion of the second half of the take, M 2synth The first half of the take motion M 1real For example, if the video is 30 fps, the motion for 10 to 15 frames is set as the motion M at the beginning of the second half of the take. 2synth If the video is 60 fps, 30 to 45 frames of motion will be the opening motion M of the second half take. 2synth This becomes:

[0029] The motion generation unit 10 generates the first half take TK 1 Take TK from the second half 2 The motion transition period is acquired as time information for realizing a smooth motion transition to the first take TK. The motion transition period means the period of the connecting portion between the takes TK where a smooth transition is required. For example, the motion transition period is 1 The number of frames t and the second take TK 2 The motion generation unit 10 generates a motion of the length indicated in the motion transition period as the beginning motion M of the second take. 2synth It is inferred as follows.

[0030] The presentation video generation unit 20 generates the second half take opening motion M 2synth The image showing the motion of the performer AC is superimposed on the image showing the motion of the performer AC to present the image I. p To facilitate checking of the motion, the presentation image generation unit 20 generates the motion M 2synth and presentation video I, which is a mirror image of the motion of performer AC.p can be generated as:

[0031] The presentation video generation unit 20 generates the second take TK 2 During the shooting of the image, the display 30 placed in front of the performer AC displays the presented image I p The performer AC displays the displayed presentation image I p Taking a look at the second half of the TK 2 This will be the first half of the performance. 1 And the second half of the take TK 2 The performance will be supported so that the performance can be smoothly connected. 2synth The motion of the performer AC to be superimposed is the first half take motion M 1real This also includes the static pose when aligning with the pose at the end of the image.

[0032] The display 30 displays various information for supporting the performer AC as support information. The support information includes the above-mentioned presentation image I p In addition, the display 30 may include background images to be used in the finished video, text information such as a scenario, etc. The display 30 has a screen large enough to display at least the entire motion. As the display 30, a known display such as an LCD (Liquid Crystal Display) or an OLED (Organic Light Emitting Diode) can be used.

[0033] The motion capture system 40 extracts multiple key points (multiple feature points indicating shoulders, elbows, wrists, hips, knees, ankles, etc.) from an image of a target person and records the motion of the target based on the relative positions of the key points. The key points may be detected based on markers attached to the target, or may be detected based on video analysis using AI (key point annotation).

[0034] The motion capture system 40 performs the second take TK 2 The motion capture system 40 detects the motion of the performer AC at time t in real time.2real,t Output as

[0035] The camera 50 captures the performance of the performer AC from in front of the performer AC. The camera 50 captures the image of the performer AC at time t as a performer image I. cam,t The output is as follows: Performer video I cam,t is displayed on the display 30 in a mirrored state. cam,t By reversing the left and right sides, the performer AC can check his or her own motion as if looking in a mirror.

[0036] 3. Display Mode of Presentation Image FIG. 4 and FIG. 5 show the display mode of presentation image I. p 10A and 10B are diagrams illustrating examples of display modes.

[0037] Presentation video I p There are two display modes for the image I: a superimposed mode and a rotated mode. p The camera image I' of the performer AC cam,t The rotation mode is a display mode in which the presented image I is superimposed on the p 4 is a diagram showing a display example in the superimposition mode. FIG. 5 is a diagram showing a display example in the rotation mode.

[0038] Camera image I' cam,t means an image based on the photographing data of the camera 50. For example, the performer image I at time t cam,t , or performer video I cam,t The image obtained by processing the camera image I' cam,t In the present disclosure, for example, a performer video I at time t is used as cam,t The image obtained by mirroring is the camera image I' cam,t The presentation image generating unit 20 converts the image of the performer AC captured by the camera 50 into the performer image I. cam,t The presentation image generation unit 20 acquires the performer image I cam,t The image obtained by mirror-reflecting the image is the camera image I'. cam,t Present as.

[0039] The presentation image generating unit 20 generates the presentation image I p The instruction information regarding the display viewpoint is set to the virtual camera setting Cedit If the performer AC has a preference for the angle of the motion shown on the display 30, the desired angle is acquired as the presentation image I. p By default, for example, the viewpoint of the camera 50 is input as the display viewpoint of the presented image I. p The correspondence between the camera coordinates and the coordinates of the display viewpoint is determined by the calibration data C obtained by calibrating the motion capture system 40 and the camera 50. def It is calculated based on the following.

[0040] The presentation image generation unit 20 acquires the viewpoint of the camera 50 that captures the performer AC. If the display viewpoint does not match the viewpoint of the camera 50, the presentation image I p The angle of view of the camera 50 also does not match the angle of view of the camera 50. Therefore, when the viewpoint of the camera 50 and the display viewpoint match, the presentation image generation unit 20 generates the camera image I′ of the performer AC. cam,t Presenting video I p When the viewpoint of the camera 50 and the display viewpoint do not match, the presentation image generation unit 20 superimposes the presentation image I p Camera footage I' of performer AC cam,t Stop superimposing (rotation mode).

[0041] The presentation image generation unit 20 can present the motion as a skeleton movement. For example, the presentation image generation unit 20 can present the motion M 2synth The ideal pose skeleton I shows the movement of the skeleton. synth The presentation image generation unit 20 presents the motion of the performer AC as a performer pose skeleton I, which indicates the movement of the skeleton of the performer AC. real Present as.

[0042] To facilitate motion checking, for example, the Ideal Pose Skeleton I synth and Performer Pose Skeleton I real The second half of the take begins with the motion M 2synthThe presentation image generating unit 20 generates the ideal pose skeleton I based on the display viewpoint instruction information. synth and Performer Pose Skeleton I real Rendering is performed.

[0043] To facilitate comparison between skeletons, the ideal pose skeleton I synth It is preferable that the second take TK is retargeted to match the physique of the performer AC. For example, the motion generation unit 10 generates the second take TK retargeted to match the physique of the performer AC. 2 Reference motion of the latter half of the reference motion M 2ref The motion generation unit 10 obtains the retargeted second-half take reference motion M 2ref and performer AC's first half take motion M 1real Based on this, the opening motion M of the second half take that matches the physique of the performer AC 2synth (Ideal Pose Skeleton I synth ) to generate the

[0044] The presentation image generator 20 generates an ideal pose skeleton I synth and Performer Pose Skeleton I real The presentation image generation unit 20 acquires a reference point (for example, a point indicating a waist joint) in each of the skeletons as a reference point. With the reference points of each skeleton aligned, the presentation image generation unit 20 generates an ideal pose skeleton I synth and Performer Pose Skeleton I real By superimposing the skeletons, the performer AC can easily understand the difference between his or her own motion and the ideal motion.

[0045] FIG. 6 is a diagram showing an example of a skeleton superimposition method.

[0046] The motion generation unit 10 generates an ideal pose at time t (the opening motion M of the second half take). 2synth,t ) and actual pose (second half take motion M 2real,t ) to the presentation image generating unit 20. The performer AC may input the virtual camera setting C edit is input to the presentation image generating unit 20. Virtual camera setting Cedit includes information about the rendering viewpoint of the pose (coordinates of the viewpoint and the line of sight (amount of rotation from the line of sight of the camera 50)). edit If there is no instruction regarding the viewpoint, the viewpoint of the camera 50 is used as the rendering viewpoint.

[0047] The presentation image generation unit 20 generates an ideal pose (the beginning motion M of the second half take) at time 0 (the beginning frame). 2synth,0 ) and the pose of the performer AC at the start of the shoot (the second half take motion M 2real,0 ) and the position of the beginning motion M of the second half take. 2synth,0 Reference point RP s and the second half of the take motion M 2real,0 Reference point RP r The plane coordinates of the second half of the take are the same. 2synth,0 In the example of Fig. 6, the position of the hip is the reference point.

[0048] If the height of the reference point is also aligned, the foot positions of each pose may not match. In this case, when the ideal pose is superimposed on the actual pose, the ideal pose may appear unnatural, either floating above the ground or sinking into the ground (see "Bad Example" in Figure 6). Therefore, it is preferable to align only in directions along coordinate axes other than height (see "Good Example" in Figure 6).

[0049] During filming, the opening motion of the second half take 2synth,t and the latter half of the take motion M 2real,t Skeleton (Ideal Pose Skeleton I) synth , Performer Pose Skeleton I real The presentation image generating unit 20 displays the virtual camera setting C edit The presentation image generating unit 20 renders each motion based on the skeletons. p In the superimposition mode, the presentation image generation unit 20 outputs the presentation image I p Further camera image I' cam,tare superimposed and displayed on the display 30.

[0050] 4. Variations of the second half take reference motions FIGS. 7 to 9 show the second half take reference motions M 2ref FIG.

[0051] Second half take reference motion M 2ref The second half of the take TK 2 This is a motion that is close to the motion that is planned to be shot. Reference motion for the second half of the take M 2ref For example, the data shown in the following examples "A" to "G" can be used. 1 The performer AC who performed in T and performer AC T A different performer AC from F It is written as follows.

[0052] <Typology of reference motions for the second half of the take> A. TK for the second half of the take during rehearsal 2 Performer AC T Motion data B. Another performer AC F The second half of the performance by TK 2 C. Motion data of the second half of the take selected by the creator from the motion data library 2 Motion data similar to the content of D. Second half take TK 2 The second half of the take was automatically searched for based on information such as the take name. 2 Motion data in the motion data library that is similar to the contents of E. The second half of the take TK, created by the creator by hand 2 Motion data similar to the contents of F. Second half take TK 2 The second half of the take was automatically generated by a machine learning model based on information such as the take name. 2 G. Motion data obtained by combining multiple data acquisition methods from types A to F above

[0053] Examples of motion data libraries for type C and type D include publicly available libraries, or proprietary libraries obtained from past filming results, etc. An example of a search for type D is when the take name is "jump," and so motion data named "jump" (motion data showing a jumping motion) is searched for in the motion data library.

[0054] For type A, performer AC during rehearsal is first half take TK 1 Performer AC T Therefore, without retargeting the motion during rehearsal, the reference motion M 2ref For type E, the creator can use the first half of the take TK. 1 Performer AC T If you create motion data that matches the physique of the character, you can use that motion data as it is without retargeting it and use it as the reference motion M for the second half of the take. 2ref (See FIG. 7).

[0055] Like Type B, take TK in the first half 1 A different performer AC F When using the motion data of the first half of the take TK, it is necessary to retarget the motion data according to the difference in physique (see Figure 8). 1 Performer AC T Since the above is different from the above, retargeting is required (see Figure 8).

[0056] For type G, the main motion data acquired is the first half of the take TK. 1 Performer AC T In the example of FIG. 9, if the creator creates motion data of type C or type D, it is necessary to retarget the data. F It has been manually adjusted to fit the body type.

[0057] 5. Processing Flow FIGS. 10 to 12 are diagrams showing an example of a processing flow of motion assistance.

[0058] The cameraman in the studio shoots the first half of a continuous motion (step S1). The motion of the performer AC in the first half of the shot is called the first half take motion M. 1real The motion generation unit 10 acquires the first half take motion M 1real And the second half take reference motion M 2ref Using the first half take motion M 1real The ideal second half take opening motion M that smoothly connects from 2synth is generated (step S2).

[0059] The cameraman starts shooting the second half of the continuous motion (step S3). The presentation image generation unit 20 generates a real-time pose of the performer AC (the second half take motion M). 2real,t ) and the pose at time t of the generated motion (the opening motion M 2synth,t ) is obtained (step S4).

[0060] The presentation image generating unit 20 acquires the current (time t) display mode (step S5). If the display mode is the superimposition mode (step S5: Yes), the presentation image generating unit 20 acquires the calibration data C def If the display mode is the rotation mode (step S5: No), the presentation image generator 20 acquires the angle of view of the camera 50 based on the virtual camera setting C edit The angle of view specified by is acquired (step S7).

[0061] The presentation image generation unit 20 generates the motion M at the beginning of the second take using the angle of view acquired in step S6 or step S7. 2synth,t and the latter half of the take motion M 2real,t The presentation image generation unit 20 renders the second half take motion M 2real,t Rendering the skeleton image I' real,t The presentation video generation unit 20 acquires the opening motion M 2synth,t Rendering the skeleton image I' synth,t is acquired (step S8).

[0062] If the display mode is the superimposition mode (step S9: Yes), the presentation image generation unit 20 outputs the real-time image of the performer AC (performer image I cam,t ) and performer video I cam,t Skeleton image I' real,t and skeleton image I' synth,t If the display mode is the rotation mode (step S9: No), the presentation image generation unit 20 generates a skeleton image I′ real,t and skeleton image I' synth,t A video image is created by superimposing the above (step S11).

[0063] The presentation image generation unit 20 outputs the mirror-inverted image of the image created in step S10 or step S11 to the display 30 (step S12). After that, the presentation image generation unit 20 advances the time t by one frame and returns to step S4. The presentation image generation unit 20 then repeats the above-described process at each time t.

[0064] 6. Example of Hardware Configuration FIG. 13 is a diagram showing an example of the hardware configuration of the information processing device 60. As shown in FIG.

[0065] The information processing of the information processing device 60 is realized by, for example, a computer 1000. The computer 1000 has a CPU (Central Processing Unit) 1100, a RAM (Random Access Memory) 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0066] The CPU 1100 operates and controls each component based on a program (program data 1450) stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the program stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0067] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 starts up, as well as programs that depend on the hardware of the computer 1000 .

[0068] The HDD 1400 is a non-transitory computer-readable recording medium that non-temporarily records programs executed by the CPU 1100 and data used by such programs. Specifically, the HDD 1400 is a recording medium that records an information processing program according to an embodiment as an example of program data 1450.

[0069] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0070] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display device, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.

[0071] For example, when the computer 1000 functions as the information processing device 60 according to the embodiment, the CPU 1100 of the computer 1000 executes an information processing program loaded onto the RAM 1200 to realize the functions of the aforementioned components. The information processing program, various models, and various data according to the present disclosure are stored in the HDD 1400. The CPU 1100 reads and executes the program data 1450 from the HDD 1400. Alternatively, the CPU 1100 may obtain these programs from another device via an external network 1550.

[0072] [7. Effects] The information processing device 60 has a motion generation unit 10 and a presentation image generation unit 20. The motion generation unit 10 generates a first half take TK. 1 The first half of the take motion M 1real And the second half of the take TK 2 The second half of the planned take reference motion M 2ref The latter half of the take TK smoothly connects 2 The motion of the beginning of the second half of the take 2synth The presentation video generation unit 20 infers that the beginning motion M 2synth The image showing the motion of the performer AC is superimposed on the image showing the motion of the performer AC to present the image I. p In the information processing method of the present disclosure, the processing of the information processing device 60 is executed by the computer 1000. The computer-readable non-transitory storage medium of the present disclosure stores a program that causes the computer 1000 to realize the processing of the information processing device 60.

[0073] This configuration provides motion support for smoothly connecting takes TK, making it possible to generate natural-looking video with minimal motion discrepancies between takes TK.

[0074] The motion generation unit 10 generates the first half take TK 1 And the second half of the take TK 2 The pose and the change of the pose at the boundary of the first half of the take motion M 1real The motion that matches the motion at the beginning of the second half of the take is M. 2synth It is inferred as follows.

[0075] According to this configuration, the motions between takes TK are smoothly connected.

[0076] The presentation video generation unit 20 generates the second half take opening motion M 2synth and presentation video I, which is a mirror image of the motion of performer AC. p Generate it as:

[0077] With this configuration, the performer AC can check his or her own motion as if looking in a mirror.

[0078] The presentation video generation unit 20 generates the second half take opening motion M 2synth The ideal pose skeleton I shows the movement of the skeleton. synth Present as.

[0079] According to this configuration, the motion M at the beginning of the second half take 2synth is accurately determined based on the movement of the skeleton.

[0080] The motion generation unit 10 generates the second take TK that is retargeted to fit the physique of the performer AC. 2 Reference motion of the latter half of the reference motion M 2ref The motion generation unit 10 obtains the retargeted second-half take reference motion M 2ref and performer AC's first half take motion M 1real Based on this, the opening motion M of the second half take that matches the physique of the performer AC 2synth Generate.

[0081] According to this configuration, performer AC performs the motion M 2synth It becomes easier to reproduce.

[0082] The presentation image generation unit 20 generates a presentation image of the motion of the performer AC as a performer pose skeleton I, which indicates the movement of the skeleton of the performer AC. real Present as.

[0083] This configuration makes it easier for the performer AC to compare his or her own pose with the ideal pose.

[0084] The presentation image generating unit 20 generates the presentation image I pThe presentation image generation unit 20 obtains instruction information relating to the display viewpoint of the ideal pose skeleton I based on the instruction information. synth and Performer Pose Skeleton I real Rendering is performed.

[0085] With this configuration, the performer AC can compare his or her own pose with the ideal pose from a preferred viewpoint.

[0086] The presentation image generating unit 20 acquires the viewpoint of the camera 50 capturing the image of the performer AC. When the viewpoint of the camera 50 and the display viewpoint match, the presentation image generating unit 20 generates a camera image I' of the performer AC. cam,t Presenting video I p is superimposed on.

[0087] This configuration clarifies the correspondence between each body part and the skeleton, making it easier to make body movements closer to ideal poses.

[0088] When the viewpoint of the camera 50 and the display viewpoint do not match, the presentation image generation unit 20 generates a presentation image I p Camera footage I' of performer AC cam,t Stop superimposing.

[0089] According to this configuration, the camera image I' cam,t This allows the display viewpoint to be changed quickly by omitting the processing (such as viewpoint conversion).

[0090] The presentation image generating unit 20 converts the image of the performer AC captured by the camera 50 into a performer image I. cam,t The presentation image generation unit 20 acquires the performer image I cam,t The image obtained by mirror-reflecting the image is the camera image I'. cam,t Present as.

[0091] With this configuration, the performer AC can check his or her own motion as if looking in a mirror.

[0092] The motion generation unit 10 generates the first half take TK 1 Take TK from the second half 2 The motion generation unit 10 obtains time information for realizing a smooth transition of the motion from the beginning of the second take to the beginning of the third take.2synth It is inferred as follows.

[0093] According to this configuration, the motions between takes TK are smoothly connected.

[0094] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0095] [Additional Notes] The present technology can also be configured as follows. (1) An information processing device having: a motion generation unit that infers, as a second-take opening motion, a motion at the beginning of a second-take that smoothly connects a first-half take motion adopted in the first-half take and a second-half take reference motion planned for the second-half take; and a presentation image generation unit that generates a presentation image by superimposing an image showing the second-half take opening motion on an image showing a performer's motion. (2) The information processing device described in (1) above, in which the motion generation unit infers, as the second-half take opening motion, a motion whose pose and how the pose changes at the boundary between the first and second takes match the first-half take motion. (3) The information processing device described in (1) or (2) above, in which the presentation image generation unit generates, as the presentation image, an image of a motion that is a mirror image of the second-half take opening motion and the performer's motion. (4) The information processing device according to any one of (1) to (3), wherein the presentation image generation unit presents the opening motion of the second half take as an ideal pose skeleton indicating skeleton movement. (5) The information processing device according to (4), wherein the motion generation unit acquires a reference motion of the second half take retargeted to match the physique of the performer as the second half take reference motion, and generates the opening motion of the second half take that matches the physique of the performer based on the retargeted second half take reference motion and the first half take motion of the performer. (6) The information processing device according to (4) or (5), wherein the presentation image generation unit presents the motion of the performer as a performer pose skeleton indicating movement of the performer's skeleton. (7) The information processing device according to (6), wherein the presentation image generation unit acquires instruction information related to a display viewpoint of the presentation image, and renders the ideal pose skeleton and the performer pose skeleton based on the instruction information. (8) The information processing device according to (7), wherein the presentation image generation unit acquires a viewpoint of a camera capturing an image of the performer, and when the viewpoint of the camera matches the display viewpoint, superimposes the camera image of the performer on the presentation image.(9) The information processing device according to (8) above, wherein the presentation image generation unit stops superimposing the camera image of the performer on the presentation image when the viewpoint of the camera and the display viewpoint do not match. (10) The information processing device according to (8) or (9) above, wherein the presentation image generation unit acquires an image of the performer taken by the camera as the performer image, and presents an image obtained by mirror-reflecting the performer image as the camera image. (11) The information processing device according to any one of (1) to (10) above, wherein the motion generation unit acquires time information for realizing a smooth transition of motion from the first take to the second take, and infers a motion of a length indicated in the time information as the opening motion of the second take. (12) An information processing method executed by a computer, comprising: inferring, as second-half take opening motion, a motion at the beginning of a second-half take that smoothly connects the first-half take motion used in the first-half take with the second-half take reference motion planned for the second-half take; and generating a presentation image by superimposing a video showing the second-half take opening motion on a video showing the performer's motion. (13) A computer-readable non-transitory storage medium storing a program that causes a computer to perform the above steps: inferring, as second-half take opening motion, a motion at the beginning of a second-half take that smoothly connects the first-half take motion used in the first-half take with the second-half take reference motion planned for the second-half take; and generating a presentation image by superimposing a video showing the second-half take opening motion on a video showing the performer's motion.

[0096] 10 Motion generation unit 20 Presentation image generation unit 50 Camera 60 Information processing device AC Performer I p Presentation video I cam,t Performer video I' cam,t Camera footage I real Performer Pose Skeleton I synth Ideal Pose Skeleton M 1real First half take motion M 2ref Second half take reference motion M 2synth Motion at the beginning of the second half of the take TK1 First half take TK 2 Second take

Claims

1. An information processing device having: a motion generation unit that infers the motion at the beginning of a second take, which smoothly connects the first half take motion adopted in the first half take and the second half take reference motion planned for the second half take, as the second half take opening motion; and a presentation image generation unit that generates a presentation image by superimposing an image showing the second half take opening motion on an image showing the actor's motion.

2. The information processing device according to claim 1, wherein the motion generation unit infers a motion whose pose and pose change pattern at the boundary between the first and second takes match the first-half take motion as the opening motion of the second-half take.

3. The information processing device according to claim 1, wherein the presentation image generation unit generates, as the presentation image, an image of the opening motion of the second take and the motion of the performer mirror-inverted.

4. The information processing device according to claim 1, wherein the presentation image generation unit presents the opening motion of the second take as an ideal pose skeleton that indicates skeleton movement.

5. The information processing device described in claim 4, wherein the motion generation unit acquires the reference motion of the second half take retargeted to match the physique of the performer as the second half take reference motion, and generates the opening motion of the second half take that matches the physique of the performer based on the retargeted reference motion of the second half take and the first half take motion of the performer.

6. The information processing device according to claim 4, wherein the presentation image generation unit presents the motion of the performer as a performer pose skeleton that indicates the movement of the performer's skeleton.

7. The information processing device according to claim 6, wherein the presentation image generation unit acquires instruction information relating to a display viewpoint of the presentation image, and renders the ideal pose skeleton and the performer pose skeleton based on the instruction information.

8. The information processing device described in claim 7, wherein the presentation image generation unit acquires the viewpoint of a camera capturing the performer, and when the viewpoint of the camera matches the display viewpoint, superimposes the camera image of the performer on the presentation image.

9. The information processing device according to claim 8, wherein the presentation image generation unit stops superimposing the camera image of the performer on the presentation image when the viewpoint of the camera and the display viewpoint do not match.

10. The information processing device according to claim 8, wherein the presentation image generation unit acquires an image of the performer captured by the camera as the performer image, and presents an image obtained by mirror-reflecting the performer image as the camera image.

11. The information processing device according to claim 1, wherein the motion generation unit acquires time information for realizing a smooth transition of motion from the first take to the second take, and infers a motion of a length indicated by the time information as the opening motion of the second take.

12. An information processing method executed by a computer, comprising: inferring the motion at the beginning of a second take that smoothly connects the first half take motion adopted in the first half take with the second half take reference motion planned for the second half take as the second half take opening motion; and generating a presentation image by superimposing an image showing the second half take opening motion on an image showing the actor's motion.

13. A computer-readable non-transitory storage medium storing a program that causes a computer to perform the following: inferring the motion at the beginning of a second half take that smoothly connects the first half take motion used in the first half take with the second half take reference motion planned for the second half take as the second half take opening motion; and generating a presentation image by superimposing an image showing the second half take opening motion on an image showing the actor's motion.

Citation Information

Patent Citations

  • Information processing method, information processing apparatus, program and computer-readable recording medium

    JP2010069102A

  • Training evaluation device, method, and program

    JP2020195573A

  • Method and system for generating a visual effect of object animation

    US20200027258A1

  • Terminal device and support method for improving form

    WO2022044399A1