Information processing method, information processing system, and program
The information processing method addresses the issue of inaccurate motion capture by using a feature point generation model to correct and transfer accurate feature point information to a character object, ensuring smooth and natural movement reproduction.
Patent Information
- Application Number
- JP2022076801
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-08
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2042-05-08
AI Technical Summary
Existing motion capture technologies fail to accurately reproduce the movements of a person in a character object, leading to incorrect and unnatural movements when motion capture fails.
An information processing method utilizing a feature point generation model to generate and correct feature point information from time-series data, which is then transferred to a character object to generate a video that accurately reproduces the actor's movements.
The method ensures high accuracy in reproducing the actor's movements, even in live streaming scenarios, by correcting errors in feature point information, resulting in smooth and natural character object movements.
Smart Images

Figure 0007755858000001 
Figure 0007755858000002 
Figure 0007755858000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing method, an information processing system, and a program. [Background technology]
[0002] 2. Description of the Related Art In the entertainment industry, a technique has been used in which a character object (avatar) reproduces the movements of a person by transcribing the movements of the person captured by motion capture.
[0003] For example, Cited Document 1 discloses a video distribution system that distributes a video including a first character generated based on the movements of a distribution user, the video distribution system including one or more computer processors, the one or more computer processors executing computer-readable instructions to accept requests from multiple users to participate in the video, select multiple participating users from among the multiple users, select one or more guest users from among the multiple participating users, send a notification regarding the selection of the one or more guest users to each of the one or more guest users, and generate a co-starring video including guest characters corresponding to at least a portion of the one or more guest users and the first character. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2022-003776 Summary of the Invention [Problem to be solved by the invention]
[0005] However, in the prior art, if the motion capture fails, there is a problem in that the incorrect movements of the person are reproduced on the character object.
[0006] The present invention has been made in consideration of the above-mentioned problems, and has as its object to allow a character object to accurately reproduce the movements of an actor. [Means for solving the problem]
[0007] An information processing method according to one embodiment is an information processing method executed by an information processing system, and includes a feature point generation process that uses a feature point generation model to generate new feature point information from time-series data of the feature point information, a video generation process that transfers the generated feature point information to a character object and generates a video to be distributed that includes the transferred character object, and a distribution process that distributes the generated video to be distributed. [Effects of the Invention]
[0008] According to one embodiment, the movement of an actor can be reproduced by a character object with high accuracy. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a diagram illustrating an example of a configuration of an information processing system according to a first embodiment. [Figure 2] FIG. 2 illustrates an example of a hardware configuration of an information processing device. [Figure 3] FIG. 2 is a diagram illustrating an example of a functional configuration of a video distribution device. [Figure 4] FIG. 10 is a schematic diagram illustrating a learning method. [Figure 5] 10 is a flowchart illustrating an example of processing executed by the information processing system. [Figure 6] FIG. 2 is a schematic diagram illustrating a process executed by the information processing system. [Figure 7] FIG. 2 is a schematic diagram illustrating a process executed by the information processing system. [Figure 8] FIG. 2 is a schematic diagram illustrating a process executed by the information processing system. [Figure 9] FIG. 10 is a schematic diagram illustrating the effect. [Figure 10]FIG. 10 is a diagram illustrating an example of a configuration of an information processing system according to a second embodiment. [Figure 11] FIG. 2 is a diagram illustrating an example of a functional configuration of a video distribution device. [Figure 12] 10 is a flowchart illustrating an example of processing executed by the information processing system. [Figure 13] FIG. 2 is a schematic diagram illustrating a process executed by the information processing system. [Figure 14] FIG. 2 is a schematic diagram illustrating a process executed by the information processing system. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, each embodiment of the present invention will be described with reference to the accompanying drawings. Note that, in the description of the specification and drawings relating to each embodiment, components having substantially the same functional configuration are designated by the same reference numerals, and redundant description will be omitted.
[0011] [First embodiment] <System configuration> First, an overview of the information processing system according to this embodiment will be described. The information processing system according to this embodiment is a system for live streaming a streaming video M including a character object that reproduces the movements of an actor based on a video of the actor. The character object referred to here is a two-dimensional or three-dimensional model that imitates the actor. The actor may be a person or an animal such as a dog or cat. This information processing system can be used, for example, to operate an avatar in a virtual space or for video streaming by a VTuber.
[0012] Fig. 1 is a diagram showing an example of the configuration of an information processing system according to this embodiment. As shown in Fig. 1, the information processing system according to this embodiment includes a video distribution device 1, an actor terminal 2, and a viewer terminal 3, which are communicably connected to each other via a network N. The network N is, for example, a wired local area network (LAN), a wireless LAN, the Internet, a public line network, a mobile data communication network, or a combination of these. In the example of Fig. 1, the information processing system includes one video distribution device 1, one actor terminal 2, and one viewer terminal 3, but may include multiple of each.
[0013] The video distribution device 1 is an information processing device that distributes a distribution video M including a character object that reproduces the movements of an actor to a viewer terminal 3. The video distribution device 1 receives an original video m in which an actor is shot as a subject from an actor terminal 2 in real time, generates a distribution video M including a character object that reproduces the movements of the actor based on the received original video m, and live-distributes the generated distribution video M to the viewer terminal 3. The video distribution device 1 may be any information processing device that is capable of generating and live-distributing a distribution video M. The video distribution device 1 is, for example, but is not limited to, a PC (Personal Computer), a smartphone, a tablet terminal, a server device, or a microcomputer.
[0014] The actor terminal 2 is an information processing device used by an actor of the video to be distributed M. The actor is the person whose movements are reproduced by a character object in the video to be distributed M, in other words, the person who operates the character object. The actor terminal 2 shoots an original video m in which the actor is the subject, and transmits the obtained original video m to the video distribution device 1 in real time. The actor terminal 2 may be any information processing device that can shoot the original video m and transmit it to the video distribution device 1 in real time. The actor terminal 2 may also be an information processing device that is connected to a camera that shoots the original video m, acquires the original video m from the camera, and transmits the acquired original video m to the video distribution device 1 in real time. The actor terminal 2 is, for example, but is not limited to, a PC, smartphone, or tablet terminal.
[0015] The viewer terminal 3 is an information processing device used by a viewer of the distributed video M. A viewer is a person who views the distributed video M, which includes character objects that reproduce the movements of actors. The viewer terminal 3 receives the distributed video M from the video distribution device 1 in real time and displays it on a display. The viewer terminal 3 can be any information processing device that can receive and display the distributed video M. The viewer terminal 3 is, for example, a PC, a smartphone, or a tablet terminal, but is not limited to these.
[0016] <Hardware configuration> Next, a description will be given of the hardware configuration of the information processing device 100. Fig. 2 is a diagram showing an example of the hardware configuration of the information processing device 100. As shown in Fig. 2, the information processing device 100 includes a processor 101, a memory 102, a storage 103, a communication I / F 104, an input / output I / F 105, and a drive device 106, which are interconnected via a bus B.
[0017] The processor 101 controls each component of the information processing device 100 and realizes the functions of the information processing device 100 by loading various programs, including an OS (Operating System), stored in the storage 103, into the memory 102 and executing the programs. The processor 101 is, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), or a combination thereof.
[0018] The memory 102 is, for example, a read-only memory (ROM), a random access memory (RAM), or a combination thereof. The ROM is, for example, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a combination thereof. The RAM is, for example, a dynamic random access memory (DRAM), a static random access memory (SRAM), or a combination thereof.
[0019] The storage 103 stores various programs including an OS and data. The storage 103 is, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), a storage class memory (SCM), or a combination of these.
[0020] The communication I / F 104 is an interface for connecting the information processing device 100 to an external device via the network N and controlling communication. The communication I / F 104 is, for example, an adapter conforming to Bluetooth (registered trademark), Wi-Fi (registered trademark), ZigBee (registered trademark), Ethernet (registered trademark), or optical communication, but is not limited to these.
[0021] The input / output I / F 105 is an interface for connecting an input device 107 and an output device 108 to the information processing device 100. The input device 107 is, for example, a mouse, a keyboard, a touch panel, a microphone, a scanner, a camera, various sensors, an operation button, or a combination thereof. The output device 108 is, for example, a display, a projector, a printer, a speaker, a vibrator, or a combination thereof.
[0022] The drive device 106 reads and writes data from and to the disk media 109. The drive device 106 is, for example, a magnetic disk drive, an optical disk drive, a magneto-optical disk drive, or a combination thereof. The disk media 109 is, for example, a compact disc (CD), a digital versatile disc (DVD), a floppy disk (FD), a magneto-optical disk (MO), a Blu-ray (registered trademark) disc (BD), or a combination thereof.
[0023] In this embodiment, the program may be written to memory 102 or storage 103 during the manufacturing stage of the information processing device 100, or may be provided to the information processing device 100 via network N, or may be provided to the information processing device 100 via a non-transitory computer-readable recording medium such as disk media 109.
[0024] <Functional configuration> Next, we will explain the functional configuration of video distribution device 1. Fig. 3 is a diagram showing an example of the functional configuration of video distribution device 1. As shown in Fig. 3, video distribution device 1 includes a communication unit 11, a storage unit 12, and a control unit 13.
[0025] The communication unit 11 is realized by the communication I / F 104. The communication unit 11 transmits and receives information between the actor terminal 2 and the viewer terminal 3 via the network N. The communication unit 11 receives the original video m from the actor terminal 2. The communication unit 11 also transmits (distributes) the distribution video M to the viewer terminal 3.
[0026] The storage unit 12 is realized by a memory 102 and a storage 103. The storage unit 12 stores an original video m, a feature point estimation model 121, feature point information 122, a feature point generation model 123, video information 124, and a distributed video M.
[0027] The original video m is a video shot with an actor as the subject. The original video m may be shot by the actor himself or by a different cameraman. The video distribution device 1 receives the original video m transmitted from the actor terminal 2 in real time and stores it sequentially in the storage unit 12.
[0028] The feature point estimation model 121 is a trained machine learning model that estimates feature point information of an actor as a subject from an image in which the actor is photographed as a subject. The feature point estimation model 121 can be any machine learning model that has been trained to output feature point information of the actor as a subject when an image in which the actor is photographed as a subject is input. The feature point estimation model 121 is, for example, a deep-learned convolutional neural network (CNN), but is not limited to this. The feature point estimation model 121 may be a skeleton detection model or a facial keypoint detection model.
[0029] The feature point information 122 is information about a plurality of feature points of an actor, and includes information indicating the coordinates of each feature point and the relationship between the feature points. The feature points of an actor are specific parts of the actor's body, and are set according to the actor and the feature point estimation model 121. The feature points of an actor may be, for example, joints, eyes, ears, nose, or mouth, but are not limited to these. The feature points may be bones detected by skeleton detection, or facial key points detected by facial key point detection. The coordinates of the feature points may be two-dimensional coordinates or three-dimensional coordinates. The feature point information 122 includes actor feature point information 122A estimated from the original video m using the feature point estimation model 121, and feature point information 122B generated from the feature point information 122A using the feature point generation model 123.
[0030] The feature point generation model 123 is a trained machine learning model that generates new feature point information 122B from the time-series data of the feature point information 122A. When time-series data of feature point information for a certain period is input, the feature point generation model 123 is trained to generate new feature point information corresponding to any of the feature point information for that period. Furthermore, when time-series data of feature point information for a certain period that includes erroneous feature point information is input, the feature point generation model 123 is trained to generate feature point information in which the erroneous feature point information has been corrected. The feature point generation model 123 is, for example, a deep-trained recurrent neural network (RNN), but is not limited to this. A training method for the feature point generation model 123 will be described in detail below.
[0031] The video information 124 is any information used to generate the streaming video M. The video information 124 includes information about character objects that reproduce the movements of actors, information about the sound of the streaming video M, information about the virtual space that constitutes the streaming video M, and information about the viewpoint of the streaming video M. The information about character objects includes information indicating the shape, structure, size, and color of the character objects, feature point information set for the character objects, and information indicating the position of the character objects in the virtual space. The information about sounds includes information about background music (BGM) and sound effects, and information about the voice of the actors. The information about the virtual space includes information indicating background character objects that constitute the virtual space, and information indicating the position, brightness, direction, and color of a light source in the virtual space. The information about the viewpoint includes information indicating the position, angle of view, zoom, and direction of the viewpoint in the virtual space. Note that the video information 124 is not limited to the above examples. The video information 124 may not include some of the above information, or may include information other than the above.
[0032] The distribution moving image M is a moving image that is generated by the moving image distribution device 1 and distributed to viewers. The moving image distribution device 1 stores the distribution moving image M that is generated in real time in the storage unit 12 sequentially.
[0033] Control unit 13 is realized by processor 101 reading and executing a program from memory 102 and working in cooperation with other hardware components. Control unit 13 controls the overall operation of video distribution device 1. Control unit 13 includes an acquisition unit 131, an estimation unit 132, a feature point generation unit 133, a video generation unit 134, a distribution unit 135, and a learning unit 136.
[0034] The acquisition unit 131 acquires the original video m received by the video distribution device 1 from the actor terminal 2 and stores it in the storage unit 12.
[0035] The estimation unit 132 uses the feature point estimation model 121 to estimate actor feature point information 122A from each frame of the original video m stored in the storage unit 12, and stores the information in the storage unit 12. As a result, the storage unit 12 stores time-series data of the actor feature point information 122A.
[0036] The feature point generation unit 133 uses the feature point generation model 123 to generate new feature point information 122B from the time-series data of the actor's feature point information 122A stored in the storage unit 12, and stores the information in the storage unit 12. The process of generating the feature point information 122B will be described in detail later.
[0037] The moving image generation unit 134 refers to the moving image information 124, transfers the new feature point information 122B generated by the feature point generation unit 133 to the character object, synthesizes the transferred character object, background, audio, BGM and sound effects, generates a distribution moving image M including the transferred character object, and stores the generated distribution moving image M in the storage unit 12. Transferring the new feature point information 122B to the character object means deforming the posture of the character object so that the feature point information of the character object corresponds to the new feature point information 122B.
[0038] The distribution unit 135 distributes (transmits) the distribution moving image M stored in the storage unit 12 to the viewer terminal 3.
[0039] The learning unit 136 trains the feature point generation model 123. In other words, the learning unit 136 generates the feature point generation model 123 by machine learning a neural network so that, when time-series data of feature point information including erroneous feature point information is input, the learning unit 136 outputs feature point information in which the erroneous feature point information has been corrected.
[0040] Here, a detailed description will be given of a learning method for the feature point generation model 123. FIG.
[0041] First, video data D1 with correct feature point information is prepared in order to train the feature point generation model 123. Data of a series of movements of an actor or character object with correct feature point information, which is distributed for a fee or free of charge, can be used as the video data D1. The series of movements can be, for example, walking, gymnastics, a specific action, a sports form, or dancing, but is not limited to these.
[0042] 4 is video data of a person (actor) raising and lowering his / her hand. This video data D1 includes frames f1 to f9 capturing a series of movements, and each frame is assigned with correct feature point information.
[0043] Next, feature point information is extracted from each frame of the video data D1, thereby extracting time-series data D2 of correct feature point information contained in each frame of the video data D1.
[0044] The time series data D2 in Fig. 4 is time series data extracted from the video data D1 in Fig. 4. The feature point information of each frame f1 to f9 of the time series data D2 corresponds to the feature point information of each frame f1 to f9 of the video data D1.
[0045] Next, part of the feature point information of the time-series data D2 is changed. The changed feature point information is different from the correct feature point information, i.e., it becomes erroneous feature point information. As a result, time-series data D3 of feature point information including the erroneous feature point information is generated.
[0046] The time series data D3 in Fig. 4 is obtained by changing a part of frame f8 of the time series data D2 in Fig. 4. This time series data D3 is composed of frames f1 to f7 and f9 that contain correct feature point information, and frame f8 that contains incorrect feature point information.
[0047] The learning unit 136 trains the feature point generation model 123 using the time-series data D2 and D3 thus prepared as training data. Specifically, the learning unit 136 trains the feature point generation model 123 so that, when time-series data of feature point information for a certain period is input, the model 123 outputs correct feature point information corresponding to any of the feature point information for that certain period. The length of the certain period can be set arbitrarily. The certain period may be represented by a time period in a video or by the number of frames. Furthermore, the output feature point information may correspond to the latest feature point information for the certain period, or may correspond to feature point information prior to the latest feature point information.
[0048] In the example of FIG. 4, one period is 3 frames. This corresponds to 0.1 seconds when the video data D1 is at 30 fps. Also, in the example of FIG. 4, the system is trained to output correct feature point information corresponding to the latest feature point information in one period. For example, when feature point information of frames f1 to f3 of the time-series data D2 is input, the system is trained to output correct feature point information of frame f3, which is the latest frame. Similarly, when feature point information of frames f6 to f8 of the time-series data D2 is input, the system is trained to output correct feature point information of frame f8, which is the latest frame.
[0049] Specifically, the learning unit 136 uses the time series data D2 as correct answer data and adjusts the parameters of the feature point generation model 123 so that the feature point information output when the feature point information of frames f(x) to f(x+2) of the time series data D3 is input approaches the feature point information of frame f(x+2) of the time series data D2.
[0050] In general, the learning unit 136 uses the time-series data D2 as correct answer data and adjusts the parameters of the feature point generation model 123 so that the feature point information output when feature point information of frames f(x) to f(x+y) of the time-series data D3 is input approaches the feature point information of frame f(x+z) of the time-series data D2. y represents the period (1≦y), and z represents the position (0≦z≦y) of the feature point information to which the output feature point information corresponds. The case where y is 2 and z is 2 corresponds to FIG. 4.
[0051] When the feature point generation model 123 is trained in this manner, the feature point generation model 123 becomes a machine learning model that, when time-series data of feature point information 122A for a certain period is input, generates new feature point information 122B corresponding to any of the feature point information for that period. The generated new feature point information 122B is feature point information that has been corrected to approach correct feature point information. In other words, when erroneous feature point information is included in the input time-series data of feature point information, the feature point generation model 123 becomes a machine learning model that outputs feature point information that has been corrected so that the erroneous feature point information approaches the correct feature point information. As a result, the time-series data of feature point information 122B generated by the feature point generation model 123 becomes time-series data of feature point information in which the erroneous feature point information has been corrected so that it approaches the correct feature point information and has been smoothed as a whole.
[0052] The functional configuration of video distribution device 1 is not limited to the above example. For example, video distribution device 1 may have some of the above functional configuration, with the actor terminal 2 or the viewer terminal 3 having the rest. For example, actor terminal 2 may have estimation unit 132, and actor terminal 2 may transmit feature point information 122A to video distribution device 1 instead of original video m. Video distribution device 1 may also have functional configurations other than those described above. Each functional configuration of video distribution device 1 may be realized by software, as described above, or by hardware such as an IC chip, a System on Chip (SoC), a Large Scale Integration (LSI), or a microcomputer.
[0053] <Processing performed by the information processing system> Next, a description will be given of processing executed by the information processing system according to this embodiment. Fig. 5 is a flowchart showing an example of processing executed by the information processing system. Figs. 6 to 8 are schematic diagrams illustrating processing executed by the information processing system.
[0054] (Step S101) When the actor starts distributing the distribution video M, the actor terminal 2 starts shooting the original video m with a camera (step S101). The actor terminal 2 may start shooting the original video m before starting distribution of the distribution video M.
[0055] (Step S102) When the actor terminal 2 starts shooting the original video m, it sequentially transmits the original video m acquired from the camera to the video distribution device 1 (step S102). The actor terminal 2 continues transmitting the original video m until the distribution of the distribution video M is completed (step S103: NO). When the actor terminal 2 completes the distribution of the distribution video M, it stops shooting and transmitting the original video m (step S103: YES).
[0056] (Step S104) The acquisition unit 131 of the moving image distribution device 1 sequentially acquires the original moving images m transmitted by the actor terminals 2 and stores them in the storage unit 12 (step S104).
[0057] 6 is video data of a person (actor) raising and lowering their hand. This original video m includes frames f1 to f9 that capture a series of actions. The acquisition unit 131 acquires frames f1 to f9 in this order and stores them in the storage unit 12.
[0058] (Step S105) The estimation unit 132 uses the feature point estimation model 121 to estimate actor feature point information 122A from each frame of the original moving image m stored in the storage unit 12 (step S105) and stores it in the storage unit 12. More specifically, the estimation unit 132 inputs the frames of the original moving image m stored in the storage unit 12 to the feature point estimation model 121, and stores the feature point information output by the feature point estimation model 121 as actor feature point information 122A estimated from the frames in the storage unit 12. By performing this for each frame, time-series data d1 of the actor feature point information 122A is stored in the storage unit 12.
[0059] The time-series data d1 in Fig. 6 includes actor feature point information 122A estimated from frames f1 to f9 of the original video m. In the example of Fig. 6, the feature point information 122A for frame f8 is incorrect (motion capture failed). Such incorrect feature point information 122A may be generated if the accuracy of the feature point estimation model 121 is not high enough, if the original video m is blurred, or if part of the actor's body is in the shadow of something.
[0060] (Step S106) The feature point generation unit 133 uses the feature point generation model 123 to generate new feature point information 122B from the time-series data d1 of the feature point information 122A stored in the storage unit 12 (step S106), and stores the generated feature point information 122B in the storage unit 12. More specifically, the feature point information generation unit 134 inputs the time-series data d1 for one period stored in the storage unit 12 to the feature point generation model 123, and stores the feature point information output by the feature point generation model 123 in the storage unit 12 as new feature point information corresponding to any of the feature point information 122B for the one period. By performing this process while shifting the periods, time-series data d2 of the generated actor's feature point information 122B is stored in the storage unit 12.
[0061] In the example of FIG. 7, one period consists of three frames, and new feature point information 122B corresponding to the latest feature point information 122A in one period is generated. As a result, new feature point information 122B corresponding to each of the feature point information 122A of frames f3 to f9 of the time-series data d1 is generated. As shown in FIG. 7, for the erroneous feature point information 122A of frame f8 of the time-series data d1, new feature point information 122B corrected so as to approximate the correct feature point information is generated. Note that new feature point information 122B corresponding to the feature point information 122A of frame f1 may be generated by inputting blank feature point information as dummy feature point information from before the frame f1 along with the feature point information 122A of frame f1 of the time-series data d1. The same applies to the feature point information 122A of frame f2 of the time-series data d1.
[0062] (Step S107) The video generation unit 134 refers to the video information 124, transfers the new feature point information 122B generated by the feature point generation unit 133 to the character object, synthesizes the transferred character object, background, audio, BGM and sound effects, generates a distribution video M including the transferred character object (step S107), and stores the generated distribution video M in the memory unit 12.
[0063] 8 is obtained by transferring feature point information 122B of frames f3 to f9 of time-series data d2 to character objects and combining the background, etc. The character objects in frames f3 to f9 of the distributed video M reproduce the movements of the actors in frames f3 to f9 of the original video m.
[0064] (Step S108) The distribution unit 135 live-distributes the distribution moving image M stored in the storage unit 12 to the viewer terminal 3 (step S108). That is, the distribution unit 135 sequentially transmits the generated distribution moving image M to the viewer terminal 3 every time the distribution moving image M is generated.
[0065] (Step S109) When the viewer terminal 3 receives the distribution video M distributed from the video distribution device 1, it plays the received distribution video M on the display (step S109). This allows the viewer to view the distribution video M on the viewer terminal 3.
[0066] <Summary> As described above, according to this embodiment, feature point information 122A is estimated based on the original moving image m of the actor, new feature point information 122B (time-series data d2) is generated from the time-series data d1 of feature point information 122A using feature point generation model 123, the generated feature point information 122B is transferred to a character object, a streaming moving image M including the transferred character object is generated, and the generated streaming moving image M can be distributed. This makes it possible to live-stream a streaming moving image M including a character object that reproduces the movements of the actor.
[0067] Here, Fig. 9 is a diagram for explaining the effect of this embodiment. The conventional moving image in Fig. 9 corresponds to a distributed moving image generated by a conventional method.
[0068] In conventional technology, feature point information 122A estimated from the original video m is directly transferred to the character object. Therefore, if motion capture fails and incorrect feature point information 122A is generated, the character object will move unnaturally, with no coherence to the previous or next frame, as in frame f8 of the conventional video. In particular, when streaming a video live, it is difficult to prevent the unnatural movement of the character object caused by a motion capture failure, since the incorrect feature point information 122A cannot be manually corrected. The more unnatural the character object's movement, the more uncomfortable it feels to the viewer, which leads to a decrease in satisfaction with the video.
[0069] In contrast, in this embodiment, feature point information 122A estimated from the original video m is corrected by feature point generation model 123, and new feature point information 122B is transferred to the character object. Therefore, even if motion capture fails and incorrect feature point information 122A is generated, the character object will move in a manner similar to the correct feature point information (i.e., similar to the actor's movement), as in frame f8 of the distribution video M. In other words, the actor's movement can be accurately reproduced by the character object. As a result, even when the distribution video M is live-streamed, it is possible to stream a distribution video M in which the unnatural movement of the character object is suppressed and the character object's movement is smooth. This can improve viewer satisfaction with the distribution video M.
[0070] In this embodiment, the distribution video M does not have to be live. When the distribution video M is distributed on demand, the distribution video M can be distributed with smooth character object movements without manually correcting the feature point information 122A.
[0071] [Second embodiment] <System configuration> The information processing system according to this embodiment is a system for live streaming a streaming video M including a character object that reproduces the movements of an actor based on sensor data from a sensor worn by the actor. Differences from the first embodiment will be described below.
[0072] FIG. 10 is a diagram showing an example of the configuration of an information processing system according to this embodiment. As shown in FIG. 10, the information processing system according to this embodiment includes a video distribution device 1, an actor terminal 2, a viewer terminal 3, and a sensor s, all of which are communicatively connected via a network N. The network N may be, for example, a wired LAN, a wireless LAN, the Internet, a public line network, a mobile data communication network, or a combination thereof. In the example of FIG. 10, the information processing system includes one video distribution device 1, one actor terminal 2, and one viewer terminal 3, but may also include multiple of each. Note that the viewer terminal 3 is the same as in the first embodiment, and therefore description thereof will be omitted.
[0073] The sensors s are sensors for measuring the actor's movements and are attached to multiple positions on the actor's body. The positions at which the sensors s are attached include, but are not limited to, the ankles, knees, hips, shoulders, head, elbows, wrists, and fingers. In the example of FIG. 10, five sensors s are attached to the actor, but the number of sensors s attached to the actor is arbitrary. Each of the sensors s attached to the actor wirelessly transmits sensor data sd to the actor terminal 2 in real time. The sensors s are, for example, acceleration sensors, but are not limited to these. The sensors s can be any sensors capable of measuring the actor's movements. Multiple types of sensors may be used in combination as the sensors s.
[0074] The video distribution device 1 is an information processing device that distributes a distribution video M including a character object that reproduces the movements of an actor to a viewer terminal 3. The video distribution device 1 receives sensor data sd of a sensor s worn by an actor from an actor terminal 2 in real time, generates a distribution video M including a character object that reproduces the movements of the actor based on the received sensor data sd, and live-distributes the generated distribution video M to the viewer terminal 3. The video distribution device 1 may be any information processing device that is capable of generating and live-distributing a distribution video M. The video distribution device 1 is, for example, but is not limited to, a PC, a smartphone, a tablet terminal, a server device, or a microcomputer.
[0075] The actor terminal 2 is an information processing device used by an actor in the streaming video M. The actor is the person whose movements are reproduced by a character object in the streaming video M, in other words, the person who operates the character object. The actor terminal 2 wirelessly receives sensor data sd from multiple sensors s worn by the actor in real time and transmits the received sensor data sd to the video streaming device 1 in real time. The actor terminal 2 can be any information processing device that can transfer the sensor data sd to the video streaming device 1 in real time. The actor terminal 2 is, for example, but is not limited to, a PC, smartphone, or tablet terminal.
[0076] <Functional configuration> Next, we will explain the functional configuration of video distribution device 1. Fig. 11 is a diagram showing an example of the functional configuration of video distribution device 1. As shown in Fig. 11, video distribution device 1 includes a communication unit 11, a storage unit 12, and a control unit 13.
[0077] The communication unit 11 is realized by the communication I / F 104. The communication unit 11 transmits and receives information between the actor terminal 2 and the viewer terminal 3 via the network N. The communication unit 11 receives sensor data sd from the actor terminal 2. The communication unit 11 also transmits (distributes) the distribution video M to the viewer terminal 3.
[0078] The storage unit 12 is realized by a memory 102 and a storage 103. The storage unit 12 stores sensor data sd, feature point information 122, a feature point generation model 123, video information 124, and a distributed video M. Since everything except the sensor data sd is the same as in the first embodiment, a description thereof will be omitted.
[0079] The sensor data sd is data measured by a sensor s attached to the actor's body. The sensor data sd is, for example, acceleration data, but is not limited to this. The sensor data sd can be any data capable of measuring the actor's movement. The video distribution device 1 receives the sensor data sd of the multiple sensors s transmitted from the actor terminal 2 in real time, and stores the data sequentially in the memory unit 12.
[0080] The control unit 13 is realized by the processor 101 reading and executing a program from the memory 102 and working in cooperation with other hardware components. The control unit 13 controls the overall operation of the video distribution device 1. The control unit 13 includes an acquisition unit 131, an estimation unit 132, a feature point generation unit 133, a video generation unit 134, a distribution unit 135, and a learning unit 136. Components other than the acquisition unit 131 and the estimation unit 132 are the same as those in the first embodiment, so a description thereof will be omitted.
[0081] The acquisition unit 131 acquires the sensor data sd of the multiple sensors s that the video distribution device 1 has received from the actor terminal 2, and stores the data in the storage unit 12.
[0082] The estimation unit 132 estimates the actor's feature point information 122A from the sensor data sd of the multiple sensors s stored in the storage unit 12, and stores the information in the storage unit 12. As a result, time-series data of the actor's feature point information 122A is stored in the storage unit 12. For example, when the sensor data sd is acceleration data, the estimation unit 132 calculates the movement distance of each sensor s from the integrated value of the acceleration of each sensor s, calculates the position of each sensor s from the integrated value of the movement distance of each sensor s, and estimates the feature point information 122A from the position of each sensor s.
[0083] <Processing performed by the information processing system> Next, the processing executed by the information processing system according to this embodiment will be described. Fig. 12 is a flowchart showing an example of the processing executed by the information processing system. Steps S207 to S209 are the same as steps S107 to S109 in the first embodiment, and therefore their description will be omitted. Figs. 13 and 14 are schematic diagrams illustrating the processing executed by the information processing system.
[0084] (Step S201) When an actor starts distributing a video to be distributed M, the actor terminal 2 starts acquiring sensor data sd from multiple sensors s (step S201). The actor terminal 2 may start acquiring sensor data sd before starting distribution of the video to be distributed M.
[0085] (Step S202) When the actor terminal 2 starts acquiring the sensor data sd, it sequentially transmits the sensor data sd acquired from the sensor s to the video distribution device 1 (step S202). The actor terminal 2 continues transmitting the sensor data sd until the distribution of the video M to be distributed is completed (step S203: NO). When the actor terminal 2 completes the distribution of the video M to be distributed, it terminates the acquisition and transmission of the sensor data sd (step S203: YES).
[0086] (Step S204) The acquisition unit 131 of the video distribution device 1 sequentially acquires the sensor data sd of the multiple sensors s transmitted by the actor terminal 2, and stores the data in the storage unit 12 (step S204).
[0087] The sensor data sd in FIG. 13 is sensor data of a person (actor) raising and lowering their hand. This sensor data sd includes sensor data sd acquired at times t1 to t9 during a series of actions. The sensor data sd at each time t includes sensor data from multiple sensors s1, s2, .... The acquisition unit 131 acquires the sensor data sd at times t1 to t9 in this order and stores it in the storage unit 12. The interval at which the sensor data sd is acquired can be set arbitrarily. For example, if the sensor data sd is acquired 30 times per second, 30 pieces of feature point information 122A can be generated per second based on the acquired sensor data sd, and therefore a 30 fps distribution video M can be generated.
[0088] (Step S205) The estimation unit 132 estimates actor feature point information 122A from the sensor data sd at each time t stored in the storage unit 12 (step S205) and stores the information in the storage unit 12. More specifically, the estimation unit 132 stores the feature point information estimated from the sensor data sd of the multiple sensors s at time t stored in the storage unit 12 as actor feature point information 122A at that time t in the storage unit 12. By performing this on the sensor data sd at each time, time-series data d1 of the actor feature point information 122A is stored in the storage unit 12.
[0089] The time-series data d1 in Fig. 13 includes actor feature point information 122A estimated from the sensor data sd at times t1 to t9. In the example of Fig. 13, the feature point information 122A at time t8 is incorrect (motion capture failed). Such incorrect feature point information 122A may be generated if the measurement accuracy of the sensor s is not sufficiently high or if an error occurs when receiving the sensor data sd.
[0090] (Step S206) The feature point generation unit 133 uses the feature point generation model 123 to generate new feature point information 122B from the time-series data d1 of the feature point information 122A stored in the storage unit 12 (step S206), and stores the generated feature point information 122B in the storage unit 12. More specifically, the feature point information generation unit 134 inputs the time-series data d1 for one period stored in the storage unit 12 to the feature point generation model 123, and stores the feature point information output by the feature point generation model 123 in the storage unit 12 as new feature point information corresponding to any of the feature point information 122B for the one period. By performing this process while shifting the periods, time-series data d2 of the generated actor's feature point information 122B is stored in the storage unit 12.
[0091] In the example of FIG. 14 , the acquisition interval of the sensor data sd matches the fps of the streaming video M, one period is three frames, and new feature point information 122B corresponding to the latest feature point information 122A in one period is generated. As a result, new feature point information 122B corresponding to each of the feature point information 122A at times t3 to t9 in the time series data d1 is generated. As shown in FIG. 14 , for the erroneous feature point information 122A at time t8 in the time series data d1, new feature point information 122B corrected to approximate the correct feature point information is generated. Note that new feature point information 122B corresponding to the feature point information 122A at time t1 in the time series data d1 may be generated by inputting blank feature point information as dummy feature point information from before that time together with the feature point information 122A at time t1 in the time series data d1. The same applies to the feature point information 122A at time t2 in the time series data d1.
[0092] Thereafter, steps S207 to S209 are executed. As a result, the viewer can view the distribution video M on the viewer terminal 3.
[0093] <Summary> As described above, according to this embodiment, feature point information 122A is estimated based on sensor data sd from multiple sensors s worn by the actor, new feature point information 122B (time-series data d2) is generated from the time-series data d1 of feature point information 122A using feature point generation model 123, the generated feature point information 122B is transferred to a character object, a streaming video M including the transferred character object is generated, and the generated streaming video M can be distributed. This makes it possible to live-stream a streaming video M including a character object that reproduces the actor's movements.
[0094] Also, in this embodiment, similar to the first embodiment, feature point information 122A estimated from sensor data sd is corrected by feature point generation model 123, and new feature point information 122B is transferred to the character object. Therefore, even if motion capture fails and incorrect feature point information 122A is generated, the character object will move in a manner similar to the correct feature point information (i.e., similar to the actor's movement), as in frame f8 of the distribution video M. In other words, the actor's movement can be accurately reproduced by the character object. As a result, even when the distribution video M is live-streamed, it is possible to stream a distribution video M in which the unnatural movement of the character object is suppressed and the character object's movement is smooth. This can improve viewer satisfaction with the distribution video M.
[0095] In this embodiment, the distribution video M does not have to be live. When the distribution video M is distributed on demand, the distribution video M can be distributed with smooth character object movements without manually correcting the feature point information 122A.
[0096] <Additional Notes> The present embodiment includes the following disclosure.
[0097] (Appendix 1) An information processing method executed by an information processing system, a feature point generation process for generating new feature point information from time-series data of the actor's feature point information using a feature point generation model; a video generation process of transferring the generated feature point information to a character object and generating a video to be distributed including the transferred character object; a distribution process for distributing the generated distribution video; An information processing method including:
[0098] (Appendix 2) An acquisition process for acquiring an original video in which the actor is shot as a subject; an estimation process of estimating feature point information of the actor from the original video and generating time-series data of the feature point information; 2. The information processing method of claim 1, further comprising:
[0099] (Appendix 3) an acquisition process for acquiring sensor data from a sensor worn by the actor; an estimation process of estimating feature point information of the actor from the sensor data and generating time-series data of the feature point information; 2. The information processing method of claim 1, further comprising:
[0100] (Appendix 4) The method further includes a learning process for training the neural network to output new feature point information when time-series data of feature point information is input. 1. The information processing method described in Appendix 1.
[0101] (Appendix 5) The learning process involves machine learning the neural network so that, when time-series data of feature point information for a certain period including erroneous feature point information is input, the neural network outputs feature point information in which the erroneous feature point information is corrected. 1. The information processing method described in Appendix 4.
[0102] (Appendix 6) The feature point generation process generates new feature point information corresponding to any of the feature point information in a certain period from time series data of the feature point information in the certain period. 1. The information processing method described in Appendix 1.
[0103] (Appendix 7) The feature point generation process generates new feature point information corresponding to the latest feature point information in a certain period from time series data of feature point information in the certain period. 1. The information processing method described in Appendix 1.
[0104] (Appendix 8) a feature point generation unit that generates new feature point information from time-series data of feature point information using a feature point generation model; a moving image generating unit that transfers the generated feature point information to a character object and generates a moving image to be distributed that includes the transferred character object; a distribution unit that distributes the generated distribution video; An information processing system comprising:
[0105] (Appendix 9) On the computer, a feature point generation process for generating new feature point information from time-series data of feature point information using a feature point generation model; a video generation process of transferring the generated feature point information to a character object and generating a video to be distributed including the transferred character object; a distribution process for distributing the generated distribution video; A program for executing an information processing method including the steps of:
[0106] The embodiments disclosed herein are illustrative in all respects and should not be considered limiting. The scope of the present invention is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims. Furthermore, the present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Explanation of symbols]
[0107] 1: Video distribution device 2: Actor terminal 3: Viewer terminal 11: Communications Department 12: Storage part 13: Control unit 100: Information processing device 121: Feature point estimation model 122: Feature point information 123: Feature point generation model 124:Video information 131: Acquisition Department 132: Estimation part 133: Feature point generation unit 134: Video generation unit 135: Distribution Department 136: Learning Department m:Original video M: Streaming video s: sensor sd: sensor data
Claims
1. An information processing method executed by an information processing system, a feature point generation process for generating new feature point information from time-series data of the actor's feature point information using a feature point generation model; a video generation process of transferring the generated feature point information to a character object and generating a video to be distributed including the transferred character object; a distribution process for distributing the generated distribution video; An information processing method including:
2. An acquisition process for acquiring an original video in which the actor is shot as a subject; an estimation process of estimating feature point information of the actor from the original video and generating time-series data of the feature point information; The information processing method according to claim 1 , further comprising:
3. an acquisition process for acquiring sensor data from a sensor worn by the actor; an estimation process of estimating feature point information of the actor from the sensor data and generating time-series data of the feature point information; The information processing method according to claim 1 , further comprising:
4. The method further includes a learning process for training the neural network to output new feature point information when time-series data of feature point information is input. The information processing method according to claim 1 .
5. The learning process involves machine learning the neural network so that, when time-series data of feature point information for a certain period including erroneous feature point information is input, the neural network outputs feature point information in which the erroneous feature point information is corrected. The information processing method according to claim 4.
6. The feature point generation process generates new feature point information corresponding to any of the feature point information in a certain period from time series data of the feature point information in the certain period. The information processing method according to claim 1 .
7. The feature point generation process generates new feature point information corresponding to the latest feature point information in a certain period from time series data of feature point information in the certain period. The information processing method according to claim 1 .
8. a feature point generation unit that generates new feature point information from time-series data of the actor's feature point information using a feature point generation model; a moving image generating unit that transfers the generated feature point information to a character object and generates a moving image to be distributed that includes the transferred character object; a distribution unit that distributes the generated distribution video; An information processing system comprising:
9. On the computer, a feature point generation process for generating new feature point information from time-series data of the actor's feature point information using a feature point generation model; a video generation process of transferring the generated feature point information to a character object and generating a video to be distributed including the transferred character object; a distribution process for distributing the generated distribution video; A program for executing an information processing method including the steps of:
Citation Information
Patent Citations
Video distribution system, video distribution method, and video distribution program live-distributing video including animation of character object generated based on movement of distribution user
JP2022003776A
Method for camera tracking for providing mixed rendering content using virtual reality and augmented reality and system using the same
KR1020210012442A
Systems and methods for applying animations or motions to a character
US20140320508A1
Systems and Methods for Motion-Controlled Animation
US20220101587A1
Generating animated digital videos utilizing a character animation neural network informed by pose and motion embeddings
US20230123820A1