Information processing device, terminal device, information processing method, and program
The system enhances camera parameter estimation using motion data and metadata to accurately synchronize live-action and reconstructed videos, addressing limitations in existing techniques by improving accuracy and speed.
Patent Information
- Application Number
- PCT/JP2025/006781
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2025-02-27
- Publication Date
- 2025-10-16
AI Technical Summary
Existing techniques for estimating camera parameters in live-action video are limited in accuracy and applicability, particularly in scenarios with camera switching, zooming, and occlusions, and are often restricted to specific contexts like sports games with visible field lines.
A system that estimates camera parameters using motion data and metadata of a moving object, employing a server to render a 3D model into a 2D image based on similarity with live-action footage, allowing synchronous playback of reconstructed and live-action videos.
Improves the accuracy and speed of camera parameter estimation, enabling seamless synchronization of live-action and reconstructed videos without incongruity, applicable to various environments beyond sports games.
Smart Images

Figure JP2025006781_16102025_PF_FP_ABST
Abstract
Description
Information processing device, terminal device, information processing method and program
[0001] The present disclosure relates to an information processing device, a terminal device, an information processing method, and a program.
[0002] Conventionally, techniques have been developed for estimating camera parameters of a camera that captured video. For example, Patent Document 1 listed below discloses a technique for estimating camera parameters based on the skeleton of a person appearing in a video.
[0003] International Publication No. 2022 / 259618
[0004] However, the technology disclosed in Patent Document 1 has only recently been developed, and there is still room for improvement in various respects, including improving the accuracy of estimating camera parameters.
[0005] Therefore, the present disclosure has been made in consideration of the above problems, and an object of the present disclosure is to provide a mechanism that can further improve the estimation accuracy of camera parameters.
[0006] According to the present disclosure, an information processing device is provided that includes a control unit that estimates camera parameters of live-action footage based on a first characteristic of a moving object present in a target space and a second characteristic of the moving object captured in the live-action footage of the target space.
[0007] Furthermore, according to the present disclosure, a terminal device is provided that includes a control unit that generates a playback screen in which a reproduced image is generated by rendering a 3D model generated based on a first feature of a moving object present in a target space into a 2D image based on camera parameters of live-action video taken of the target space, and plays the reproduced image in synchronization with the live-action video.
[0008] Furthermore, according to the present disclosure, an information processing method is provided that includes estimating camera parameters of live-action footage based on a first characteristic amount of a moving object present in a target space and a second characteristic amount of the moving object captured in the live-action footage of the target space.
[0009] Furthermore, according to the present disclosure, an information processing method is provided that includes generating a playback screen in which a reproduced image is generated by rendering a 3D model generated based on a first feature of a moving object present in a target space into a 2D image based on camera parameters of live-action video taken of the target space, and playing the reproduced image in synchronization with the live-action video.
[0010] Furthermore, according to the present disclosure, a program is provided for causing a computer to function as a control unit that estimates camera parameters of live-action footage based on a first characteristic of a moving object present in a target space and a second characteristic of the moving object captured in the live-action footage of the target space.
[0011] Furthermore, according to the present disclosure, a program is provided for causing a computer to function as a control unit that generates a playback screen in which a reproduced image is generated by rendering a 3D model generated based on a first feature of a moving object existing in a target space into a 2D image based on camera parameters of live-action footage of the target space, and the reproduced image is played back in synchronization with the live-action footage.
[0012] FIG. 1 is a diagram for explaining an overview of a system 1 according to an embodiment of the present disclosure. FIG. 2 is a block diagram showing a configuration related to a reproduced video of the system 1 according to a first embodiment of the present disclosure. FIG. 3 is a block diagram showing a configuration related to an actual video of the system 1 according to the same embodiment. FIG. 4 is a block diagram showing an example configuration of a server 2 according to the same embodiment. FIG. 5 is a diagram for explaining camera parameters in the same embodiment. FIG. 6 is a block diagram showing an example configuration of a user terminal 3 according to the same embodiment. FIG. 7 is a diagram showing an example of a playback screen displayed by the user terminal 3 according to the same embodiment. FIG. 8 is a diagram showing an example of a playback screen displayed by the user terminal 3 according to the same embodiment. FIG. 9 is a block diagram showing an example configuration of a camera parameter estimation unit 34 according to the same embodiment. FIG. 10 is a diagram showing a specific example of processing by the camera parameter estimation unit 34 according to the same embodiment. FIG. 11 is a diagram showing a specific example of processing by the camera parameter estimation unit 34 according to the same embodiment. FIG. 12 is a block diagram showing an example of a configuration of a playback speed estimation unit 35 according to the same embodiment. FIG. 13 is a diagram showing a specific example of processing by the playback speed estimation unit 35 according to the same embodiment. FIG. 14 is a block diagram showing an example of processing by the playback speed estimation unit 35 according to the same embodiment. FIG. 15 is a block diagram showing an example of the flow of camera parameter estimation processing executed by the server 2 according to the same embodiment. Fig. 10 is a block diagram showing an example of the flow of synchronized playback processing executed by a user terminal 3 according to the embodiment. Fig. 11 is a block diagram showing a configuration related to reproduced video of a system 1 according to a second embodiment of the present disclosure. Fig. 12 is a block diagram showing a configuration related to live-action video of a system 1 according to the embodiment. Fig. 13 is a block diagram showing an example of a hardware configuration of an information processing device according to each of the above embodiments.
[0013] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0014] The explanation will be given in the following order: 1. Technical features 2. First embodiment 2.1. Configuration example 2.2. Detailed configuration 2.3. Processing flow 3. Second embodiment 4. Hardware configuration example 5. Supplementary information
[0015] 1. Technical Features> Fig. 1 is a diagram illustrating an overview of a system 1 according to an embodiment of the present disclosure. As shown in Fig. 1, the system 1 generates a reproduced video based on motion data and metadata obtained from a moving object such as a human, and plays the reproduced video together with live-action video obtained by actually capturing the moving object.
[0016] The reconstructed video may be a 3D computer graphics (3DCG) animation of a 3D model of a moving object. In particular, the 2D (2-dimensional) images for each video frame constituting the reconstructed video are generated by rendering a 3D model based on motion data into a 2D image using the same camera parameters as those of the live-action video. This allows a user to view live-action video and reconstructed video that were shot or generated using the same camera parameters.
[0017] Furthermore, the system 1 synchronously plays back the reproduced video and the live-action video. Synchronous playback of the reproduced video and the live-action video means that the acquisition time of the motion data that is the basis of the reproduced video and the shooting time of the live-action video match or nearly match. This allows the user to enjoy the reproduced video and the live-action video without any sense of incongruity.
[0018] Various prior art techniques have been proposed for estimating camera parameters of live-action video by image analysis, including the technique disclosed in Patent Document 1. One example is a technique for estimating camera parameters based on line recognition and vanishing point recognition in live-action video. Another related technique is a technique for estimating 3D structure from semantic segmentation of 2D images.
[0019] However, each of the prior arts has various drawbacks. For example, live-action footage from which camera parameters can be accurately estimated is sometimes limited to sports games, such as soccer, in which field lines may appear, shot at an appropriate angle of view that includes the lines. For another example, it is difficult to accurately estimate camera parameters for live-action footage that includes camera switching, zooming, and overlapping people (i.e., occlusion).
[0020] Taking these circumstances into consideration, a system 1 according to an embodiment of the present disclosure has been invented. The system 1 estimates camera parameters of live-action video using motion data and metadata of a moving object. This enables the system 1 to further improve the accuracy of camera parameter estimation while eliminating the disadvantages imposed by prior art techniques that estimate camera parameters through image analysis.
[0021] The technical features of the system 1 will be described below with reference to Fig. 1 again. As shown in Fig. 1, the system 1 includes a server 2 and a user terminal 3.
[0022] (Server 2) The server 2 is an information processing device having a control unit that controls various processes to enable synchronous playback of live-action footage and reproduced footage. The function of the server 2 as a control unit is realized by cooperation between hardware such as a processor and programs possessed by the server 2. The server 2 generates synchronous playback information based on the motion data and metadata of a moving object and outputs it to the user terminal 3. The server 2 may be, for example, a cloud server.
[0023] The server 2 estimates camera parameters of the live-action video based on a first feature amount of a moving object present in the target space and a second feature amount of the moving object captured in the live-action video captured of the target space. The target space is not limited to real space and may be a virtual space, or a space in which the real space and the virtual space interfere with each other. The moving object may be, for example, a human, an animal, a ball, a car, or a plant. By using the first feature amount and the second feature amount in combination, the server 2 can eliminate the inconveniences associated with prior art techniques that estimate camera parameters through image analysis.
[0024] The first feature amount is motion data including information indicating the three-dimensional position and posture of each of one or more parts constituting a moving object. Motion data has a smaller data volume and is more portable than images. Therefore, camera parameters can be estimated at a higher speed than when camera parameters are estimated by applying image analysis to live-action video. As a result, it is possible to quickly generate a reproduced video corresponding to the live-action video and play the reproduced video synchronously.
[0025] The second feature amount is the pose of the moving object shown in the 2D image. The server 2 renders a 3D model of the moving object based on the first feature amount into a 2D image based on candidate camera parameters, and estimates camera parameters of the live-action video based on the similarity between the pose of the moving object shown in the rendered 2D image and the pose of the moving object shown in the 2D image constituting the live-action video. For example, the server 2 estimates the candidate camera parameter with the highest similarity or when the similarity exceeds a predetermined threshold as the camera parameters of the live-action video. This configuration makes it possible to improve the accuracy of estimating the camera parameters of the live-action video.
[0026] The server 2 sets a first search range including one or more camera parameter candidates based on metadata of the live-action video, and estimates camera parameters from within the set first search range. The metadata includes information for identifying the shooting location of the live-action video. The server 2 sets the first search range to the vicinity of the shooting location of the live-action video indicated by the metadata. The server 2 then searches for camera parameters within the first search range by repeatedly setting camera parameter candidates and calculating similarities. This configuration reduces the amount of calculation required for searching for camera parameters, enabling high-speed estimation of camera parameters.
[0027] The first feature amount may be associated with identification information of the moving object. The server 2 may then calculate the similarity between moving objects associated with the same identification information. In particular, the server 2 may calculate the similarity between poses of moving objects associated with the same identification information that appear in a 2D image based on the motion data and a 2D image constituting the live-action video. The identification information of the moving object may be, for example, a player's name or uniform number. The identification information of the moving object that appears in the 2D image constituting the live-action video may be included in metadata or may be estimated by image analysis. This configuration makes it possible to further improve the estimation accuracy of the camera parameters.
[0028] The server 2 renders a 3D model of the moving object corresponding to the first feature acquired at an acquisition time corresponding to the shooting time of the live-action video into a 2D image based on candidate camera parameters, and calculates the similarity. The acquisition time corresponding to the shooting time refers to an acquisition time that coincides or nearly coincides with the shooting time. This configuration makes it possible to appropriately calculate the similarity, and as a result, to appropriately estimate the camera parameters.
[0029] The server 2 sets a second search range including one or more candidate acquisition times based on the playback speed of the live-action video, and estimates the acquisition time from within the set second search range. For example, if the playback speed is 1x, the server 2 may determine that the live-action video is video broadcast in real time and set the second search range to a range corresponding to several frames before and after the current time. On the other hand, if the playback speed is less than 1x, the server 2 may determine that the live-action video is a slow-motion playback of past video and set the second search range to a range from the current time to a predetermined time in the past. If the metadata includes information for identifying the shooting time of the live-action video, the server 2 may set the second search range further based on the metadata. Then, the server 2 may search for the acquisition time within the second search range by repeatedly setting candidate acquisition times and calculating similarities based on the first feature values acquired at the candidate acquisition times. This configuration reduces the amount of calculation required for searching for the acquisition time and enables high-speed estimation of the acquisition time.
[0030] The server 2 generates synchronized playback information that associates the live-action video, a first feature acquired at an acquisition time corresponding to the shooting time of the live-action video, and camera parameters of the live-action video. More specifically, the synchronized playback information associates a 2D image constituting the live-action video, a first feature acquired at an acquisition time corresponding to the shooting time of the 2D image, and camera parameters of the 2D image. With this configuration, the user terminal 3 can use the synchronized playback information to synchronously play back the live-action video and the reproduced video.
[0031] The synchronized playback information may include the playback speed of the live-action video. With this configuration, the user terminal 3 can display the playback speed when synchronously playing back the live-action video and the reproduced video.
[0032] (User Terminal 3) The user terminal 3 is a terminal device that controls the process of generating a playback screen in which live-action video and reproduced video are synchronously played back. The function of the user terminal 3 as a control unit is realized by cooperation between hardware such as a processor possessed by the user terminal 3 and a program. The user terminal 3 generates reproduced video based on the motion data of a moving object and synchronous playback information, and synchronously plays back the reproduced video and live-action video. The user terminal 3 may be, for example, a smartphone or a tablet terminal.
[0033] The user terminal 3 generates a playback screen in which a reproduced video is generated by rendering a 3D model generated based on the first feature amount of a moving object existing in the target space into a 2D image based on camera parameters of the live-action video captured of the target space, and the reproduced video is played back in synchronization with the live-action video. With this configuration, a user using the user terminal 3 can view the live-action video and the reproduced video that are played back in synchronization.
[0034] The user terminal 3 generates a playback screen that simultaneously displays the live-action video and the reproduced video, or that switchably displays either the live-action video or the reproduced video. With this configuration, the user can simultaneously view and compare the live-action video and the reproduced video, or view either one as needed.
[0035] The user terminal 3 generates a playback screen that plays back the reproduced video in synchronization with the live-action video based on the synchronous playback information. Specifically, the user terminal 3 plays back the live-action video while playing back the reproduced video generated based on the first feature acquired at the acquisition time corresponding to the shooting time of the live-action video and the camera parameters of the live-action video. With this configuration, the user terminal 3 does not need to estimate the acquisition time and the camera parameters, which reduces the processing load required for synchronous playback of the reproduced video and the live-action video.
[0036] The playback screen may include the playback speed of the live-action video. With this configuration, the user can view the reproduced video while being aware of the playback speed of the live-action video.
[0037] The technical features of the system 1 have been described above.
[0038] 2. First Embodiment The first embodiment is an embodiment in which live-action video and reconstructed video are synchronously played back for a soccer match held at a soccer stadium.
[0039] 2.1. Example of configuration (Example of configuration related to distribution of reproduced video) Fig. 2 is a block diagram showing the configuration related to reproduced video of the system 1 according to this embodiment. As shown in Fig. 2, the system 1 includes multiple cameras 11, a tracking system 12, a streamer 13, a converter 14, a metadata labeler 15, multiple players 16, and multiple renderers 17.
[0040] A plurality of cameras 11 are arranged around the soccer field, and simultaneously capture images of the soccer match from various angles, and output the acquired video data to a tracking system 12.
[0041] The tracking system 12 acquires motion data of the match based on the video data obtained by the multiple cameras 11 and outputs the data to the streamer 13. The motion data of the match may include, for example, the three-dimensional position and orientation of each part of a soccer player (e.g., bones of a skeletal model) and the three-dimensional position of a soccer ball. The motion data is output as time-series data together with time data indicating the time when the motion data was acquired.
[0042] The streamer 13 distributes a stream of motion data of the match. Specifically, the streamer 13 distributes the motion data of all players and the ball at each time as a stream.
[0043] The converter 14 converts the motion data distributed from the streamer 13 into a predetermined format and distributes it. One example of the predetermined format is a format suitable for distribution and playback, such as the MPEG-DASH standard or a proprietary format based on the MPEG-DASH standard. The converter 14 is connected to the Internet and provides an interface that allows multiple players 16 to simultaneously obtain the converted data.
[0044] The metadata labeler 15 distributes match metadata. In particular, the metadata labeler 15 outputs the match metadata in association with the motion data output from the converter 14. The metadata is distributed as time-series data along with time data indicating the time the metadata was added. Note that the metadata may be added manually or automatically via the Internet, etc. The metadata labeler 15, like the converter 14, is connected to the Internet and provides an interface that allows multiple players 16 to simultaneously obtain the metadata.
[0045] When a soccer match begins, the match motion data and match metadata are distributed over the Internet with a delay of a few seconds.
[0046] The player 16 acquires the distributed motion data and metadata.
[0047] The renderer 17 renders and displays a reproduced image based on the stream of motion data acquired by the player 16. More specifically, the renderer 17 renders a 3D model based on the motion data into a 2D image and displays it in chronological order.
[0048] (Configuration example for broadcasting live-action video) Fig. 3 is a block diagram showing the configuration for live-action video of the system 1 according to this embodiment. As shown in Fig. 3, the system 1 includes multiple cameras 21, a switcher 22, an editor 23, an encoder 24, a transmitter 25, a receiver 26, a decoder 27, and a renderer 28.
[0049] The cameras 21 are placed near the field, and simultaneously capture images of the soccer match from various angles, and output video data of the captured live images (i.e., time-series data of 2D images).
[0050] The switcher 22 selects and outputs the video data of the live-action video to be broadcast from the video data of the live-action video captured by the plurality of cameras 21 .
[0051] The editor 23 edits the video data of the live-action video. For example, the editor 23 sets the playback speed of the live-action video.
[0052] The encoder 24 encodes the video data of the live-action video that has been edited by the editor 23 into a predetermined format suitable for broadcasting, and outputs the encoded data to the transmitter 25 .
[0053] The transmitter 25 broadcasts the encoded video data of the actual video.
[0054] The receiver 26 receives the video data of the broadcasted live-action video and outputs it to the decoder 27 .
[0055] The decoder 27 decodes the video data of the live-action video and outputs it to the renderer 28 .
[0056] The renderer 28 renders and displays live-action video based on the decoded data.
[0057] (Configuration Example of Server 2) Fig. 4 is a block diagram showing a configuration example of the server 2 according to this embodiment. As shown in Fig. 4, the server 2 includes a player 31, a receiver 32, a decoder 33, a camera parameter estimation unit 34, a playback speed estimation unit 35, and a synchronized playback information generation unit 36.
[0058] The player 31 acquires the motion data and metadata delivered by the converter 14 and metadata labeler 15 shown in Fig. 2 and outputs them to the camera parameter estimation unit 34. The player 31 corresponds to the player 16 shown in Fig. 2.
[0059] 3 and outputs the video data of the live-action video broadcast by the transmitter 25 shown in FIG. 3 to the decoder 33. The receiver 32 corresponds to the receiver 26 shown in FIG.
[0060] The decoder 33 decodes the video data of the live-action video and outputs the decoded video data to the camera parameter estimation unit 34 and the playback speed estimation unit 35. The decoder 33 corresponds to the decoder 27 shown in FIG.
[0061] The camera parameter estimation unit 34 estimates the camera parameters of the live-action video (i.e., the camera parameters of the camera 21 that captured the live-action video) based on the motion data, metadata, and video data of the live-action video. The detailed configuration of the camera parameter estimation unit 34 will be described in detail later.
[0062] An example of camera parameters for a real-life video will be described with reference to FIG.
[0063] FIG. 5 is a diagram illustrating camera parameters in this embodiment. As shown in FIG. 5, the camera parameters may include a three-dimensional position (x, y, z) consisting of the x-coordinate, y-coordinate, and z-coordinate of the camera 21 relative to the origin O of the world coordinate system. The camera parameters may also include an orientation (θ, ω, γ) consisting of the roll angle θ, pitch angle ω, and yaw angle γ of the camera 21 in the world coordinate system. Furthermore, the camera parameters may also include the angle of view α of the camera 21. The camera parameters of the live-action video may include at least one of the three-dimensional position, orientation, and angle of view of the camera 21 that captured the live-action video.
[0064] The playback speed estimation unit 35 estimates the playback speed of the live-action video. For example, the playback speed estimation unit 35 estimates a playback speed of 1x for live-action video broadcast in real time, a playback speed greater than 1x for live-action video replayed at 2x or 3x speed, etc., and a playback speed less than 1x for live-action video replayed in slow motion. The detailed configuration of the playback speed estimation unit 35 will be described in detail later.
[0065] The synchronized playback information generating unit 36 generates synchronized playback information based on the camera parameters estimated by the camera parameter estimating unit 34 and the playback speed estimated by the playback speed estimating unit 35. Then, the synchronized playback information generating unit 36 transmits the generated synchronized playback information.
[0066] An example of the synchronous playback information is shown in Table 1 below.
[0067]
[0068] The video frame number is an identification number for a 2D image that constitutes the live-action video. The motion frame number is an identification number for each piece of motion data that constitutes the motion data stream. The video frame number, motion frame number, and camera parameters stored in the same row in Table 1 indicate a 2D image that constitutes the live-action video, the motion data acquired at the acquisition time corresponding to the shooting time of the 2D image, and the camera parameters of the 2D image.
[0069] (Configuration Example of User Terminal 3) Fig. 6 is a block diagram showing a configuration example of the user terminal 3 according to this embodiment. As shown in Fig. 6, the user terminal 3 includes a receiver 41, a decoder 42, a synchronous player 43, and a renderer 44. These components may be incorporated into, for example, an application or a browser.
[0070] The receiver 41 receives video data of broadcast live-action video and outputs it to the decoder 42. The receiver 41 corresponds to the receiver 26 shown in FIG.
[0071] The decoder 42 decodes the video data of the live-action video and outputs it to the synchronous player 43. The decoder 42 corresponds to the decoder 27 shown in FIG.
[0072] The synchronization player 43 receives the motion data and metadata delivered by the converter 14 and metadata labeler 15 shown in Fig. 2, as well as the synchronization playback information sent by the synchronization playback information generator 36 shown in Fig. 4. Based on the synchronization playback information, the synchronization player 43 outputs motion data synchronized with the live-action video (i.e., acquired at an acquisition time corresponding to the shooting time of the live-action video), camera parameters of the live-action video, and a playback speed of the live-action video. Specifically, the synchronization player 43 outputs the motion data, camera parameters, and playback speed of the motion frame number corresponding to the video frame number of the live-action video output from the decoder 42 in the synchronization playback information.
[0073] The renderer 44 renders and displays the live-action video and the reproduced video. Specifically, the renderer 44 renders and displays the live-action video based on the video data output from the decoder 42. At the same time, the renderer 44 renders and displays a 2D image of a 3D model based on the motion data output from the synchronous player 43 based on the camera parameters output from the synchronous player 43. Furthermore, the renderer 44 renders and displays the playback speed output from the synchronous player 43.
[0074] An example of a playback screen displayed by the user terminal 3 will be described with reference to Fig. 7 to Fig. 9. Fig. 7 to Fig. 9 are diagrams showing an example of a playback screen displayed by the user terminal 3 according to this embodiment.
[0075] As shown in Fig. 7, the user terminal 3 may simultaneously display a reproduced video V1 and an actual video V2 on a single playback screen D1. As shown in Fig. 7, the reproduced video V1 is rendered using the same camera parameters as the actual video V2. As shown in Fig. 8, when the playback speed specified in the synchronized playback information is other than 1x, information D21 indicating the playback speed may be clearly displayed on the playback screen D2.
[0076] 9, one playback screen D3 may display either a reproduced video V1 or an actual video V2. The user terminal 3 may seamlessly switch the video to be played back based on a user operation.
[0077] Furthermore, the user terminal 3 may accept camera parameter manipulation by the user. In this case, the user terminal 3 renders a 3D model based on the motion data into a 2D image based on the camera parameters specified by the user and displays the 2D image as the reproduced video V1. With this configuration, the user can freely move the viewpoint and zoom in / out while viewing the reproduced video V1 that is played back in synchronization with the live-action video V2.
[0078] 10 is a block diagram showing an example of the configuration of the camera parameter estimation unit 34 according to this embodiment. As shown in Fig. 10, the camera parameter estimation unit 34 includes an input interface 110, a camera parameter analysis block 120, a motion data analysis block 130, an image analysis block 140, a buffer memory 150, and an output interface 160.
[0079] (Input Interface 110) The input interface 110 is an interface that accepts input of information to the camera parameter estimation unit 34. As shown in Fig. 10 , the input interface 110 includes a metadata input unit 111, a motion data input unit 112, and a video data input unit 113.
[0080] The metadata input unit 111 acquires metadata and outputs it to the camera parameter analysis block 120 .
[0081] The motion data input unit 112 acquires motion data and outputs it to the motion data analysis block 130 .
[0082] The video data input unit 113 acquires video data of live-action video and outputs it to the image analysis block 140 .
[0083] (Camera Parameter Analysis Block 120) The camera parameter analysis block 120 is a processing block that analyzes camera parameters. As shown in Fig. 10, the camera parameter analysis block 120 includes a metadata analysis unit 121 and a camera parameter search unit 122.
[0084] The metadata analysis unit 121 analyzes the metadata input by the metadata input unit 111 and sets a search range (corresponding to the first search range) of camera parameters for the live-action video. The metadata includes information for identifying the shooting time and shooting location of the live-action video. Therefore, the metadata analysis unit 121 sets the search range to include camera parameters that can capture the shooting location identified by the metadata.
[0085] The metadata may include absolute information such as a timestamp, latitude, longitude, and altitude, or may include information that can be processed into absolute information such as "a match between soccer team DEF and soccer team GHI held at ABC Stadium on February 24, 2024." The metadata may also include information that can be used to identify moving objects in live-action footage, such as the names of the teams playing the game, the names of the players, each player's jersey number, the color of each player's uniform, score information, referee information such as fouls, and the position of the ball at the time of filming. The metadata may also include data other than the data described above.
[0086] The metadata analysis unit 121 may set the search range of the camera parameters further based on the live-action video. As an example, when a ball appears in the live-action video, the metadata analysis unit 121 may set the range in which the ball fits within the angle of view as the search range. As another example, the metadata analysis unit 121 may set the range in which the players appearing in the live-action video fit within the angle of view as the search range.
[0087] The camera parameter searching unit 122 searches for camera parameters of the live-action video within a search range. Specifically, the camera parameter searching unit 122 outputs candidate camera parameters within the search range to the pose comparing unit 132. Then, the camera parameter searching unit 122 identifies candidate camera parameters whose comparison results by the pose comparing unit 132 satisfy predetermined conditions as camera parameters of the live-action video.
[0088] (Image analysis block 140) The image analysis block 140 is a processing block that analyzes 2D images for each video frame that constitutes video data of live-action video. As shown in Fig. 10 , the image analysis block 140 includes a video data decoding unit 141 and a pose estimation unit 142.
[0089] The video data decoding unit 141 decodes the video data of the live-action video input by the video data input unit 113. For example, the video data decoding unit 141 decodes the video data of the live-action video into a predetermined format.
[0090] The buffer memory 150 buffers the 2D images for each video frame that constitute the live-action video decoded by the video data decoding unit 141 .
[0091] The pose estimation unit 142 estimates the pose of a moving object appearing in a 2D image constituting the live-action video. The pose here may refer to the position and orientation in the 2D image of each of one or more parts constituting the moving object appearing in the 2D image. For example, the pose estimation unit 142 applies any pose recognition algorithm such as OpenPose to the 2D image buffered by the buffer memory 150 to estimate the pose of the moving object appearing in the live-action video. The pose estimation unit 142 outputs the estimated pose of the moving object to the pose comparison unit 132.
[0092] (Motion Data Analysis Block 130) The motion data analysis block 130 is a processing block that analyzes motion data. As shown in Fig. 10, the motion data analysis block 130 includes a motion data decoding unit 131 and a pose comparison unit 132.
[0093] The motion data decoding unit 131 decodes the motion data. For example, the motion data decoding unit 131 decodes the motion data into a predetermined format, such as setting the number of bones to a predetermined value.
[0094] The pose comparison unit 132 compares the pose of the moving object in the 2D image constituting the live-action video with the pose of the 3D model based on the motion data. More specifically, the pose comparison unit 132 renders a 2D image of the 3D model based on the motion data decoded by the motion data analysis block 130, based on the candidate camera parameters output from the camera parameter search unit 122. The pose comparison unit 132 then calculates the similarity between the pose of the moving object estimated by the pose estimation unit 142 and the pose of the 3D model in the 2D image rendered based on the motion data, and outputs the similarity to the camera parameter search unit 122.
[0095] If the similarity calculated by the pose comparison unit 132 is equal to or smaller than a predetermined threshold, the camera parameter search unit 122 continues to calculate the similarity by the pose comparison unit 132 while changing the candidate camera parameters. Then, the camera parameter search unit 122 identifies the candidate camera parameters for which the similarity calculated by the pose comparison unit 132 exceeds the predetermined threshold as the camera parameters of the live-action video. In this way, the camera parameter search unit 122 sequentially searches for camera parameters of the live-action video by changing the candidate camera parameters until the similarity calculated by the pose comparison unit 132 exceeds the predetermined threshold.
[0096] The threshold and the change amount of the candidate camera parameters may be constant or may be changed as appropriate. For example, the threshold may be set low and the change amount may be set large at the start of the search. If the similarity subsequently exceeds the threshold, the threshold may be updated to a higher value and the change amount may be updated to a smaller value, and the search may continue.
[0097] (Supplementary Note) While the above describes the search for camera parameters, the same method may be used to search for the acquisition time of motion data corresponding to the shooting time of live-action video. For example, the metadata analysis unit 121 may set a search range (corresponding to a second search range) for the acquisition time of motion data corresponding to the shooting time of live-action video, based on the metadata. Then, the camera parameter search unit 122 may search within the search range for the acquisition time of motion data corresponding to the shooting time of live-action video.
[0098] Specifically, the camera parameter search unit 122 may output candidate camera parameters within the first search range and candidate acquisition times within the second search range to the pose comparison unit 132. In this case, the pose comparison unit 132 renders a 3D model based on motion data acquired at the candidate acquisition times into a 2D image based on the candidate camera parameters, and calculates the similarity. The camera parameter search unit 122 then identifies a combination of candidate camera parameters and candidate acquisition times whose similarity exceeds a predetermined threshold as the acquisition time of the motion data corresponding to the camera parameters and shooting time of the live-action video.
[0099] (Output Interface 160) The output interface 160 is an interface that outputs information from the camera parameter estimation unit 34. As shown in FIG.
[0100] The camera parameter output unit 161 outputs the camera parameters of the actual video image identified by the camera parameter search unit 122 .
[0101] Specific Examples FIGS. 11 to 14 are diagrams for explaining specific examples of processing by the camera parameter estimation unit 34 according to this embodiment.
[0102] 11 , the metadata analysis unit 121 sets a search range 51 for camera parameters of the live-action video. Next, the camera parameter search unit 122 sets candidate camera parameters from within the camera parameter search range 51. Then, the pose comparison unit 132 renders a 3D model 52 based on the motion data into a 2D image 53 based on the candidate camera parameters.
[0103] As shown in FIG. 12, the pose estimation unit 142 analyzes the 2D image 54 of the real-life video to estimate poses 541 to 544 of the moving object appearing in the 2D image 54 of the real-life video.
[0104] Then, pose comparison unit 132 compares the pose of 3D model 52 shown in 2D image 53 based on the motion data with poses 541 to 544 of the moving object shown in 2D image 54 of the live-action video. Camera parameter search unit 122 searches for camera parameters of the live-action video based on the comparison result by pose comparison unit 132 while changing candidate camera parameters.
[0105] For example, as shown in Fig. 13, when poses 531-534 of 3D model 52 shown in 2D image 53 based on motion data are not similar to poses 541-544 of the moving object shown in 2D image 54 of live-action video, camera parameter searching unit 122 changes the candidate camera parameters and continues the search. On the other hand, as shown in Fig. 14, when poses 531-534 of 3D model 52 shown in 2D image 53 based on motion data are similar to poses 541-544 of the moving object shown in 2D image 54 of live-action video, camera parameter searching unit 122 identifies the candidate camera parameters as the camera parameters of the live-action video.
[0106] (2) Playback Speed Estimation Unit 35 Fig. 15 is a block diagram showing an example of the configuration of the playback speed estimation unit 35 according to this embodiment. As shown in Fig. 15, the playback speed estimation unit 35 includes a motion database 210, a playback speed learning block 220, an input interface 230, a playback speed estimation block 240, a trained network 250, and an output interface 260.
[0107] (Trained Network 250) The trained network 250 is a network that outputs the playback speed of a video when the video shows a moving object. The trained network 250 may be configured, for example, by a neural network or a recurrent neural network. The trained network 250 is trained by the training unit 223, which will be described later.
[0108] (Motion Database 210) The motion database 210 is a database that stores time-series data of motion data that is being played back in real time (i.e., at a playback speed of 1x).
[0109] (Playback Speed Learning Block 220) The playback speed learning block 220 is a processing block that constructs the trained network 250. As shown in FIG. 15 , the playback speed learning block 220 includes a motion data decoding unit 221, a variation generating unit 222, and a learning unit 223.
[0110] The motion data decoding unit 221 decodes the time-series data of the motion data. Then, the motion data decoding unit 221 associates the decoded time-series data of the motion data with a playback speed of 1x and outputs the data to the motion data decoding unit 221 and the variation generating unit 222.
[0111] The variation generation unit 222 generates variations of the time-series data of the motion data. For example, the variation generation unit 222 generates time-series data of the motion data at a playback speed other than 1x, associates it with the playback speed, and outputs it to the learning unit 223.
[0112] The learning unit 223 learns the parameters of the trained network 250 based on a combination of the playback speed and the time-series data of the motion data. In detail, the learning unit 223 learns the parameters of the trained network 250 using, as training data, a combination of the playback speed and the 3DCG animation generated based on the time-series data of the motion data.
[0113] (Input Interface 230) The input interface 230 is an interface that accepts information input to the playback speed estimation unit 35. As shown in FIG.
[0114] The video data input unit 231 acquires video data of live-action video and outputs it to the playback speed estimation block 240 .
[0115] (Playback Speed Estimation Block 240) The playback speed estimation block 240 is a processing block that estimates the playback speed of live-action video. As shown in Fig. 15, the playback speed estimation block 240 includes a video data decoding unit 241 and an inference unit 242.
[0116] The video data decoding unit 241 decodes the video data of the actual video and outputs it to the inference unit 242 .
[0117] The inference unit 242 estimates the playback speed of the live-action video by performing inference using the trained network 250. Specifically, the inference unit 242 estimates the pose of a moving object appearing in a 2D image for each video frame that constitutes the live-action video. The inference unit 242 then inputs an animation consisting of time-series data of the estimated pose of the moving object into the trained network 250, and acquires the output playback speed as an estimated result of the playback speed of the live-action video. The inference unit 242 outputs the estimated playback speed of the live-action video to the output interface 260.
[0118] (Output Interface 260) The output interface 260 is an interface that outputs information from the playback speed estimation unit 35. As shown in FIG.
[0119] The playback speed output unit 261 outputs the playback speed output from the inference unit 242 .
[0120] Specific Example FIGS. 16 and 17 are diagrams for explaining a specific example of processing by the playback speed estimation unit 35 according to this embodiment.
[0121] 16, the learning unit 223 generates a 3DCG animation 61 based on time-series data of motion data at a playback speed of 1x and various playback speeds. Then, the learning unit 223 learns parameters of the trained network 250 using combinations of the playback speeds and the 3DCG animation 61 as training data.
[0122] 17 , the inference unit 242 estimates the pose of a moving object appearing in a broadcast live-action video for each 2D image constituting the live-action video, and generates an animation 62 that is time-series data of the pose of the estimated moving object. The inference unit 242 then inputs the generated animation 62 to the trained network 250, thereby estimating the playback speed.
[0123] 2.3. Processing Flow> (Camera Parameter Estimation Processing) FIG. 18 is a block diagram showing an example of the flow of camera parameter estimation processing executed by the server 2 according to this embodiment.
[0124] As shown in FIG. 18, first, the server 2 acquires video data, metadata, and motion data of live-action video (step S102).
[0125] Next, the server 2 selects a video frame to be processed (step S104). The video frame number of the video frame to be processed is incremented each time this step is executed, with the initial value being -1.
[0126] Next, the server 2 inputs one frame of the 2D image to be processed (step S106), and estimates the pose of the moving object shown in the 2D image (step S108).
[0127] In parallel, the server 2 estimates the playback speed of the live-action video (step S110).
[0128] Next, the server 2 sets candidate motion data acquisition times corresponding to the shooting time of the live-action video based on the estimated playback speed of the live-action video (step S112). As an example, if the playback speed is 1x, the server 2 may determine that the live-action video is video broadcast in real time and set a range of times corresponding to several frames before and after the current time as candidate motion data acquisition times corresponding to the shooting time of the live-action video. On the other hand, if the playback speed is less than 1x, the server 2 may determine that the live-action video is a slow-motion playback of past video and set a range of times from the current time to a predetermined time in the past as candidate motion data acquisition times corresponding to the shooting time of the live-action video. Furthermore, if the playback speed is more than 1x, the server 2 may determine that the live-action video is a high-speed replay of past video and set a range of times from the current time to a predetermined time in the past as candidate motion data acquisition times corresponding to the shooting time of the live-action video.
[0129] Next, the server 2 sets a search range for the camera parameters based on the metadata (step S114).
[0130] Next, the server 2 acquires one frame of motion data acquired at the candidate motion data acquisition time corresponding to the shooting time of the live-action video set in step S112 (step S116).
[0131] Next, the server 2 determines whether or not there are any unverified camera parameters within the search range of the camera parameters set in step S114 (step S118). Note that verification here refers to the calculation of the similarity and comparison with the threshold in step S124.
[0132] If it is determined that there are no unverified camera parameters (step S118: NO), the process proceeds to step S128.
[0133] If it is determined that there are unverified camera parameters (step S118: YES), the server 2 sets the unverified camera parameters as candidates for camera parameters (step S120).
[0134] Next, the server 2 renders the 3D model based on the motion data acquired in step S116 as a pose on a 2D image based on the candidate camera parameters set in step S120 (step S122).
[0135] Next, the server 2 calculates the similarity between the pose estimated in step S108 and the pose rendered in step S122. Then, the server 2 determines whether the calculated similarity is equal to or greater than a threshold value (step S124).
[0136] If it is determined that the similarity is less than the threshold value (step S124: NO), the process returns to step S118 again.
[0137] If it is determined that the similarity is equal to or greater than the threshold value (step S124: YES), the server 2 outputs the candidate camera parameters used for rendering in step S122 as the camera parameters for the live-action video (step S126).
[0138] At this time, the server 2 can generate synchronized playback information that associates the video frame selected in step S104, the motion data acquired in step S116, and the camera parameters output in step S126.
[0139] Thereafter, the server 2 determines whether or not the next video frame exists (step S128).
[0140] If it is determined that the next video frame exists (step S128: YES), the process returns to step S104.
[0141] On the other hand, if it is determined that the next video frame does not exist (step S128: NO), the process ends.
[0142] (Synchronized Playback Processing) FIG. 19 is a block diagram showing an example of the flow of synchronized playback processing executed by the user terminal 3 according to this embodiment.
[0143] As shown in FIG. 19, the user terminal 3 acquires synchronized playback information (step S202).
[0144] Next, the user terminal 3 reads synchronized playback information for one frame (step S204).
[0145] Next, the user terminal 3 acquires the video frame number of the synchronized playback information read in step S204 (step S206).
[0146] Next, the user terminal 3 acquires the motion data of the motion frame number corresponding to the video frame number acquired in step S206 (step S208).
[0147] Next, the user terminal 3 acquires the camera parameters corresponding to the video frame number acquired in step S206 (step S210).
[0148] Next, the user terminal 3 renders a reproduced image of the 3D model based on the motion data acquired in step S208, based on the camera parameters acquired in step S210 (step S212).
[0149] On the other hand, the user terminal 3 acquires the video data of the live-action video of the video frame number acquired in step S206 (step S214).
[0150] Next, the user terminal 3 renders the live-action video based on the video data of the live-action video acquired in step S214 (step S216).
[0151] The user terminal 3 then synchronously displays the reproduced video rendered in step S212 and the live-action video rendered in step S216 on the playback screen (step S218).
[0152] On the other hand, the user terminal 3 acquires the playback speed of the synchronized playback information read in step S204 (step S220).
[0153] Next, the user terminal 3 displays the playback speed acquired in step S220 on the playback screen (step S222).
[0154] Thereafter, the user terminal 3 determines whether or not synchronized playback has ended (step S224).
[0155] If it is determined that the synchronous playback should not be ended (step S224: NO), the process returns to step S204 again.
[0156] On the other hand, if it is determined that synchronous playback is to be ended (step S224: YES), the process ends.
[0157] 3. Second Embodiment A system 1 according to a second embodiment is an embodiment in which live-action video and replayed video are synchronously played back for a user wearing a motion sensor. Specifically, the system 1 synchronously plays back live-action video of a user wearing a motion sensor captured by a user terminal 3 and replayed video based on motion data obtained by the motion sensor. The system 1 according to this embodiment can be used, for example, to improve form in sports instruction.
[0158] 20 is a block diagram showing a configuration related to the reproduction video of the system 1 according to this embodiment. As shown in FIG. 20, the system 1 includes a plurality of motion sensors 71, a data receiver 72, a motion capture device 73, and a transmitter 74.
[0159] The motion sensors 71 are attached to various parts of the user's body. For example, the motion sensors 71 are attached to the user's head, both arms, both hands, waist, and both feet. The motion sensors 71 acquire sensor data related to the user's motion data and transmit the sensor data to the data receiver 72. The motion sensors 71 may include, for example, an acceleration sensor and an angular velocity sensor, and the sensor data may be, for example, time-series data of acceleration and angular velocity.
[0160] The data receiver 72 receives sensor data from the multiple motion sensors 71 and outputs it to the motion capture 73 .
[0161] The motion capture 73 acquires motion data of the user based on sensor data acquired by the multiple motion sensors 71 .
[0162] The transmitter 74 transmits time-series data of the motion data acquired by the motion capture 73 .
[0163] (Configuration Example for Distribution of Live-Action Video) Fig. 21 is a block diagram showing the configuration for live-action video of the system 1 according to this embodiment. As shown in Fig. 21, the system 1 includes a camera 81, an encoder 82, and a transmitter 83. These components can be arranged in the user terminal 3.
[0164] The camera 81 captures an image of the user wearing the motion sensor 71 and outputs image data of the captured image.
[0165] The encoder 82 encodes the video data of the live video captured by the camera 81 into a predetermined format suitable for distribution and outputs it to the transmitter 83 .
[0166] The transmitter 83 transmits the encoded video data of the actual video.
[0167] Before transmitting the video data of the live-action video, editing such as cutting and changing the playback speed may be performed.
[0168] The user terminal 3 may also transmit metadata associated with the video data of the live-action video. The metadata may include, for example, the shooting time of the live-action video and location information at the time of shooting. The metadata may also include information that can be used to identify moving objects in the live-action video, such as the name of the subject.
[0169] (Configuration Example of Server 2) The configuration example of the server 2 is as described in the first embodiment with reference to FIG. 4 and the like.
[0170] That is, the server 2 estimates the camera parameters of the live-action video (i.e., the camera parameters of the camera 81 that captured the live-action video) based on the video data of the live-action video captured by the camera 81 and the motion data captured by the motion sensor 71. The server 2 then generates synchronized playback information that associates the live-action video, motion data acquired at the acquisition time of the motion data corresponding to the capture time of the live-action video, and the camera parameters of the live-action video. The server 2 then distributes the generated synchronized playback information in association with the live-action video and the motion data.
[0171] The server 2 may also estimate the playback speed of the live-action video, and then generate synchronized playback information including the estimated playback speed.
[0172] (Configuration Example of User Terminal 3) The configuration example of the user terminal 3 is as described in the first embodiment with reference to FIG. 6 and the like.
[0173] That is, the user terminal 3 receives the live-action video, motion data, and synchronous playback information, and synchronously plays back the live-action video and the reproduced video based on the received information. Specifically, the user terminal 3 generates a reproduced video by rendering a 3D model based on the motion data acquired at the acquisition time of the motion data corresponding to the shooting time of the live-action video into a 2D image based on the camera parameters of the live-action video, and plays back the reproduced video together with the live-action video.
[0174] (Detailed Configuration) The detailed configurations of the camera parameter estimation unit 34 and the playback speed estimation unit 35 according to this embodiment are as described in the first embodiment with reference to FIGS. 10 and 15, etc.
[0175] (Processing Flow) The processing flow by the server 2 and the user terminal 3 according to this embodiment is the same as that described in the first embodiment with reference to FIGS. 18 and 19 .
[0176] 4. Example of Hardware Configuration Finally, the hardware configuration of the information processing device according to each of the above-described embodiments will be described with reference to Fig. 22. Fig. 22 is a block diagram showing an example of the hardware configuration of the information processing device according to each of the above-described embodiments. Note that the information processing device 900 shown in Fig. 22 may realize, for example, the server 2 shown in Fig. 4 or the user terminal 3 shown in Fig. 5. Information processing by the server 2 or the user terminal 3 according to each of the above-described embodiments is realized by cooperation between software and hardware described below.
[0177] 22 , the information processing device 900 includes a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, a RAM (Random Access Memory) 903, and a host bus 904a. The information processing device 900 also includes a bridge 904, an external bus 904b, an interface 905, an input device 906, an output device 907, a storage device 908, a drive 909, a connection port 910, and a communication device 912. The information processing device 900 may include a processing circuit such as an electric circuit, a DSP, or an ASIC instead of or in addition to the CPU 901.
[0178] The CPU 901 functions as an arithmetic processing device and control device, and controls the overall operation of the information processing device 900 in accordance with various programs. The CPU 901 may also be a microprocessor. The ROM 902 stores programs and calculation parameters used by the CPU 901. The RAM 903 temporarily stores programs used in execution by the CPU 901 and parameters that change as appropriate during execution. The CPU 901 may form, for example, a control unit that controls the overall processing of the server 2 or a control unit that controls the overall processing of the user terminal 3.
[0179] The CPU 901, ROM 902, and RAM 903 are interconnected by a host bus 904a, which includes a CPU bus and the like. The host bus 904a is connected to an external bus 904b, such as a PCI (Peripheral Component Interconnect / Interface) bus, via a bridge 904. Note that the host bus 904a, bridge 904, and external bus 904b do not necessarily need to be configured separately, and these functions may be implemented on a single bus.
[0180] The input device 906 is realized by a device into which a user inputs information, such as a mouse, keyboard, touch panel, button, microphone, switch, or lever. The input device 906 may also be, for example, a remote control device that uses infrared or other radio waves, or an externally connected device such as a mobile phone or PDA that is compatible with the operation of the information processing device 900. Furthermore, the input device 906 may include, for example, an input control circuit that generates an input signal based on information input by the user using the above-mentioned input means and outputs the signal to the CPU 901. By operating the input device 906, the user of the information processing device 900 can input various data to the information processing device 900 and instruct processing operations.
[0181] The output device 907 is formed by a device capable of visually or audibly notifying the user of acquired information. Examples of such devices include display devices such as CRT display devices, liquid crystal display devices, plasma display devices, EL display devices, laser projectors, LED projectors, and lamps, audio output devices such as speakers and headphones, and printer devices. The output device 907 outputs, for example, results obtained from various processes performed by the information processing device 900. Specifically, the display device visually displays the results obtained from various processes performed by the information processing device 900 in various formats, such as text, images, tables, and graphs. On the other hand, the audio output device converts audio signals consisting of reproduced audio data, acoustic data, etc. into analog signals and outputs them audibly. The display device can, for example, synchronously reproduce live-action video and reproduced video on the user terminal 3.
[0182] The storage device 908 is a data storage device formed as an example of a storage unit of the information processing device 900. The storage device 908 is realized, for example, by a magnetic storage device such as an HDD, a semiconductor storage device, an optical storage device, or a magneto-optical storage device. The storage device 908 may include a storage medium, a recording device that records data on the storage medium, a reading device that reads data from the storage medium, and a deletion device that deletes data recorded on the storage medium. This storage device 908 stores programs executed by the CPU 901, various data, and various data acquired from outside. The storage device 908 may temporarily store motion data, metadata, and video data in the server 2 or the user terminal 3, for example.
[0183] The drive 909 is a reader / writer for a storage medium, and is built into or externally attached to the information processing device 900. The drive 909 reads information recorded on a removable storage medium such as an attached magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, and outputs the information to the RAM 903. The drive 909 can also write information to the removable storage medium.
[0184] The connection port 910 is an interface connected to an external device, and is a connection port for connecting to an external device that can transmit data via, for example, a USB (Universal Serial Bus).
[0185] The communication device 912 is, for example, a communication interface formed by a communication device for connecting to the network 920. The communication device 912 is, for example, a communication card for a wired or wireless local area network (LAN), a long term evolution (LTE), Bluetooth (registered trademark), or a wireless USB (WUSB). The communication device 912 may also be a router for optical communication, a router for an asymmetric digital subscriber line (ADSL), or a modem for various communication purposes. The communication device 912 can transmit and receive signals, for example, between the Internet and other communication devices in accordance with a predetermined protocol such as TCP / IP. The communication device 912 can transmit and receive synchronized playback information between the server 2 or the user terminal 3.
[0186] The network 920 is a wired or wireless transmission path for information transmitted from devices connected to the network 920. For example, the network 920 may include public network such as the Internet, a telephone network, or a satellite communication network, as well as various LANs (Local Area Networks) and WANs (Wide Area Networks) including Ethernet (registered trademark). The network 920 may also include a dedicated network such as an IP-VPN (Internet Protocol-Virtual Private Network).
[0187] The above describes an example of a hardware configuration capable of realizing the functions of the information processing device 900 according to each of the above embodiments. Each of the above components may be realized using general-purpose components, or may be realized by hardware specialized for the function of each component. Therefore, the hardware configuration used can be changed as appropriate depending on the technical level at the time of implementing each of the above embodiments.
[0188] <5. Supplementary Information> Although preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0189] For example, in the above embodiment, an example has been described in which video data of live-action video is distributed and played back, but audio data may be distributed and played back together with the video data.
[0190] For example, in the above embodiment, an example has been described in which the system 1 is applied to live broadcasts of sports games or sports coaching, but the present technology is not limited to such examples. As an example, the system 1 may be applied to live broadcasts of music concerts, or recording / playback of stage performances such as dance or theater. As another example, the system 1 may be applied to computer games or AR (Augmented Reality) applications that seamlessly switch between live-action footage and 3D CG animation.
[0191] Note that each device described in this specification may be realized as a single device, or some or all of them may be realized as separate devices. Furthermore, at least a portion of the processing described in this specification as being performed by a specific device may be performed by any other device. As an example, the user terminal 3 may include a player 31, a receiver 32, a decoder 33, a camera parameter estimation unit 34, a playback speed estimation unit 35, and a synchronized playback information generation unit 36, and generate synchronized playback information. As another example, the server 2 may include a receiver 41, a decoder 42, a synchronized player 43, and a renderer 44, and the user terminal 3 may display video rendered by the server 2 on the Web.
[0192] The series of processes performed by each device described herein may be implemented using software, hardware, or a combination of software and hardware. The software programs may be stored in advance, for example, on a recording medium (more specifically, a non-transitory computer-readable storage medium) internal or external to each device. Each program is then loaded into a random access memory (RAM) and executed by a processing circuit such as a central processing unit (CPU). The recording medium may be, for example, a magnetic disk, an optical disk, a magneto-optical disk, or a flash memory. The computer program may also be distributed, for example, via a network, without using a recording medium. The computer may be an application-specific integrated circuit (ASIC), a general-purpose processor that executes functions by loading a software program, or a computer on a server used in cloud computing. The series of processes performed by each device described herein may be centrally processed by a single computer or distributed across multiple computers. Furthermore, in each of the above embodiments, two or more communication means present in a single device may be physically implemented on a single medium.
[0193] Furthermore, the processes described herein using flowcharts or sequence diagrams do not necessarily have to be performed in the order shown. Some process steps may be performed in parallel. Furthermore, additional process steps may be employed, and some process steps may be omitted.
[0194] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0195] Note that the following configurations also fall within the technical scope of the present disclosure. (1) An information processing device comprising: a control unit that estimates camera parameters of live-action video based on a first feature amount of a moving object present in a target space and a second feature amount of the moving object captured in the live-action video captured of the target space. (2) The information processing device described in (1), wherein the first feature amount includes information indicating a three-dimensional position and posture of each of one or more parts constituting the moving object. (3) The information processing device described in (2), wherein the control unit renders a 3D model of the moving object based on the first feature amount into a 2D image based on candidates for the camera parameter, and estimates the camera parameters of the live-action video based on a similarity between a pose of the moving object captured in the rendered 2D image and a pose of the moving object captured in a 2D image constituting the live-action video. (4) The information processing device according to (3), wherein the control unit continues calculating the similarity while changing the candidate camera parameters when the similarity is equal to or less than a predetermined threshold, and identifies the candidate camera parameters whose similarity exceeds the predetermined threshold as the camera parameters of the live-action video. (5) The information processing device according to (3) or (4), wherein the control unit sets a first search range including one or more candidate camera parameters based on metadata of the live-action video, and estimates the camera parameters from within the set first search range. (6) The information processing device according to (5), wherein the metadata includes at least one of information for specifying a shooting location of the live-action video, information for specifying a shooting time of the live-action video, or information usable to identify the moving object appearing in the live-action video. (7) The information processing device according to any one of (3) to (6), wherein the first feature amount is associated with identification information of the moving object, and the control unit calculates the similarity for moving objects associated with the same identification information. (8) The information processing device according to any one of (3) to (7), wherein the control unit renders a 3D model of the moving object corresponding to the first feature amount acquired at an acquisition time corresponding to a shooting time of the live-action video into a 2D image based on the candidate camera parameters, and calculates the similarity.(9) The information processing device according to (8), wherein the control unit sets a second search range including one or more candidates for the acquisition time based on a playback speed of the live-action video, and estimates the acquisition time from within the set second search range. (10) The information processing device according to any one of (1) to (9), wherein the control unit generates synchronized playback information that associates the live-action video, the first feature amount acquired at an acquisition time corresponding to a shooting time of the live-action video, and the camera parameters of the live-action video. (11) The information processing device according to (10), wherein the synchronized playback information includes a playback speed of the live-action video. (12) The information processing device according to any one of (1) to (11), wherein the camera parameters include at least one of a three-dimensional position, an attitude, and an angle of view of the camera. (13) A terminal device comprising: a control unit that generates a playback screen on which a reproduction video generated by rendering a 3D model, generated based on a first feature amount of a moving object existing in a target space, into a 2D image based on camera parameters of live-action video captured of the target space is played back in synchronization with the live-action video. (14) The terminal device described in (13), wherein the control unit generates the playback screen on which the live-action video and the reproduction video are displayed simultaneously, or on which either the live-action video or the reproduction video is switchably displayed. (15) The terminal device described in (13) or (14), wherein the control unit generates the playback screen on which the reproduction video is played back in synchronization with the live-action video based on synchronized playback information that associates the live-action video, the first feature amount acquired at an acquisition time corresponding to the capture time of the live-action video, and the camera parameters of the live-action video. (16) The terminal device described in (15), wherein the synchronized playback information includes a playback speed of the live-action video, and the playback screen includes the playback speed of the live-action video. (17) An information processing method including: estimating camera parameters of a live-action video captured of the target space based on a first feature amount of a moving object present in the target space and a second feature amount of the moving object captured in the live-action video.(18) An information processing method including: generating a playback screen in which a reproduced video generated by rendering a 3D model generated based on a first feature amount of a moving object existing in a target space into a 2D image based on camera parameters of live-action video captured of the target space is played back in synchronization with the live-action video. (19) A program for causing a computer to function as: a control unit that estimates camera parameters of the live-action video based on a first feature amount of a moving object existing in the target space and a second feature amount of the moving object appearing in the live-action video captured of the target space. (20) A program for causing a computer to function as: a control unit that generates a playback screen in which a reproduced video generated by rendering a 3D model generated based on the first feature amount of a moving object existing in the target space into a 2D image based on camera parameters of the live-action video captured of the target space is played back in synchronization with the live-action video.
[0196] 1 System 2 Server 3 User terminal
Claims
1. An information processing device comprising: a control unit that estimates camera parameters of live-action footage based on a first feature of a moving object present in a target space and a second feature of the moving object captured in the live-action footage of the target space.
2. The information processing device according to claim 1, wherein the first feature amount includes information indicating the three-dimensional position and orientation of each of one or more parts that make up the moving object.
3. The information processing device described in claim 2, wherein the control unit renders a 3D model of the moving object based on the first feature into a 2D image based on the candidate camera parameters, and estimates the camera parameters of the live-action video based on the similarity between the pose of the moving object shown in the rendered 2D image and the pose of the moving object shown in the 2D images constituting the live-action video.
4. The information processing device described in claim 3, wherein the control unit continues to calculate the similarity while changing the candidate camera parameters when the similarity is below a predetermined threshold, and identifies the candidate camera parameters whose similarity exceeds the predetermined threshold as the camera parameters of the live-action video.
5. The information processing device according to claim 3, wherein the control unit sets a first search range including one or more candidates for the camera parameters based on metadata of the live-action video, and estimates the camera parameters from within the set first search range.
6. An information processing device as described in claim 5, wherein the metadata includes at least one of information for identifying the location where the live-action video was shot, information for identifying the time when the live-action video was shot, or information that can be used to identify the moving object appearing in the live-action video.
7. The information processing device according to claim 3, wherein the first feature amount is associated with identification information of the moving object, and the control unit calculates the similarity for the moving objects associated with the same identification information.
8. The information processing device according to claim 3, wherein the control unit renders a 3D model of the moving object corresponding to the first feature acquired at an acquisition time corresponding to the shooting time of the live-action video into a 2D image based on the candidate camera parameters, and calculates the similarity.
9. The information processing device according to claim 8, wherein the control unit sets a second search range including one or more candidates for the acquisition time based on the playback speed of the live-action video, and estimates the acquisition time from within the set second search range.
10. The information processing device of claim 1, wherein the control unit generates synchronized playback information that associates the live-action video, the first feature acquired at an acquisition time corresponding to the shooting time of the live-action video, and the camera parameters of the live-action video.
11. The information processing device according to claim 10, wherein the synchronized playback information includes a playback speed of the live-action video.
12. The information processing device according to claim 1, wherein the camera parameters include at least one of a three-dimensional position, an orientation, and an angle of view of the camera.
13. A terminal device comprising: a control unit that generates a playback screen in which a reproduced image is generated by rendering a 3D model generated based on a first feature amount of a moving object existing in a target space into a 2D image based on camera parameters of live-action video taken of the target space, and plays the reproduced image in synchronization with the live-action video.
14. The terminal device according to claim 13, wherein the control unit generates the playback screen that simultaneously displays the live-action video and the reproduced video, or that switchably displays either the live-action video or the reproduced video.
15. The terminal device described in claim 13, wherein the control unit generates the playback screen in which the reproduced video is played back in synchronization with the live-action video based on synchronized playback information that associates the live-action video with the first feature acquired at an acquisition time corresponding to the shooting time of the live-action video and the camera parameters of the live-action video.
16. The terminal device according to claim 15, wherein the synchronized playback information includes a playback speed of the live-action video, and the playback screen includes a playback speed of the live-action video.
17. An information processing method including: estimating camera parameters of live-action footage based on a first feature amount of a moving object present in a target space and a second feature amount of the moving object captured in the live-action footage of the target space.
18. An information processing method including: generating a playback screen in which a reproduced image is generated by rendering a 3D model generated based on a first feature amount of a moving object existing in a target space into a 2D image based on camera parameters of live-action video of the target space, and playing the reproduced image in synchronization with the live-action video.
19. A program for causing a computer to function as a control unit that estimates camera parameters of live-action footage based on a first feature of a moving object present in a target space and a second feature of the moving object captured in the live-action footage of the target space.
20. A program for causing a computer to function as a control unit that generates a playback screen in which a reproduced image is generated by rendering a 3D model based on a first feature amount of a moving object existing in a target space into a 2D image based on camera parameters of live-action footage of the target space, and plays the reproduced image in synchronization with the live-action footage.
Citation Information
Patent Citations
Image processing device, image processing system, and image processing method
JP2013182523A
Camera calibration device, camera calibration method and program
JP2024021218A
Information processing device, information processing method, and program
WO2023210187A1