3D Video Highlights from Camera Source
The system efficiently converts 2D video content into 3D format by applying 3D motion data to animated objects and using a streaming engine, reducing resource demands and enabling real-time 3D rendering on diverse displays.
Patent Information
- Application Number
- JP2024577277
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-01
- Filing Date
- 2023-06-30
- Publication Date
- 2025-07-17
AI Technical Summary
Conventional 3D production is time-consuming and resource-intensive, requiring significant computing resources and dedicated GPU hardware, making it inefficient for real-time conversion of 2D video content into immersive 3D formats.
A system that generates 3D video segments from 2D video content using a video manager and 3D pose estimation engine, applying 3D motion data to animated objects, and utilizes a streaming engine for real-time rendering without the need for dedicated GPU hardware, enabling 3D video content display on various devices.
Enables fast and accurate conversion of 2D video content into 3D format, reducing computing resource requirements and allowing real-time rendering of 3D video content on both 2D and 3D displays, enhancing user interaction and customization.
Smart Images

Figure 2025522846000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 367,525, filed on July 1, 2022, entitled "THREE DIMENSIONAL(3D) HIGHLIGHTS", the disclosure of which is hereby incorporated by reference in its entirety.
Background Art
[0002] 3D production can be a process of creating and producing content in three dimensions (e.g., height, width, and depth), and can include various techniques and technologies for capturing, generating, manipulating, and / or presenting content in a 3D format. In 3D production, content can be created or captured in a way that simulates depth perception, providing a more immersive and realistic experience. Conventional 3D production can be time - consuming and can involve a number of technical processes that may use relatively large amounts of computing resources (e.g., memory, the capabilities of a central processing unit (CPU) and / or a graphics processing unit (GPU)).
Summary of the Invention
[0003] In some aspects, the technology described herein relates to a method that includes generating a three - dimensional (3D) video segment from two - dimensional (2D) video content captured by a camera system. Generating includes obtaining 3D motion data of an object detected within the 2D video content from a 3D pose estimation engine, and generating an animated object based on the 3D motion data such that the movement of the animated object corresponds to the movement of the object within the 2D video content. The method further includes generating 3D video content from the 3D video segment, and transmitting the 3D video content to a user device for display.
[0004] In some aspects, the techniques described herein relate to a non-transitory computer-readable medium storing executable instructions that cause at least one processor to perform operations including transmitting, via a network, a three-dimensional (3D) viewing request for two-dimensional (2D) video content captured by a camera system in 3D format to a video manager executable by at least one server computer, the 3D viewing request being a request to view the 2D video content, and further including receiving, via the network, 3D video content from the video manager, the 3D video content being generated from 3D video segments generated by the video manager using the 2D video content, and further including starting to display the 3D video content on an interface of a user device, the 3D video content including animated objects, the movement of the animated objects corresponding to the movement of objects in the 2D video content.
[0005] In some aspects, the techniques described herein relate to an apparatus including at least one processor and a non-transitory computer-readable medium storing executable instructions that cause the at least one processor to receive, from a user device, a three-dimensional (3D) viewing request to view two-dimensional (2D) video content in 3D format, retrieve, from a video database, a 3D video segment corresponding to the 2D video content, the 3D video segment including animated objects, the movement of the animated objects corresponding to the movement of objects in the 2D video content, and further cause the at least one processor to generate 3D video content from the 3D video segment and transmit the 3D video content to the user device for display.
[0006] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0007]
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 1E
Figure 1F
Figure 1G
Figure 1H
Figure 1I
Figure 1J
Figure 1K
Figure 1L
Figure 1M
Figure 1N
Figure 2A
Figure 2B
Figure 2C
Figure 2D
Figure 2E
Figure 2F
Figure 2G
Figure 2H
Figure 2I
Figure 2J
Figure 2K
Figure 2L
Figure 2M
Figure 3A
Figure 3B
Figure 4A
Figure 4B
Figure 4C
Figure 5A
Figure 5B
Figure 5C
Figure 6
Figure 7
Figure 8
[0008] The present disclosure can generate a three-dimensional (3D) video segment (e.g., 3D model data) from two-dimensional (2D) video content (e.g., sports video), and render 3D video content that is transmitted to and displayed on an interface of an application executable by a user device in substantially real time. The 3D video content includes images that replay a scene captured in the 2D video content (e.g., a scene where a basketball player dunks a basketball, a scene where a football team scores, etc.) in a 3D format. In some examples, the 3D video content includes virtual reality (VR) content configured to be displayed on a 3D display (e.g., a VR wearable device). In some examples, the 3D video content includes augmented reality (AR) content configured to be displayed on a 2D display. The system includes a video manager that executes a technical process to increase the speed and / or accuracy of 3D production while reducing the amount of computing resources (e.g., memory, central processing unit (CPU) power, graphics processing unit (GPU) power, etc.) to enable the display of 3D video content on a 2D or 3D display.
[0009] For example, a video manager can quickly and accurately generate 3D video segments (e.g., 3D video highlights) from 2D video content by obtaining 3D motion data of objects detected within the 2D video content from a 3D pose estimation engine. The 3D motion data includes the position and orientation of the objects in 3D space, which captures the movement of the objects in 3D space. The video manager generates animated objects based on the 3D motion data, and the movement of the animated objects corresponds to the movement of the objects within the 2D video content. For example, realistic motions captured by the 3D motion data are applied to computer-generated objects (e.g., 3D object models), thereby converting static models into dynamic models that move in the same / similar manner as the objects within the 2D video content. The system can enable real-time rendering of 3D video segments on a user device by using a streaming engine (e.g., a cloud-based 3D pixel streaming engine) to generate (e.g., render) 3D video content from the 3D video segments, and can regenerate (e.g., re-render) the 3D video content according to any user selection. Thereby, the need for the user device to have a dedicated GPU device can be minimized, reduced, or eliminated.
[0010] The video manager includes a 3D video generator configured to generate 3D video segments (e.g., 3D model data) from 2D video content. The 2D video content can be any type of 2D video content captured by a camera system. In some examples, the 2D video content is a video file or a portion of the video content from a video file. In some examples, the 2D video content includes video highlights from a sports event. However, note that the systems described herein are not limited to sports highlights, and the techniques described herein can be used for any type of 2D video content captured from a camera system.
[0011] In some examples, a user can view 2D video content using their user device and initiate a 3D viewing request to view the 2D video content in 3D format. For example, the interface can include UI elements that, when selected, cause a 3D viewing request to be sent to the video manager. In response to the 3D viewing request, the 3D video generator can generate a 3D video segment (e.g., generate the 3D video segment on-the-fly or near real-time). In some examples, in response to the 3D viewing request, the 3D video generator can retrieve a 3D video segment (e.g., a previously generated 3D video segment) from a video database.
[0012] In some examples, the 2D video content may include television content (e.g., live television content), and the 3D video generator may generate one or more 3D video segments from one or more 2D video segments while the live television content is being broadcast. In some examples, the 3D video generator may identify that the 2D video content includes a key event (e.g., a key play such as a player scoring), select the 2D video segment that includes the key event, and generate a 3D video segment from the 2D video segment. In some examples, during the live broadcast, the 3D video generator may receive 2D video segments from an external service, and the external service may include logic for selecting 2D video segments that include key events.
[0013] The 3D video generator may use a machine learning (ML) model(s) to obtain 3D motion data (e.g., 3D pose data, 3D body position data) of related object(s) detected within the 2D video content. In some examples, the 3D video generator may obtain the 3D motion data from an existing (known) ML pose estimation model (e.g., a 3D human pose estimation model and / or a 3D object (e.g., ball) tracking model). The 3D motion data may include position information and rotation information that describe the movement and orientation of an object within 3D space. In some examples, the 3D video generator includes an ML human pose estimation model configured to detect the movement and orientation of the pose of a human object (e.g., represented by key points such as ankles, shoulders, neck, hands, etc.) within 3D space. In some examples, the 3D video generator includes an object detection and tracking model configured to detect the movement and orientation of a non-human object (e.g., a ball) within 2D or 3D space. In some examples, the object detection and tracking model may provide the 2D position of the non-human object. In some examples, the object detection and tracking model may provide the height from the ground for each frame with respect to a player.
[0014] A 3D video generator can generate animated objects using 3D motion data and include the animated objects in a 3D video segment. In some examples, the 3D video generator can generate an animated object by applying 3D motion data to a 3D object model, which, in some examples, can represent an object. In some examples, the 3D video generator can select an existing 3D object model from an object model database (e.g., a model inventory) corresponding to the detected object, or generate the 3D object model itself using 2D video content with one or known mesh generation techniques. The actions performed by the 3D video generator can enable the animated objects represented in the 3D video content to have smooth and realistic movements that reflect the movement of the objects in the 2D video content. For example, if the object is a basketball player, the 3D video generator can generate 3D motion data regarding the actions of the player in 3D space from 2D video content. In some examples, the 3D generator can obtain (or generate) a 3D model object representing a basketball player (which can be a specific known basketball player or a general basketball player), apply the 3D motion data to the 3D object model to generate an animated object, and the 3D object model is animated according to the movement of the player.
[0015] The system may include a streaming engine configured to generate 3D video content from 3D video segments. In some examples, the streaming engine is a 3D pixel streaming engine. The streaming engine may enable near real-time streaming and rendering of 3D graphics and interactive content over a network, which may include one or more rendering operations such as geometry processing, shading and material calculations, camera and viewpoint calculations, frame composition, video encoding, and / or network transmission. For example, the streaming engine includes one or more GPUs configured to perform rendering operations on 3D video segments to generate images (e.g., video frames) of 3D video content. The streaming engine may transmit the 3D video content to a user device for display. In some examples, since the rendering operations are performed by the streaming engine, 3D video content can be sent to the user without requiring a user device with special GPU hardware. The streaming engine may receive information indicating one or more user selections for user control to adjust and / or customize the playback of 3D video segments, and may regenerate the 3D video content from the 3D video segments according to the user selection(s).
[0016] The interface may include one or more user controls that enable a user to modify the playback of a 3D video segment. Examples of user controls include adjusting the viewing angle (e.g., selecting a quarter-back point of view (POV), selecting the POV of the ball, etc.), adding virtual content (e.g., animation effects, statistical data, graphics, etc.), adjusting the display speed (e.g., slowing down or speeding up), and modifying animated objects (e.g., selecting a different model, modifying more features of an object model). In some examples, a user may use the user controls to create customized 3D video content and use sharing controls to share the customized 3D video content with one or more other devices. In some examples, the interface includes a search field that receives a search query from the user, and the interface may display search results that identify one or more 3D video segments in response to the search query. For example, a user may retrieve 3D highlights of a particular game, player, team, type of movement, etc. These features and other features are further described with reference to the figures.
[0017] Figures 1A - 1N illustrate a system 100 for generating a 3D video segment 122 and delivering it to a user device 152, according to an aspect. The system 100 includes a video manager 102 configured to generate the 3D video segment 122 from 2D video content 134 using one or more machine learning models 105, thereby enabling the 3D video content 121 generated from the 3D video segment 122 to be displayed on an interface 140 on a display 123 of the user device 152. The video manager 102 can implement technical features that can increase the speed and / or accuracy of 3D production while reducing the amount of computing resources (e.g., memory, central processing unit (CPU) resources, graphics processing unit (GPU) resources) to enable the display of the 3D video segment 122 on the display 123, which can be a 2D display or a 3D display.
[0018] For example, the video manager 102 can quickly and accurately generate a 3D video segment 122 (e.g., a 3D video highlight) from the 2D video content 134 by obtaining the 3D motion data 110 of the object 108 detected within the 2D video content 134 from the 3D pose estimation engine 188. The 3D motion data 110 includes the position and orientation of the object in 3D space, which captures the movement of the object 108 in 3D space. The video manager 102 generates an animated object 114 based on the 3D motion data 110, and the movement of the animated object 114 corresponds to the movement of the object 108 within the 2D video content 134. For example, the realistic motion captured by the 3D motion data 110 is applied to a computer-generated object (e.g., the 3D object model 116), thereby converting the static model into a dynamic model that moves in the same / similar manner as the object 108 within the 2D video content 134. The system 100 can enable real-time rendering of the 3D video segment on the user device 152 by generating (e.g., rendering) the 3D video content 121 from the 3D video segment 122 and using a streaming engine 120 (e.g., a cloud-based 3D pixel streaming engine) to regenerate (e.g., re-render) the 3D video content according to any user selection 171. This can minimize, reduce, or eliminate the need for the user device 152 to have a dedicated GPU device.
[0019] The 3D video segment 122 includes 3D model data 125 that includes one or more animated objects 114 that move in a 3D scene in a manner corresponding to the movement(s) of the object(s) 108 within the 2D scene of the 2D video content 134. In some examples, the object 108 can be the structure of the object 108 (e.g., a human body, or a non-human object such as a ball or another movable object). The 3D video segment 122 can include one or more static structures and / or other model data that defines the environment of the 3D scene. In some examples, one or more machine learning models 105 (e.g., the 3D pose estimation engine 188) can detect 3D motion data 110 regarding the movement and orientation of the object 108 in 3D space from the 2D position of the object 108 in 3D space in the frames of the 2D video content 134. The video manager 102 can generate the animated object 114 based on the 3D motion data 110. The animated object 114 can move in 3D space in a manner corresponding to the movement of the object 108 within the 2D video content 134. In some examples, the video manager 102 can apply the 3D motion data 110 to the 3D object model 116, and the 3D object model 116 creates the animated object 114 that moves within the 3D scene in a manner corresponding to the movement of the object 108 in the 2D scene.
[0020] For example, video manager 102 may detect 3D motion data 110 regarding the motion of object 108 from 2D video content 134. In some examples, 3D motion data 110 is 3D pose data over time. 3D motion data 110 includes position information and rotation information that describe the movement and orientation of object 108 in 3D space. In some examples, 3D motion data 110 includes the 3D position coordinates (e.g., x value, y value, and z value) and rotation information (e.g., Euler angles, quaternions, and / or rotation matrices) of one or more keypoints 170 on object 108 over a period of time. Keypoints 170 may represent different parts of the structure of object 108, and 3D motion data 110 may capture the movement and orientation of keypoints 170 over time. Video manager 102 may generate animated object 114 by applying 3D motion data 110 to 3D object model 116 representing object 108. In some examples, 3D object model 116 is selected from object model database 180, and 3D motion data 110 is applied to the selected 3D object model 116. In some examples, 3D object model 116 is generated from 2D video content 134 using one or more known mesh generation techniques. Animated object 114 may move within the 3D scene in a manner corresponding to the movement of object 108 within the 2D scene.
[0021] The 3D video segment 122 can be a 3D video highlight (e.g., a 3D sports highlight). However, the techniques described herein can be applied to any type of underlying video content having one or more objects 108 moving within 2D video content 134. The video manager 102 can convert a 2D video highlight into a 3D video highlight, thereby enabling the user to replay the video highlight in 3D format and view the highlight from different angles and / or speeds, and / or change the video aspect of the 3D video segment 122. Such changes can include, for example, adding graphics, statistical data, and / or animation effects to the highlight, customizing animated object(s) 114 (e.g., changing a player's clothing, style, or other feature), and / or customizing another part of the 3D scene. For example, the 3D video segment 122 can be visually explored by the user based on user selections 171 received via one or more user controls 142 on the interface 140. The user can view the highlight from different camera viewpoints and / or speeds and can also create customized 3D video content 121a (resulting in graphics and / or animation effects) that can be shared with another user.
[0022] The video manager 102 includes a 3D video generator 104 configured to receive 2D video content 134 and generate 3D video segments 122 based on the 2D video content 134. The 2D video content 134 includes video data (and in some examples, audio data) captured from the camera system 132. The camera system 132 can include one or more camera devices configured to capture video in two dimensions, representing a scene as a flat image or a sequence with height and width. Each frame of the 2D video content 134 includes a 2D array of pixels, and each pixel includes color and luminance information. In some examples, the camera system 132 includes an audio system having one or more microphones configured to capture audio data. In some examples, the camera system 132 does not include dedicated capture equipment (such as a stereo camera or depth sensing technology) for obtaining additional stereoscopic information for a 3D experience. In some examples, the 2D video content 134 includes television content 157 (such as live television content).
[0023] The 2D video content 134 is displayed on the display 123 of the user device 152. In some examples, the display 123 is a 2D display. In some examples, the display 123 is a 3D display, and the 2D video content 134 is displayed on the 3D display in a 2D format. In some examples, while the camera system 132 is generating the 2D video content 134, the 2D video content 134 is streamed to the user device 152 via a media platform (such as a streaming platform). In some examples, the 2D video content 134 is stored on a remote server computer and streamed from the remote server computer to the user device 152. In some examples, the 2D video content 134 is stored on the user device 152.
[0024] The application 146 executable by the user device 152 may receive the 2D video content 134 and display the 2D video content 134 on the display 123. The application 146 may be a native application executable by the operating system 186 of the user device 152. In some examples, the application 146 is a streaming application. In some examples, the application 146 is a video sharing application. In some examples, the application 146 is a browser application. The browser application may render (or execute the application) a web page within a browser tab that streams the 2D video content 134 to the user device 152 from a streaming platform (also referred to as a media platform or media provider).
[0025] The 3D video generator 104 may obtain the 2D video content 134 from a streaming platform that distributes and / or stores the 2D video content 134 captured from the camera system 132. In some examples, while the camera system 132 is generating the 2D video content 134, the 3D video generator 104 may obtain the 2D video content 134 from the streaming platform. In some examples, the 3D video generator 104 may receive the 2D video content 134 from a remote server computer that stores the 2D video content 134. In some examples, the 2D video content 134 is associated with a resource identifier (e.g., a uniform resource locator (URL)) that identifies the location of the 2D video content 134. In some examples, the 3D video generator 104 may retrieve the 2D video content 134 using the resource identifier. In some examples, the 3D video generator 104 may receive the resource identifier and / or the 2D video content 134 from the user device 152.
[0026] In some examples, a user may view 2D video content 134 within interface 140 of application 146. Interface 140 may include UI element 107 configured to enable the user to view 2D video content 134 in a 3D format. In some examples, UI element 107 is a UI control that enables the user to view 2D highlights in a 3D format. In some examples, selection of UI element 107 causes a 3D viewing request 184 to be generated and sent to video manager 102. 3D viewing request 184 may include information identifying 2D video content 134. In some examples, 3D viewing request 184 includes a resource identifier of 2D video content 134. In some examples, 3D viewing request 184 may include 2D video content 134.
[0027] In some examples, application 146 may generate and send 3D viewing request 184. In some examples, user device 152 includes a 3D client manager 155 configured to generate and send 3D viewing request 184. In some examples, 3D client manager 155 is a program separate from application 146. In some examples, 3D client manager 155 is an operating system program. In some examples, 3D client manager 155 is a component (e.g., a sub-component) of a browser application. 3D client manager 155 and application 146 may communicate with each other via an application programming interface (API) or an inter-process communication (IPC) link. In some examples, 3D client manager 155 is configured to receive an indication of user selection 171 to user control 142, thereby causing 3D client manager 155 to generate and send 3D viewing request 184. In some examples, 3D client manager 155 is included as part of application 146.
[0028] In response to the 3D viewing request 184, the 3D video generator 104 obtains the 2D video content 134 and generates the 3D video segment 122. In some examples, the 3D video generator 104 may generate a single 3D video segment 122 from the 2D video content 134. In some examples, the 3D video generator 104 may generate multiple 3D video segments 122 from the 2D video content 134. In some examples, the 3D video generator 104 may retrieve the 2D video content 134 using the resource identifier of the 2D video content 134. In some examples, the 3D video generator 104 may obtain the 2D video content 134 from the 3D viewing request 184. In some examples, the video manager 102 may use the information from the 3D viewing request 184 to determine whether the 3D video segment 122 has already been generated for the 2D video content 134. For example, the video manager 102 may execute a query against a video database 196 that stores multiple 3D video segments 122 (e.g., performed using the resource identifier, other identifiers that can uniquely identify the 2D video content 134, and / or other information included in the 3D viewing request 184). If the video manager 102 determines that the 3D video segment 122 has not been generated from the 2D video content 134, the video manager 102 may cause the 3D video generator 104 to generate the 3D video segment 122.
[0029] In some examples, the 3D video generator 104 can generate one or more 3D video segments 122 from 2D video content 134 without prompting from the user. As shown in FIG. 1B, the video manager 102 can include a segment selector 128 configured to identify one or more 2D video segments 134a from 2D video content 134. The 2D video segments 134a can be examples of 2D video content 134. In some examples, the 2D video segments 134a are shorter video clips (e.g., highlights) from longer video content (e.g., 2D video content 134). The 2D video content 134 can include television content 157. The television content 157 can be a program associated with a television channel. The program can be a sports program. In some examples, the media platform (or media provider) is configured to stream the television content 157 via the network 150 and / or is configured to broadcast the television content 157 using radio waves. In some examples, the television content 157 includes live television content. The live television content can be digital data that is streamed or broadcast when the 2D video content 134 is captured by the camera system 132.
[0030] The segment selector 128 detects that a portion of the TV content 157 includes the key event 127 and identifies the 2D video segment 134a from a part of the TV content 157. For example, when the TV content 157 is captured by the camera system 132 and streamed / broadcast to the viewer, the segment selector 128 may receive the TV content 157. The segment selector 128 may detect whether the 2D video content 134 includes the key event 127, and if it does, may identify the 2D video segment 134a. The key event 127 may be a specific sports action or event (e.g., a goal, a foul, a pass, etc.). In response to detecting the key event 127 within the 2D video content 134, the segment selector 128 may determine the start and end of the scene and select that portion for the 2D video segment 134a. The 2D video segment 134a may be a clip or a portion of the TV content 157 that includes a highlight (e.g., the key event 127). The 2D video segment 134a may include a key play from a sports event. The 2D video segment 134a may be a video clip of a basketball player dunking a basketball or a football player scoring a goal.
[0031] The segment selector 128 executes one or more existing event detection algorithms to analyze the video footage of the TV content 157 to determine whether the 2D video content 134 represents the key event 127, and if it does, may identify the 2D video segment 134a as a highlight of the 2D video content 134.
[0032] The segment selector 128 may include an event recognition model (e.g., a convolutional neural network (CNN) or a recurrent neural network (RNN)) configured (e.g., trained) to recognize key events 127 within one or more types of sports events. In the case of basketball, the segment selector 128 detects a key event 127 when a player scores a point in basketball, performs a specific action such as dunking the basketball, making an assist, stealing the ball, etc. In the case of football, the segment selector 128 may detect a key event 127 such as a player who catches the ball beyond a threshold number of yards, a player who scores a touchdown, etc. In some examples, the segment selector 128 includes a motion analysis model configured to analyze the motion pattern of a detected object (e.g., object 108 detected by an object tracker(s) 106) to detect a key event 127 (e.g., a key movement) within the 2D video content 134. Techniques such as optical flow that estimate the motion of pixels between consecutive frames may be used to calculate the magnitude and direction of the motion of each object 108. A change in motion, a sudden acceleration, or a specific motion pattern may indicate a key event 127. In some examples, the segment selector 128 includes a rule-based model that analyzes the 2D video content 134 and applies specific rules or criteria to detect a key event 127. For example, in basketball, detecting a slam dunk or a three-point shot may be based on height, the trajectory of the ball, and the shooting location. By defining rules or thresholds based on the rules of the sport, the segment selector 128 may detect a key event 127 within the 2D video content 134.
[0033] As shown in FIG. 1B, the segment selector 128 can identify the 2D video segments 134a-1 and 134a-2 from the 2D video content 134. Next, the 3D video generator 104 can generate a 3D video segment 122-1 corresponding to the 2D video segment 134a-1 and a 3D video segment 122-2 corresponding to the 2D video segment 134a-2. Two 3D video segments 122 are shown in FIG. 1G, but the 3D video generator 104 can generate any number of 3D video segments 122 from the TV content 157, which may depend on the number of key events 127 detected within the 2D video content 134.
[0034] Referring again to FIG. 1A, the 3D video generator 104 can include one or more object trackers 106 configured to detect one or more types of objects 108 within the 2D video content 134 and generate 3D motion data 110 for the detected object(s) 108 within the 2D video content 134. In some examples, the object tracker(s) 106 can include separate object trackers 106 configured to detect a particular type (or classification) of object 108 and generate 3D motion data 110 for that type of object 108, where the 3D motion data 110 can vary between different types of objects 108. In some examples, a single object tracker 106 can detect multiple different types of objects 108 and generate 3D motion data 110 for the different types of objects 108. In some examples, a first object tracker is configured to generate 3D motion data 110 for one or more types of objects 108 (e.g., players), and a second object tracker is configured to generate 2D motion data for one or more types of objects 108 (e.g., balls).
[0035] In some examples, object tracker 106 includes a 3D pose estimation engine 188 for detecting human or non-human poses. In some examples, 3D pose estimation engine 188 includes an existing 3D pose estimation model (e.g., an embodied model, a PyMAF model). 3D motion data 110 may include position information and rotation information that describe the movement and orientation of object 108 in 3D space. In some examples, 3D pose estimation engine 188 includes an ML human pose estimation model configured to detect the movement and orientation of the poses of human objects (e.g., represented by key points 170 such as ankles, shoulders, necks, hands, etc.) in 3D space. In some examples, object tracker 106 includes an object detection and tracking model configured to detect the movement and orientation of non-human objects (e.g., balls) in 2D or 3D space. In some examples, the object detection and tracking model may provide the 2D position of the non-human object. In some examples, the object detection and tracking model may provide the height above the ground for each frame with respect to the player.
[0036] As shown in FIG. 1C, 2D video content 134 may visually represent objects 108-1 and 108-2 within the frame of 2D video content 134. In some examples, object 108-1 is a person. In some examples, object 108-2 is a non-human object such as a ball. Although FIG. 1 depicts two objects 108 within 2D video content 134, object tracker(s) 106 may detect and generate 3D motion data 110 for any number (or type) of objects 108 within 2D video content 134 that includes one object 108 or any number of objects 108 greater than two. Objects 108 may represent various physical objects such as players, balls, and / or other physical elements represented in 2D video content 134.
[0037] In some examples, the object tracker 106 can detect that the object 108-1 is an object 108 of a first type (e.g., the object 108-1 has a human body) and can generate 3D motion data 110 from the 3D video content 134. In some examples, the object tracker 106 can detect whether a person is a player or a non-player (e.g., a coach or a referee, training staff, etc.), whether the person is a specific type of player (e.g., a guard, a center, a lineman, a receiver, a quarterback), and / or whether the person is a known entity (e.g., a specific person such as Khris Middleton, Jrue Holiday, Jordan Love, etc.), and can detect further subtypes of the first type. In some examples, the object tracker 106 (or, in some examples, another object tracker) can detect that the object 108-2 is an object 108 of a second type (e.g., a ball) and can generate 3D motion data 110 from the 2D video content 134. In some examples, the object tracker 106 (or another object tracker) can detect another type of object 108 and can generate 3D motion data 110 from the 2D video content 134.
[0038] Object tracker 106 includes a 3D pose estimation engine 188 configured to generate 3D motion data 110 for objects 108 (e.g., object 108-1, object 108-2) detected within 2D video content 134. The 3D pose estimation engine 188 may include one or more machine learning (ML) models 105. In some examples, the 3D motion data 110 is the 3D pose of the object 108 over time. The 3D motion data 110 includes position information and rotation information that describe the movement and orientation of the object 108 within 3D space. In some examples, the 3D motion data 110 includes the 3D position coordinates (e.g., x value, y value, and z value) and rotation information (e.g., Euler angles, quaternions, and / or rotation matrices) of at least one of a plurality of keypoints 170 over a period of time. The keypoints 170 may represent different parts of the structure of the object 108.
[0039] As shown in FIG. 1D, the object 108 may be defined by one or a set of keypoints 170, and the 3D position and orientation of the object are estimated by the 3D pose estimation engine 188. The keypoints 170 may be associated with different parts of the object 108 (e.g., in the case of a human body, the parts may be ankles, waist, head, shoulders, elbows, etc.). The keypoints 170 may include keypoints from keypoint 170-1, keypoint 170-2, and keypoint 170-3 to keypoint 170-N. It should be noted that the 3D pose estimation engine 188 may track a single keypoint 170 (e.g., a center point on the object 108), or multiple keypoints 170 such as any number two or more.
[0040] The 3D pose estimation engine 188 generates 3D motion data 110 by estimating the 3D locations of the keypoints 170 from the 2D locations of the keypoints 170 in the 2D video content 134. The keypoints 170 may include parts that form a pose. In some examples, in the case of a human body, the keypoints 170 may include the head, neck, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right hip, left hip, right knee, left knee, right ankle, and / or left ankle. In some examples, the keypoints 170 may include the nose, left eye, and / or right eye. The 3D location of the keypoint 170 may refer to the 3D spatial position (e.g., position coordinates (e.g., X, Y, Z values)) of the keypoint 170 in 3D space, and in some examples, may refer to rotation information regarding the orientation of the keypoint 170 in 3D space. In some examples, the rotation information may include Euler angles, quaternions, and / or rotation matrices.
[0041] In some examples, the 3D pose estimation engine 188 estimates the 2D pose of an object 108 (e.g., a person) in the input data (e.g., the 2D video content 134). 2D pose estimation involves the detection and localization of the keypoints 170 (e.g., major body joints) within each frame of the input data. This can be achieved using various techniques such as convolutional neural network (CNN) or pose estimation algorithms based on graphical models. In some examples, the 3D pose estimation engine 188 may estimate the depth of the keypoints 170 using the 2D pose, which may include triangulating the 2D positions of the keypoints 170 from multiple camera views or utilizing a depth map to infer depth information. Once the 2D positions and stereo information (if available) are obtained, the 3D pose estimation engine 188 may perform 3D pose estimation. There are different approaches to 3D pose estimation, including model-based methods, direct regression methods, and learning-based methods.
[0042] In some examples, the 3D motion data 110 generated by the 3D pose estimation engine 188 may not include the position(s) of the object(s) 108 relative to the static structure within the 2D video content 134 (e.g., the position of a player on a court or field). As shown in FIG. 1E, the object tracker 106 may include a relative position detector 172, which is configured to generate relative position information 129 regarding the position of the object 108 (e.g., one or more keypoints 170 of the object 108) with respect to one or more static structures 108a representing physical structures (e.g., a court, a field, a target, etc.) within the 2D video content 134. In some examples, the relative position detector 172 controls one or more aspects of the 3D pose estimation engine 188 to calculate the 3D motion data 110 once or multiple times when one or more keypoints 170 of the object 108 contact a part(s) of the static structure(s) 108a, or to select a part(s) of the 3D motion data 110 once or multiple times when one or more keypoints 170 of the object 108 contact a part(s) of the static structure(s) 108a.
[0043] The static structure 108a represents a physical space having a surface defined by a width (direction A1) and a length (direction A2), and in some examples, represents a height defined in a direction orthogonal to the surface (direction A3 represented as dots extending in and out of the page). The relative position information 129 may include the 3D position of the object 108-1 relative to another object, i.e., the static structure 108a.
[0044] The relative position detector 172 can detect the static structure 108a from the 2D video content 134 and, over time, calculate the position (e.g., 3D position) of the object 108 relative to the static structure 108a from the 2D video content 134. In some examples, the relative position detector 172 can use inverse kinematics (IK) techniques to calculate the position of the object relative to the static structure 108a from the video frames of the 2D video content 134. In some examples, the relative position detector 172 can use IK techniques to calculate the position of the object relative to the static structure 108a at one or more key events (e.g., the position of each foot on the court when a player's foot touches the ground within the 2D video content 134, the time and location of the start and / or end of any jump when the ball is passed and released). In some examples, the relative position detector 172 can use a camera calibration tool and camera parameters from the camera system 132 to calculate the position of the object 108 relative to the static structure 108a from the video frames of the 2D video content 134.
[0045] Referring again to FIG. 1A, the 3D video generator 104 includes a 3D object engine 112 configured to generate animated objects 114 using 3D motion data 110. In some examples, the 3D object engine 112 can apply the 3D motion data 110 to a 3D object model 116 that represents a 3D version of the object 108. In some examples, the 3D object model 116 does not represent the object 108 but represents the user of the user device 152. In some examples, the 3D object model 116 can be any type of computer-generated object configured to be enhanced by the 3D motion data 110. In some examples, if the object 108 is a basketball player, the 3D object model 116 is a model representation of a basketball player.
[0046] The 3D object model 116 may include information regarding the geometry, topology, and appearance of the 3D object. The 3D object model 116 may define the shape and structure of the 3D object and may include information regarding vertices, edges, and / or faces that form the surface of the object. The 3D object model 116 may include information regarding the connections and relationships between the geometric elements of the model and may define how vertices, edges, and surfaces are connected to form the structure of the object. The 3D object model 116 may include texture coordinates that define how a texture or image is mapped onto the surface of the model and may provide a correspondence between points on the 3D surface and pixels within a 2D texture image. In some examples, the 3D object model 116 may include information regarding normals (e.g., vectors perpendicular to the surface at each vertex or face) that determine the orientation and direction of the surface and indicate how light interacts with the surface during shading calculations. The 3D object model 116 may include information regarding material properties that describe the visual appearance and characteristics of the surface of the 3D object and may include information such as color, reflectivity, transparency, gloss, and other parameters that affect how the surface interacts with light.
[0047] In some examples, the 3D object model 116 is initially configured as a static model. However, when the 3D object engine 112 applies 3D motion data 110 from the object tracker(s) 106, the static model is converted into an animated object 114, thereby generating a dynamic object. In some examples, the animated object 114 may be referred to as an animated mesh or an animated rig. Applying the 3D motion data 110 to the 3D object model 116 may include adding the 3D motion data 110 to the 3D object model 116 to generate the animated object 114, and the 3D object model 116 is configured to move in the manner indicated by the 3D motion data 110.
[0048] In other words, as shown in FIG. 1F, the animated object 114 is defined by a 3D object model 116, which is augmented with 3D motion data 110 generated by a 3D pose estimation engine 188. The animated object 114 is configured to have a 3D pose in a video frame of 3D video content 121 corresponding to the pose in a video frame of 2D video content 134, thereby playing back the movement of the object in 3D. The animated object 114 may also include any of the information described with reference to the 3D object model 116.
[0049] In some examples, the 3D object engine 112 selects a 3D object model 116 that represents an object 108 detected in the 2D video content 134. As shown in FIG. 1G, the 3D object engine 112 may include an object model selector 178. For an object 108-1 detected within the 2D video content 134, the object model selector 178 may select a 3D object model 116-1 that represents the object 108 from an object model database 180 that stores a plurality of 3D object models 116. The object model database 180 may be referred to as a 3D object model inventory. The 3D object models 116 stored in the object model database 180 may be different versions of the 3D model. The 3D object models 116 stored in the object model database 180 may include a 3D object model 116-1 and a 3D object model 116-2. The 3D object model 116-2 may have at least one feature different from the 3D object model 116-1. However, the 3D object models 116 stored in the object model database 180 may include any number of 3D models (e.g., ranging from dozens to hundreds, thousands). In some examples, the object model selector 178 may select one of the 3D object models 116 stored in the object model database 180 based on the object data 174 associated with the object 108-1.
[0050] In some examples, the object data 174 includes information detected by the object tracker 106, such as a classification type, e.g., whether the object 108 is a player or non-player, a particular type of player, and / or a particular known player, etc. In some examples, when the object data 174 indicates that the object 108 is a player, the object model selector 178 may obtain an object model 116-1 representing a general player. In some examples, when the object data 174 indicates that the object 108 is a particular type of player (e.g., a receiver, running back, center, or guard), the object model selector 178 may obtain an object model 116-1 representing the particular type of player. In some examples, when the object data 174 indicates that the object 108 is a particular known player (e.g., Kris Middleton), the object model selector 178 may obtain an object model 116-1 representing that particular known player.
[0051] In some examples, the 3D object engine 112 generates a 3D object model 116 using the video image of the object 108 within the 2D video content 134. For example, as shown in FIG. 1I, the 3D object engine 112 may include a mesh generator 182, and the mesh generator 182 is configured to receive the 2D video content 134 including the object 108 and generate the 3D object model 116 to have features that mimic the features of the object within the 2D video content 134. The mesh generator 182 may create the 3D object model 116 from the 2D video content 134 using one or more existing 3D reconstruction techniques. In some examples, the 3D object model 116 includes a 3D mesh representing the structure of the object 108.
[0052] In some examples, the mesh generator 182 may obtain or determine the intrinsic parameters of the camera system 132, such as the focal length, the principal point, and the lens distortion. Using this information, the 3D geometry can be accurately projected onto the 2D image coordinates. To reconstruct the 3D geometry, the mesh generator 182 may extract features or keypoints from the 2D video frames. These features can be points, edges, corners, or other characteristic elements. A feature tracking algorithm is used to match corresponding features across different frames, enabling the tracking of the movement of the features over time. In some examples, the mesh generator 182 may include structure from motion (SfM) techniques. SfM is a technique used to estimate the pose of the camera and reconstruct the 3D structure of a scene from a set of 2D images. It uses the tracked features and the calibration information of the camera to determine the position and orientation of the camera at different points in time. The SfM algorithm estimates the 3D structure by triangulating corresponding feature points from multiple camera viewpoints.
[0053] Once an initial sparse 3D structure is estimated, the mesh generator 182 can refine the initial sparse 3D structure using a dense reconstruction technique. Dense reconstruction can reconstruct the 3D geometry of the object 108 at a higher level of detail by estimating the depth values of the pixels (e.g., all pixels) within the 2D video frame. Techniques such as depth mapping, stereo matching, or structure refinement from motion can be used for dense reconstruction. After obtaining a dense point cloud representing the 3D structure of the object, the mesh generator 182 can generate a 3D mesh representing the surface of the object. Surface reconstruction techniques such as Delaunay triangulation or Poisson surface reconstruction can be applied to convert the point cloud into a mesh representation. These algorithms create a set of connected triangles that approximate the surface geometry of the object. Depending on the quality and accuracy of the initial mesh, additional refinement steps may be required. Mesh refinement techniques such as smoothing, regularization, or texture mapping can be applied to improve the visual quality of the mesh, remove noise, and align it more precisely with the true geometry of the object.
[0054] Referring back to FIG. 1A, the 3D video segment 122 includes animated objects 114 (e.g., those defined by 3D motion data 110 and an extended 3D object model 116) of one or more objects 108 detected within the 2D video content 134. In some examples, the 3D video segment 122 includes an animated object 114 for a single object 108 (e.g., a single player) detected within the 2D video content 134. In some examples, the 3D video segment 122 includes multiple animated objects 114 for multiple objects 108 (e.g., multiple players) detected within the 2D video content 134. The 3D video segment 122 may include another computer-generated model object and content representing a 3D scene. For example, if the 3D scene is related to a basketball event, the 3D video generator 104 may generate the 3D video segment 122 to include another computer-generated model object (e.g., a static model and / or an animation model) representing another object within the 3D scene. Examples of the 3D scene include a court (e.g., an outdoor basketball court, a basketball within an arena in an area), a hoop, fans, lighting and shadows indicating a time of day (e.g., night or midday).
[0055] The video manager 102 system may include a streaming engine 120 configured to generate 3D video content 121 from the 3D video segments 122. In some examples, the streaming engine 120 renders video frames from the 3D video segments 122, and the 3D video content 121 includes the rendered video frames. In some examples, the streaming engine 120 encodes and compresses the rendered video frames, and the 3D video content 121 is a video stream having the encoded and compressed video frames.
[0056] The streaming engine 120 can enable the almost real-time streaming and rendering of 3D graphics and interactive content via the network 150, which may include one or more rendering operations such as geometry processing, shading and material calculation, camera and viewpoint calculation, frame composition, video encoding, and / or network transmission. For example, the streaming engine 120 includes one or more GPUs 126 configured to perform rendering operations on the 3D video segment 122 to generate images (e.g., video frames) of the 3D video content 121. The streaming engine 120 can transmit the 3D video content 121 to the user device 152 for display. In some examples, since the rendering operations are performed by the streaming engine 120, the 3D video content 121 can be transmitted to the user without the need for a user device with special GPU hardware. In some examples, the 3D video segment 122 is a video stream.
[0057] In some examples, the application 146 receives the 3D video content 121, decodes the 3D video content 121, and displays the 3D video content 121 on the interface 140 of the application 146. In some examples, the 3D client manager 155 receives the 3D video content 121, decodes the 3D video content 121, and provides the decoded (rendered) video frames to the application 146 for display on the interface 140. In some examples, the 3D client manager 155 is a program configured to operate as a mediator between the application 146 and the video manager 102. Thus, the 3D client manager 155 can perform operations associated with communication with the video manager 102 and decoding of the 3D video content 121, thereby improving the performance of the application 146.
[0058] The streaming engine 120 may receive information indicating one or more user selections 171 for user controls 140 to adjust and / or customize the playback of the 3D video segment 122. The streaming engine 120 may regenerate the 3D video content 121 from the 3D video segment 122 according to the user selection(s) 171. The streaming engine 120 may send the 3D video content 121-1 to the user device 152 for display at the interface 140. The interface 140 may include one or more user controls 142 that enable the user to view, modify, and share the playback of the 3D video segment 122. In response to the user selection 171 for the user control 142, the streaming engine 120 may receive information regarding the user selection 171, generate the 3D video content 121-2 from the 3D video segment 122 according to the user selection 171, and send the 3D video content 121-2 to the user device 152 for display at the interface 140.
[0059] As shown in FIGS. 1I and 1J, the user control 142 may include one or more customization controls 142a that may modify the playback of the 3D video segment 122. The customization control 142a may include a viewing angle control 131 configured to enable the user to adjust the viewing angle 111 of the 3D video segment 122. The viewing angle control 131 may provide a plurality of viewing angles 111 including the viewing angle 111a and the viewing angle 111b. Although two viewing angles 111 are shown in FIG. 1I, note that the 3D video segment 122 may be replayed from many viewing angles 111 (including all viewing angles 111). In some examples, the viewing angles 111 include a predetermined set of viewing angles 111.
[0060] The customization control 142a may include a virtual content control 135 configured to enable a user to add virtual content 124 to the playback of the 3D video segment 122. For example, the virtual content control 135 may enable the user to add one or more graphics 113 to the 3D video segment 122. The graphics 113 may include statistical data associated with the 3D scene (e.g., the movement speed of a player, the height a player jumped, the number of yards of a passing play, etc.) and other graphics. In some examples, the virtual content control 135 may enable the user to add an animation effect 115 to the 3D scene.
[0061] The customization control 142a may include an object customization control 137 configured to enable a user to modify one or more features 117 of the animated object 114. For example, the user may use the object customization control 137 to change certain aspects of the 3D object model 116 (e.g., change the clothing or other attributes of the underlying avatar), or use a different 3D object model 116 (e.g., use a different avatar). The customization control 142a may include a display speed control 139 configured to enable the user to adjust the display speed 119 of the 3D video segment 122. The display speed control 139 may provide a plurality of display speeds 119 including a viewing angle 119a and a viewing angle 119b. Note that although two display speeds 119 are shown in FIG. 1I, the 3D video segment 122 may be replayed at many display speeds 119. In some examples, the display speed 119 includes a predetermined set of display speeds 119.
[0062] In some examples, interface 140 includes a search field 147 that enables a user to submit a search query 149 from the user, and in response to submission of the search query 149, interface 140 may display a search result 159 that identifies one or more 3D video segments 122 responsive to the search query 149. For example, a user may retrieve 3D highlights for a particular game, player, team, type of movement, etc.
[0063] For example, as shown in FIG. 1K, video manager 102 may include a metadata generator 173. Metadata generator 173 may receive content data 165 regarding 2D video content 134 and generate (and / or associate) metadata 175 with 3D video segments 122. Metadata generator 173 may generate metadata 175 regarding 2D video content 134 and store metadata 175 within 3D video segments 122 in video database 196. In some examples, metadata generator 173 may generate metadata 175 when 3D video generator 104 generates 3D video segments 122. In some examples, metadata generator 173 may also associate corresponding 2D video segment 134a used to generate 3D video segment 122 with 3D video segment 122. In some examples, 3D video segment 122 includes an identifier (e.g., a resource identifier or another identifier) that identifies the location of 2D video segment 134a used to generate 3D video segment 122. In some examples, 2D video segment 134a includes an identifier (e.g., a resource identifier or another identifier) that identifies the location of corresponding 3D video segment 122.
[0064] Content data 165 may include information associated with 2D video content 134, such as a title, location, time, team, central player, description of the underlying event, and / or any metadata included as part of the 2D video content 134. In some examples, content data 165 may include any information detected by object tracker(s) 106. Metadata 175 may include event data 179 regarding time, location, duration, team, record, central player, and / or any information regarding the underlying event. Metadata 175 may include object data 181 regarding object 108 detected by object tracker 106. For example, object data 181 may identify a particular player or information regarding that player. In some examples, metadata 175 includes movement type data 183 that classifies a particular movement of object 108. For example, object tracker 106 may be able to classify a particular player's movement regarding a known sports move (e.g., a 360 - degree dunk, or a windmill dunk, etc.).
[0065] As shown in FIG. 1K, video database 196 may store 3D video segment 122 - 1, and video database 196 may store metadata 175 generated by metadata generator 173 associated with 3D video segment 122 - 1. In some examples, video database 196 may also associate 2D video segment 134a - 1 with 3D video segment 122 - 1, where 2D video segment 134a - 1 was used to generate 3D video segment 122 - 1. Video database 196 may store 3D video segment 122 - 2, and video database 196 may store metadata 175 generated by metadata generator 173 associated with 3D video segment 122 - 2. In some examples, video database 196 may also associate 2D video segment 134a - 2 with 3D video segment 122 - 2, where 2D video segment 134a - 2 was used to generate 3D video segment 122 - 2.
[0066] The video manager 102 may include a video search engine 194. The video search engine 194 receives a search query 149 from the user device 152 and terms (s) included in the search query 149, searches the metadata 175, and may identify one or more 3D video segments 122 in the video database 196 in response to the search query 149. The video search engine 194 may provide search result (s) 159 that identify the 3D video segment (s) 122 in response to the search query 149. In some examples, the search result (s) 159 may also identify a 2D video segment 134a for the corresponding 3D video segment 122 included in the search result (s) 159. Thus, the user may view the original 2D video and explore the highlights in 3D format.
[0067] In some examples, referring to FIGS. 1I and 1L, the interface 140 may include a sharing control 145 configured to enable a user to share the customized 3D video content 121a with one or more other users (e.g., user device 152b). The user may use the user control 142 to modify aspects of the 3D video segment 122, create the customized 3D video content 121a, and use the sharing control 145 to send the customized 3D video content 121a to the user device 152b. In some examples, the user-customized 3D video content 121a is stored in the video manager 102, and the user may create a message with a resource identifier (e.g., a selectable URL link) that identifies the location of the customized 3D video content 121a and send the message to the user device 152b. In some examples, the user device 152 may share the customized 3D video content 121a with another user by posting a message with the customized 3D video content 121a to the social platform 198. In some examples, the user device 152 is associated with a user account 149a of the social platform 198, and the user device 152b is associated with a user account 149b of the social platform 198. In some examples, the user may use the user device 152b to discover a message with the customized 3D video content 121a on the social platform 198.
[0068] Referring to FIG. 1M, in some examples, the video manager 102 may be executable by one or more server computers 160. The server computer(s) 160 may include one or more processors 161 and one or more memory devices 163.
[0069] As shown in FIG. 1M, the user device 152 may include an application 146 and a 3D client manager 155. The user device 152 may include one or more processors 101 and one or more memory devices 103. In some examples, the 3D client manager 155 is configured to communicate with the application 146, detect a user selection 171 for one or more user controls 142, and send information regarding the user selection 171 to the video manager 102. The 3D client manager 155 may receive 3D video content 121 from the video manager 102, decode the 3D video content 121, and provide the 3D video content 121 to the application 146. In some examples, the 3D client manager 155 is a component that is part of the application 146. In some examples, the 3D client manager 155 is a component separate from the application 146. In some examples, the 3D client manager 155 is a component of the operating system 186 or a browser application.
[0070] As shown in FIG. 1N, in some examples, the video manager portion 102-1 is executable by a server computer(s) 160, and the video manager portion 102-2 is executable by the user device 152. For example, some of the operations of the video manager 102 may be executable by a server computer(s) 160, and some of the operations of the video manager 102 may be executable by the user device 152. In some examples, the video manager portion 102-1 includes a 3D video generator 104, and the video manager portion 102-2 includes a streaming engine 120. In some examples, the video manager portion 102-1 includes an object tracker(s) 106 and a 3D object engine 112. In some examples, the video manager portion 102-1 includes an object tracker(s) 106, and the video manager portion 102-2 includes a 3D object engine 112.
[0071] The user device 152 can be any type of computing device that includes one or more processors 101, one or more memory devices 103, a display 123, and an operating system 186 configured to execute (or assist in the execution of) one or more applications (including application 146). In some examples, the display 123 is a 2D display. In some examples, the display 123 is a 3D display. In some examples, the operating system 186 is the application 146. In some examples, the application 146 is an application executable by the operating system 186. In some examples, the user device 152 is a laptop computer. In some examples, the user device 152 is a desktop computer. In some examples, the user device 152 is a tablet computer. In some examples, the user device 152 is a smartphone. In some examples, the user device 152 is a wearable device. In some examples, the user device 152 is a virtual reality (VR) device (which may include a headset). In some examples, the user device 152 is an augmented reality (AR) device. In some examples, the display 123 is the display of the user device 152.
[0072] The processor(s) 101 can be formed on a substrate configured to execute one or more machine-executable instructions, or software, firmware, or a combination thereof. The processor(s) 101 can be semiconductor-based. That is, the processor can include a semiconductor material capable of performing digital logic. The memory device(s) 103 can include a main memory that stores information in a format readable and / or executable by the processor(s) 101. The memory device(s) 103 can store the application 146, the operating system 186, and / or the 3D client manager 155, which, when executed by the processor 101, perform the specific operations described herein. In some examples, the memory device(s) 103 includes a non-transitory computer-readable medium that includes executable instructions to cause at least one processor (e.g., processor 101) to perform an operation.
[0073] The server computer(s) 160 can be a computing device that takes various forms of devices, such as, for example, a standard server, a group of such servers, or a rack server system. In some examples, the server computer(s) 160 can be a single system that shares components such as a processor and memory. In some examples, the server computer(s) 160 can be multiple systems that do not share a processor and memory. The network 150 can include the Internet and / or another type of data network, such as a local area network (LAN), a wide area network (WAN), a cellular network, a satellite network, or another type of data network. The network 150 can also include any number of computing devices (e.g., computers, servers, routers, network switches, etc.) configured to receive and / or transmit data within the network 150. The network 150 can further include any number of wired and / or wireless connections.
[0074] The video manager 102 (or a part thereof) may be executable by one or more server computers 160. The server computer(s) 160 may include one or more processors 161 formed on a substrate, an operating system (not shown), and one or more memory devices 163. The memory device(s) 163 may represent any type (or types) of memory (e.g., RAM, flash, cache, disk, tape, etc.). In some examples (not shown), the memory device may include external storage, e.g., memory that is physically distant from but accessible from the server computer(s) 160. The processor(s) 161 may be formed on a substrate configured to execute one or more machine-executable instructions, or a part of software, firmware, or a combination thereof. The processor(s) 161 may be semiconductor-based. That is, the processor may include semiconductor material capable of performing digital logic. The memory device(s) 163 may store information in a format readable and / or executable by the processor(s) 161. The memory device(s) 163 may store the video manager 102 (or a part thereof) that performs the specific operations described herein when executed by the processor(s) 161. In some examples, the memory device(s) 163 includes a non-transitory computer-readable medium containing executable instructions to cause at least one processor (e.g., the processor(s) 161) to perform operations.
[0075] Figures 2A - 2M show examples of one or more interfaces 240 of an application for displaying a 3D video segment and user controls for controlling the 3D video segment, according to an aspect.
[0076] As shown in FIG. 2A, interface 240 may display 2D video segment 234. 2D video segment 234 may be a highlight from a football program. In some examples, the football program is still live (e.g., being captured by camera system 132 of FIGS. 1A-1N). 2D video segment 234 may be related to a previous key play in the football program. 2D video segment 234 may be one of several key plays in the football program. 2D video segment 234 may be an example of 2D video segment 134a of FIGS. 1A-1N. For example, segment selector 128 may detect key event 127 within the 2D video and select 2D video segment 234 as a highlight included in the key play of the football program. In some examples, in response to 2D video segment 234 being detected as a highlight, 3D video generator 104 of FIGS. 1A-1N may generate a 3D video segment. 2D video segment 234 includes a plurality of objects 208. Object 208 is the body structure of a football player.
[0077] As shown in FIG. 2A, interface 240 includes UI element 207a, which, when selected, enables the user to view the 2D video segment 234 in 3D format. UI element 207a can be an example of UI element 107 of FIGS. 1A-1N. The user can select UI element 207a at any time during the display of the 2D video segment 234 (or after the 2D video segment 234 has ended). In some examples, in response to the selection of UI element 207a, as shown in FIG. 2B, interface 240 can display UI information elements 213 (e.g., "Swipe to rotate") and UI information element 215 ("Pinch to zoom") regarding user selections that can adjust the view of the 3D video segment. Interface 240 can also display UI element 207b ("Explore in 3D"), which, when selected, causes a 3D viewing request (e.g., 3D viewing request 184 of FIGS. 1A-1N) to be sent to streaming engine 120 (e.g., streaming engine 120 of FIGS. 1A-1N). In some examples, the selection of UI element 207a causes a 3D viewing request to be sent to the streaming engine.
[0078] As shown in FIG. 2C, interface 240 displays 3D video content 221a. The 3D video content 221a includes a plurality of animated objects 214 that move in a manner corresponding to the movement of the objects 208 within the 2D video segment 234. The animated objects 214 can include players and balls. Each animated object 214 can be an example of the animated object 114 of FIGS. 1A - 1N. In response to a 3D viewing request, the streaming engine can obtain a 3D video segment corresponding to the 2D video segment 234 (e.g., the 3D video segment 122 of FIGS. 1A - 1N). The streaming engine can generate the 3D video content 221a from the 3D video segment and transmit the 3D video content 221a to the user device for display on the interface 240. In some examples, for the initial view, the streaming engine can select the default settings of the user control 242 and generate the 3D video content 221a from the 3D video segment according to the default settings.
[0079] The interface 240 includes a plurality of user controls 242 for adjusting the playback of the 3D video segment. The user controls 242 include a viewing angle control 231 configured to allow the user to adjust the viewing angle and a display speed control 239 configured to allow the user to adjust the display speed of the playback. The user controls 242 can include a statistical setting control 235a, which, when enabled, inserts statistical data (including graphics) into the display of the 3D video segment. The user controls 242 can include an effect setting control 235b, which, when enabled, inserts one or more animation effects into the display of the 3D video segment. The user controls 242 can include a sharing control 245, which, when selected, provides one or more interfaces that allow the user to share the 3D video content with another user.
[0080] In response to the selection of the viewing angle control 231, as shown in FIG. 2D, the interface 240 displays a menu at a plurality of viewing angles 211. The viewing angles 211 may include a viewing angle 211a (drone camera), a viewing angle 211b (bird's-eye view), a viewing angle 211c (in-game), a viewing angle 211d (quarterback view), a viewing angle 211e (ball view), and a viewing angle 211f (on the sideline).
[0081] In response to the selection of the viewing angle 211d, the application may send information regarding the user selection (e.g., viewing angle 211d) to the streaming engine. As shown in FIG. 2E, the streaming engine may generate 3D video content 221b from 3D video segments according to the user selection and send it to the application for display on the interface 240. The 3D video content 221b represents highlights from a quarterback's perspective.
[0082] In response to the selection of the viewing angle 211e, the application may send information regarding the user selection (e.g., viewing angle 211e) to the streaming engine. As shown in FIG. 2F, the streaming engine may generate 3D video content 221c from 3D video segments according to the user selection and send it to the application for display on the interface 240. The 3D video content 221c represents highlights from a ball's perspective. In some examples, the user may adjust the display speed of the playback. For example, referring to FIG. 2F, the user used the display speed control 239 to adjust the display speed to a slow motion speed to view the 3D video segments in slow motion from a ball's perspective. In some examples, the application may send information regarding the user selection (e.g., slow motion speed), and the streaming engine may adjust the playback to the slow motion speed. In some examples, the playback speed is controlled by the application.
[0083] In response to the selection of the statistics setting control 235a, the application may send information regarding the user selection (e.g., the statistics setting control 235a) to the streaming engine. As shown in FIGS. 2G and 2H, the streaming engine may generate 3D video content 221d from 3D video segments according to the user selection and send the 3D video content 221d to the application for display on the interface 240. The 3D video content 221d represents an animated object 214 with virtual content 224 (e.g., graphics and / or statistical data regarding the play) added to the replay.
[0084] In response to the selection of the effect setting control 235b, as shown in FIG. 2I, the interface 240 may display a menu with a plurality of effect options 237. The effect options 237 may include an effect option 237a (celebration), an effect option 237b (8-bit mode), and an effect option 237c (retro mode). In response to the selection of the effect option 237a, the application may send information regarding the user selection (e.g., the effect option 237a) to the streaming engine. As shown in FIG. 2E, the streaming engine may generate 3D video content 221e from 3D video segments according to the user selection and send the 3D video content 221e to the application for display on the interface 240. The 3D video content 221e represents an animated object 214 with virtual content 224 added to the 3D scene.
[0085] In response to the selection of the co - control 245, the application may display a clip creator interface 271 that enables the user to create a video clip 269 that includes a portion of the 3D video content 221e. The video clip 269 with a portion of the 3D video content 221e can be shared with one or more other users. The video clip 269 can be an example of the customized 3D video content 121a of FIGS. 1A - 1N. The clip creator interface 271 may include one or more movable elements 275 that enable the user to define the start and end of the video clip, and a progress indicator 273 that indicates the time position of the currently displayed video frame. The user can use the progress indicator 273 to move forward and backward within the video clip. After the user finishes editing the video clip, the user may select the sharing element 245a, and the application may display a sharing interface 299 with multiple sharing options for sharing the video clip 269. The user can download the video clip 269, post the video clip 269 to one or more social platforms or media platforms so that other users can discover it, send the video clip 269 via email or direct message, and / or upload it to an online storage system. In some examples, when the user posts the video clip 269 to a social platform or media platform, another user may discover a message identifying the video clip 269 and can select the video clip 269 to view the user's customized 3D highlight video.
[0086] Figures 3A and 3B show an example of the application interface 340. The interface 340 may display highlights of a basketball game. For example, the interface 340 may display a 2D video segment 334 of a basketball play. The interface 340 may display a UI element 307, which, when selected, enables the 2D video segment 334 to be viewable in 3D format.
[0087] In some examples, in response to the selection of the UI element 307, the application (or a 3D client manager that communicates with the application (e.g., the 3D client manager 155 in FIGS. 1A - 1N)) may send a 3D viewing request (e.g., the 3D viewing request 184 in FIGS. 1A - 1N) to a video manager (e.g., the video manager 102 in FIGS. 1A - 1N). In some examples, the video manager is configured to generate a 3D video segment (e.g., the 3D video segment 122 in FIGS. 1A - 1N) from the 2D video segment 334. In some examples, the video manager may retrieve a pre - created 3D video segment from a video database. In response to the 3D viewing request, as shown in FIG. 3B, the interface 340 may display 3D video content 321. The 3D video content 321 includes animated objects 314 (e.g., the animated objects 114 in FIGS. 1A - 1N) that move in a manner corresponding to the movement of the players within the 2D video segment 334.
[0088] Figures 4A - 4C show a comparison diagram of a display 401 of 2D video content 434 having an object 408 (e.g., a basketball player) within a timeline 405 and a display 403 of 3D video content 421 having an animated object 414. Figure 4A shows a frame of the 2D video content 434 together with a corresponding frame of the 3D video content 421 at a point in the timeline 405. Figure 4B shows frames of the 2D video content 434 together with corresponding frames of the 3D video content 421 at successive points in the timeline 405. Figure 4C shows frames of the 2D video content 434 together with corresponding frames of the 3D video content 421 at successive points in the timeline 405. As shown in FIGS. 4A - 4C, the animated object 414 can move in a manner corresponding to the movement of the object 408.
[0089] Figures 5A to 5C show an example of 3D video content 521 created from 2D video content 534, and the animated object 514 within the 3D video content 521 can move in a manner corresponding to the movement of the object 508 within the 2D video content 534. Figure 5A shows a video frame 535 from the 2D video content 534. The 2D video content 534 includes an object 508 (e.g., a basketball player performing a dunk). Figure 5B shows the 3D video content 521 with an animated object 514 configured to move in a manner corresponding to the movement of the object 508 within the 2D video content 521. The 3D video content 521 also includes other computer-generated content such as a virtual object 540 (e.g., a basketball backboard and hoop), a virtual object 516 (e.g., a ball), a virtual object 530 (e.g., a court), and a virtual object 520 (e.g., background graphics). The animated object 514 can be the user's avatar extended with 3D motion data, thereby enabling the animated object 514 to move in the same / similar manner as the basketball player within the 3D video content 521. In some examples, one or more aspects of the animated object 514 can be modified by the user (e.g., changing clothing, adding accessories, etc.). Referring to Figure 5C, the user can add one or more animation effects 515 to the animated object 514 and / or the 3D scene.
[0090] FIG. 6 is a flowchart 600 representing an exemplary operation of a system for generating and delivering 3D video segments. Flowchart 600 may represent the operations of a method implemented on a computer. Although flowchart 600 is described with respect to system 100 of FIGS. 1A - 1N, flowchart 600 may be applicable to any of the embodiments described herein. The flowchart 600 of FIG. 6 shows operations in sequence, but this is merely an example, and it will be recognized that additional or alternative operations may be included. Further, the operations and related operations of FIG. 6 may be executed in an order different from that shown, or may be executed in parallel or in an overlapping manner.
[0091] Operation 602 includes obtaining two - dimensional (2D) video content 134 captured by camera system 132. Operation 604 includes generating a three - dimensional (3D) video segment 122 from the 2D video content 134, where the 3D video segment 122 includes an animated object 114 configured to move in a manner corresponding to the movement of an object 108 within the 2D video content 134. In some examples, generating includes obtaining 3D motion data 110 of an object 108 detected within the 2D video content 134 from a 3D pose estimation engine 188 and generating an animated object 114 based on the 3D motion data 110, and the motion of the animated object 114 corresponds to the movement of the object 108 within the 2D video content 134. Operation 606 includes generating 3D video content 121 from the 3D video segment 122. Operation 608 includes transmitting the 3D video content 121 to a user device 152 for display.
[0092] FIG. 7 is a flowchart 700 depicting an exemplary operation of a system for rendering 3D video segments. Flowchart 700 may represent the operations of a method implemented on a computer. Although flowchart 700 is described with respect to system 100 of FIGS. 1A - 1N, flowchart 600 may be applicable to any of the embodiments described herein. The flowchart 700 of FIG. 7 shows the operations in sequence, but this is merely an example, and it should be recognized that additional or alternative operations may be included. Further, the operations and related operations of FIG. 7 may be executed in an order different from that shown, or may be executed in parallel or in an overlapping manner.
[0093] Operation 702 includes transmitting, via network 150, a three-dimensional (3D) viewing request 184 to a video manager 102 executable by at least one server computer 160, where the 3D viewing request 184 is a request to view two-dimensional (2D) video content 134 captured by camera system 132 in a 3D format. Operation 704 includes receiving, via network 150, 3D video content 121 from video manager 102, where the 3D video content 121 is generated from a 3D video segment 122 generated by video manager 102 using 2D video content 134. Operation 706 includes starting the display of 3D video content 121 on interface 140 of user device 152, where the 3D video content 121 includes an animated object 114 that moves in a manner corresponding to the movement of an object 108 within 2D video content 134.
[0094] FIG. 8 is a flowchart 800 that illustrates an exemplary operation of a system for extracting and delivering 3D video segments. The flowchart 800 may represent the operation of a method implemented on a computer. Although the flowchart 800 is described with respect to the system 100 of FIGS. 1A-1N, the flowchart 800 may be applicable to any of the embodiments described herein. The flowchart 800 of FIG. 8 shows the operations in order, but this is merely an example, and it will be recognized that additional or alternative operations may be included. Further, the operations of FIG. 8 and related operations may be executed in an order different from that shown, or may be executed in parallel or in an overlapping manner.
[0095] Operation 802 includes receiving, from the user device 152, a 3D viewing request 184 to view 2D video content 134 in 3D format. Operation 804 includes retrieving, from the video database 196, a 3D video segment 122 corresponding to the 2D video content 134, where the 3D video segment 122 includes an animated object 114, and the movement of the animated object 114 corresponds to the movement of an object 108 within the 2D video content 134. Operation 806 includes generating 3D video content 121 from the 3D video segment 122. Operation 808 includes transmitting the 3D video content 121 to the user device 152 for display.
[0096] Clause 1. A method comprising generating a three-dimensional (3D) video segment from two-dimensional (2D) video content captured by a camera system, said generating comprising obtaining 3D motion data of an object detected within the 2D video content from a 3D pose estimation engine, and generating the animated object based on the 3D motion data such that movement of the animated object corresponds to movement of the object within the 2D video content, said method further comprising generating 3D video content from the 3D video segment and transmitting the 3D video content to a user device for display.
[0097] Clause 2. The method according to clause 1, further comprising receiving, from the user device, a 3D viewing request for viewing at least a portion of the 2D video content in 3D format, and in response to the 3D viewing request, generating the 3D video content from the 3D video segment.
[0098] Clause 3. The method according to clause 1, further comprising storing the 3D video segment in a video database, receiving, from the user device, a 3D viewing request for viewing at least a portion of the 2D video content in 3D format, retrieving the 3D video segment from the video database in response to the 3D viewing request, and generating the 3D video content from the 3D video segment.
[0099] Clause 4. The method according to any one of clauses 1 to 3, wherein the 2D video content includes television content, further comprising detecting that a portion of the television content includes a key event, identifying a 2D video segment from the portion of the television content, and generating the 3D video segment from the 2D video segment.
[0100] Clause 5. The 3D video content is the first 3D video content, and the method further includes receiving, from the user device, information indicating a user selection for adjusting the playback of the 3D video segment; generating second 3D video content based on the 3D video segment and the user selection; and transmitting the second 3D video content to the user device for display, and the method according to any one of Clauses 1 to 4.
[0101] Clause 6. The method according to Clause 5, wherein the user selection includes at least one of adjusting a viewing angle or adjusting a feature of the animated object.
[0102] Clause 7. The method according to Clause 5 or 6, wherein the user selection includes a selection of virtual content, and the virtual content includes at least one of statistical data or an animation effect.
[0103] Clause 8. Generating the 3D motion data includes estimating a 3D position of at least one of a plurality of key points of the object from 2D positions of the object in a plurality of frames of 2D video content, and the method according to any one of Clauses 1 to 7.
[0104] Clause 9. The method according to any one of Clauses 1 to 8, further including generating the animated object by applying the 3D motion data to a 3D object model representing the object detected in the 2D video content.
[0105] Clause 10. The method according to Clause 9, including selecting the 3D object model based on object data associated with the object from an object model database, or generating the 3D object model based on the 2D video content.
[0106] Clause 11. The method according to any one of Clauses 1 to 10, wherein the 3D motion data includes position information and rotation information describing the movement and orientation of the object in 3D space.
[0107] Clause 12. A non-transitory computer-readable medium storing executable instructions for causing at least one processor to perform operations, the operations including transmitting, via a network, a three-dimensional (3D) viewing request to a video manager executable by at least one server computer, the 3D viewing request being a request to view two-dimensional (2D) content captured by a camera system in a 3D format, the operations further including receiving, via the network, 3D video content from the video manager, the 3D video content being generated from 3D video segments generated by the video manager using the 2D video content, the operations further including starting display of the 3D video content on an interface of a user device, the 3D video content including animated objects, and movement of the animated objects corresponding to movement of objects in the 2D video content.
[0108] Clause 13. The non-transitory computer-readable medium according to Clause 12, wherein the 3D video content is first 3D video content, the interface includes user controls configured to enable a user to adjust playback of the 3D video segments, and the operations further include transmitting, via the network, information indicating a user selection to the user controls and receiving, via the network, second 3D video content with an image according to the user selection.
[0109] Clause 14. The non-transitory computer-readable medium according to Clause 13, wherein the user controls include viewing angle controls configured to enable a user to adjust a viewing angle of the 3D video segments.
[0110] Clause 15. The user control includes virtual content control configured to enable the user to add virtual content to the playback of the 3D video segment, and the non-transitory computer-readable medium according to Clause 13 or 14.
[0111] Clause 16. The user control includes object customization control configured to enable the user to modify the characteristics of the animated object, and the non-transitory computer-readable medium according to any one of Clauses 13 to 15.
[0112] Clause 17. The 3D video segment is a first 3D video segment, the interface includes a search field configured to enable the user to submit a search query, and the operation further includes receiving at least one search result that identifies a second 3D video segment responsive to the search query in response to the submission of the search query, and the non-transitory computer-readable medium according to any one of Clauses 12 to 16.
[0113] Clause 18. The user device is a first user device, the interface includes sharing control configured to enable the user to share customized 3D video content with a second user device, and the non-transitory computer-readable medium according to any one of Clauses 12 to 17.
[0114] Clause 19. An apparatus comprising at least one processor and a non-transitory computer-readable medium storing executable instructions, the executable instructions causing the at least one processor to receive, from a user device, a three-dimensional (3D) viewing request to view two-dimensional (2D) video content in 3D format, retrieve from a video database a 3D video segment corresponding to the 2D video content, the 3D video segment including animated objects, the movement of the animated objects corresponding to the movement of objects in the 2D video content, the executable instructions further causing the at least one processor to generate 3D video content from the 3D video segment and transmit the 3D video content to the user device for display.
[0115] Clause 20. The apparatus according to Clause 19, wherein the 3D video content is first 3D video content, and the executable instructions cause the at least one processor to receive, from the user device, information indicating a user selection for adjusting playback of the 3D video segment, generate second 3D video content based on the 3D video segment and the user selection, and transmit the second 3D video content to the user device for display.
[0116] Clause 21. The apparatus according to Clause 20, wherein the user selection includes at least one of adjustment of a viewing angle or adjustment of characteristics of the animated objects.
[0117] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include embodiments in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, where the at least one programmable processor can be special purpose or general purpose and coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0118] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in a high-level procedural programming language, and / or an object-oriented programming language, and / or assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic circuits (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0119] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Similarly, other types of devices can be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and the input from the user can be received in any form, including acoustic input, speech input, or tactile input.
[0120] The systems and techniques described herein can be implemented on a computing system that includes back-end components (e.g., data servers, etc.), or the computing system includes middleware components (e.g., application servers), or the computing system includes front-end components (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0121] A computing system can include clients and servers. The clients and servers are generally remote from each other and typically interact via a communication network. The client - server relationship is created by computer programs that run on each computer and have a client - server relationship with each other.
[0122] In this specification and the appended claims, the singular forms "a", "an", and "the" do not exclude a plural reference unless the context clearly dictates otherwise. Further, conjunctions such as "and", "or", and "and / or" are inclusive unless the context clearly dictates otherwise. For example, "A and / or B" includes A alone, B alone, and A and B. Additionally, the connecting lines or connectors shown in the various figures presented are intended to represent exemplary functional relationships and / or physical or logical couplings between various elements. Many alternative or additional functional relationships, physical connections, or logical connections may exist in a practical device. Further, an article or component is not essential for the practice of the embodiments disclosed herein unless it is specifically described as "essential" or "important".
[0123] Terms such as, but not limited to, generally, substantially, and typically are used herein to indicate that their exact values or ranges are not required and need not be specified. When used herein, the above - mentioned terms are ones that can be readily and immediately understood in meaning by one of ordinary skill in the art.
[0124] Further, the use of terms such as above, below, upper, lower, side, end, front, back, etc. in this specification is made with reference to the orientation currently under consideration or illustrated. It should be understood that if they are considered with respect to other orientations, such terms need to be modified accordingly.
[0125] Furthermore, in this specification and the appended claims, the singular forms "a", "an", and "the" do not exclude plural references unless the context clearly dictates otherwise. Additionally, conjunctions such as "and", "or", and "and / or" are inclusive unless the context clearly dictates otherwise. For example, "A and / or B" includes A alone, B alone, and both A and B.
[0126] Specific exemplary methods, apparatuses, and articles of manufacture are described herein, but the scope of this patent is not limited thereto. It should be understood that the terminology used herein is for the purpose of describing particular embodiments and is not intended to be limiting. On the contrary, this patent is directed to all methods, apparatuses, and articles of manufacture fairly falling within the scope of the claims of this patent.
Claims
1. A method comprising: generating a three-dimensional (3D) video segment from two-dimensional (2D) video content captured by a camera system, said generating comprising: obtaining 3D motion data of an object detected within said 2D video content from a 3D pose estimation engine; and generating an animated object based on said 3D motion data such that movement of the animated object corresponds to movement of the object within said 2D video content, said method further comprising: generating 3D video content from said 3D video segment; and transmitting said 3D video content to a user device for display; A method.
2. Receiving, from the user device, a 3D viewing request for viewing at least a portion of the 2D video content in 3D format; and In response to the 3D viewing request, generating the 3D video content from the 3D video segment; The method according to claim 1, further comprising.
3. Storing the 3D video segment in a video database; Receiving, from the user device, a 3D viewing request for viewing at least a portion of the 2D video content in 3D format; In response to the 3D viewing request, Retrieving the 3D video segment from the video database; and Generating the 3D video content from the 3D video segment; The method according to claim 1, further comprising.
4. The 2D video content includes television content, Detecting that a portion of the television content includes a key event; Identifying a 2D video segment from said portion of the television content; and Generating the 3D video segment from the 2D video segment; The method according to any one of claims 1 to 3, further comprising.
5. The 3D video content is a first 3D video content, and the method further comprises: Receiving, from the user device, information indicating a user selection for adjusting playback of the 3D video segment; and Generating a second 3D video content based on the 3D video segment and the user selection; Transmitting to the user device for displaying the second 3D video content; The method according to any one of claims 1 to 4, comprising:
6. The method according to claim 5, wherein the user selection includes at least one of adjusting a viewing angle or adjusting a feature of the animated object.
7. The method according to claim 5 or 6, wherein the user selection includes a selection of virtual content, and the virtual content includes at least one of statistical data or an animation effect.
8. The method according to any one of claims 1 to 7, wherein generating the 3D motion data includes estimating a 3D position of at least one of a plurality of key points of the object from 2D positions of the object in a plurality of frames of 2D video content.
9. The method according to any one of claims 1 to 8, further comprising generating the animated object by applying the 3D motion data to a 3D object model representing the object detected in the 2D video content.
10. Selecting the 3D object model from an object model database based on object data associated with the object, or Generating the 3D object model based on the 2D video content, The method according to claim 9, further comprising:
11. The method according to any one of claims 1 to 10, wherein the 3D motion data includes position information and rotation information describing movement and orientation of the object in 3D space.
12. A non-transitory computer-readable medium storing executable instructions for causing at least one processor to perform operations, the operations including: Transmitting, via a network, a three-dimensional (3D) viewing request to a video manager executable by at least one server computer, the 3D viewing request being a request to view two-dimensional (2D) video content captured by a camera system in a 3D format, and the operations further including: Receiving, via the network, 3D video content from the video manager, the 3D video content being generated from 3D video segments generated by the video manager using the 2D video content, the operation further comprising starting display of the 3D video content on an interface of a user device, the 3D video content including animated objects, movement of the animated objects corresponding to movement of objects within the 2D video content, a non-transitory computer-readable medium. **Claim 13** The 3D video content is first 3D video content, the interface including user controls configured to enable a user to adjust playback of the 3D video segments, the operation further comprising transmitting, via the network, information indicative of a user selection to the user controls; and receiving, via the network, second 3D video content with an image according to the user selection, The non-transitory computer-readable medium according to claim 12. **Claim 14** The non-transitory computer-readable medium according to claim 13, wherein the user controls include viewing angle controls configured to enable the user to adjust a viewing angle of the 3D video segments. **Claim 15** The non-transitory computer-readable medium according to claim 13 or 14, wherein the user controls include virtual content controls configured to enable the user to add virtual content to the playback of the 3D video segments. **Claim 16** The non-transitory computer-readable medium according to any one of claims 13 to 15, wherein the user controls include object customization controls configured to enable the user to modify characteristics of the animated objects. **Claim 17** The 3D video segment is a first 3D video segment, the interface including a search field configured to enable a user to submit a search query, the operation further comprising A non - transitory computer - readable medium according to any one of claims 12 - 16, comprising receiving, in response to submission of the search query, at least one search result that identifies a second 3D video segment responsive to the search query.
18. The user device is a first user device, and the interface includes sharing control configured to enable a user to share customized 3D video content with a second user device. A non - transitory computer - readable medium according to any one of claims 12 - 17.
19. An apparatus, comprising: at least one processor; and a non - transitory computer - readable medium storing executable instructions, the executable instructions causing the at least one processor to: receive, from a user device, a three - dimensional (3D) viewing request to view two - dimensional (2D) video content in 3D format; retrieve, from a video database, a 3D video segment corresponding to the 2D video content, the 3D video segment including animated objects, wherein movement of the animated objects corresponds to movement of objects in the 2D video content, and the executable instructions further cause the at least one processor to: generate 3D video content from the 3D video segment; transmit the 3D video content to the user device for display. An apparatus.
20. The 3D video content is first 3D video content, and the executable instructions cause the at least one processor to: receive, from the user device, information indicating a user selection to adjust playback of the 3D video segment; generate second 3D video content based on the 3D video segment and the user selection; transmit the second 3D video content to the user device for display. The apparatus according to claim 19, comprising instructions for causing the above - mentioned operations.
21. The apparatus according to claim 20, wherein the user selection includes at least one of adjustment of a viewing angle or adjustment of characteristics of the animated objects.
Citation Information
Patent Citations
Method and system for generating an image
US20200035019A1
Techniques for rendering three-dimensional animated graphics from video
US20200193671A1