Method, computing device, and computer-readable storage hardware for mixed reality animation

By building a 3D model in a mixed reality system and mapping animation features using user input, the problem of animation production complexity in the existing technology is solved, intuitive real-time animation production is achieved, and efficiency and flexibility are improved.

CN112585647BActive Publication Date: 2025-08-05MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980036140.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-05-31
Filing Date
2019-05-20
Publication Date
2025-08-05
Estimated Expiration
2039-05-20

AI Technical Summary

Technical Problem

It is difficult for existing mixed reality systems to produce mixed reality animations in an intuitive and effective way, and users need to manually code elements such as animated 3D characters.

Method used

The 3D model is constructed through a mixed reality system to receive physical scene videos, provide pose updates using spatial sensing, user input defines the input path, and map animation features to the 3D path. The animation command is executed under the path guidance.

Benefits of technology

It realizes intuitive and real-time animation production in mixed reality systems, reducing the complexity of traditional programming and improving the efficiency and flexibility of animation production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112585647B_ABST
    Figure CN112585647B_ABST
Patent Text Reader

Abstract

A mixed reality system including a display and a camera is configured to receive a video of a physical scene from the camera and construct a 3D model of the physical scene based on the video. Spatial sensing provides posture (positioning and orientation) updates corresponding to the physical posture of the display. A first user input allows the user to define an input path. The input path can be displayed as a graphical path or line. The input path is mapped to a 3D path in the 3D model. A second user input defines animation features associated with the 3D path. Animation features include objects (e.g., characters), animation commands, etc. Animation commands can be manually mapped to points on the 3D path and executed during the animation of the object guided by the 3D path.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Mixed reality systems are becoming increasingly accessible thanks to improvements in hardware and software. Improved processing power, especially for handheld devices with integrated cameras, is making real-time mixed reality rendering possible. Mixed reality systems with advanced programming suites are easing the challenges of developing mixed reality applications.

[0002] A mixed reality system typically builds a three-dimensional (3D) model of the physical scene being viewed with a camera. By analyzing the camera's video output and by tracking the camera's spatial movement, the mixed reality system can maintain a continuously changing transformation for alignment between the camera's changing physical pose (position and orientation) and the current view of a 3D model that can be drawn and displayed. The mixed reality system draws its elements for the 3D model from a virtual view that corresponds to the camera's physical pose. The user will see a drawing of the model superimposed on or fused with the physical scene; the physical scene is viewed on a display showing the video from the camera, or directly through a translucent surface. In short, the mixed reality system presents virtual and real visual information in a unified manner that gives the person the perception that they form a single space regardless of the movement of the display.

[0003] Displaying virtual animations is a common use of mixed reality systems. Although mixed reality systems have become economical and practical, it is still impossible to produce mixed reality animations in real time in an intuitive and effective way. Previously, users had to program animations using traditional 3D programming techniques. For example, if you want to animate a 3D character model, the lifespan, position, movement, orientation, interaction with the 3D model of the physical scene (the apparent interaction with the physical scene), logic, and behavior have mostly been manually coded in advance for arbitrary scene geometry.

[0004] The following discusses mixed reality animation techniques that can avoid this difficulty. Summary of the Invention

[0005] The following "Summary" is included only to introduce some of the concepts discussed in the "Detailed Description" below. This "Summary" is not comprehensive and is not intended to delineate the scope of the claimed technical solutions as set forth by the ultimately filed claims.

[0006] A mixed reality system including a display and a camera is configured to receive a video of a physical scene from the camera and construct a 3D model of the physical scene based on the video. Spatial sensing provides posture (positioning and orientation) updates corresponding to the physical posture of the display. A first user input allows a user to define an input path. The input path can be displayed as a graphical path or line. The input path is mapped to a 3D path in the 3D model. A second user input defines animation features associated with the 3D path. Animation features include objects (e.g., characters), animation commands, etc. Animation commands can be mapped to points on the 3D path and executed during the animation of the object guided by the 3D path.

[0007] Many of the attendant features will be explained below with reference to the following detailed description considered in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The present specification will be better understood from the following detailed description read with reference to the accompanying drawings, wherein like reference numerals are used to refer to like parts throughout the description of the drawings.

[0009] Figure 1 A mixed reality configuration is shown.

[0010] Figure 2 Another mixed reality configuration is shown.

[0011] Figure 3 It shows how a mixed reality system builds a 3D model of a physical scene and renders views of the 3D model.

[0012] Figure 4 A process for defining animation paths in a mixed reality presentation is shown.

[0013] Figure 5 A process for defining animation paths in a mixed reality presentation is shown.

[0014] Figure 6 A user interface for interactively defining an animation associated with an animation path is shown.

[0015] Figure 7 A process for executing an animation path is shown.

[0016] Figure 8 An example simulation of an animation is shown.

[0017] Figure 9 Details of a computing device are shown upon which the above-described embodiments may be implemented. DETAILED DESCRIPTION

[0018] Figure 1-Figure 3The following illustrates the types of mixed reality systems to which the embodiments described herein may be applied. The term "mixed reality" as used herein refers to augmenting real-time video ( Figure 1 ) and augmenting the direct view of reality with computer-generated graphics ( Figure 2 ).

[0019] Figure 1 A mixed reality presentation is shown in which the eyes of a viewer or user 100 receive a mixture of: (i) real-world light 102 reflected from a physical scene 104 and (ii) computer-rendered light 106. That is, the user perceives the mixed reality as a composite of computer-generated light and real-world light. Real-world light 102 is light from an ambient source (artificial or natural) that has been reflected from the physical scene 104 and passed to the eyes of the user 100 as is; real-world light is not computer-rendered light and may be passed to the eyes directly, by reflection, and / or by transmission through a transparent or optically transforming material. In contrast, computer-rendered light 106 is emitted by any type of display 108, which converts a computer-generated video signal 110 into light that forms an image corresponding to the content of the video signal 110.

[0020] Display 108 may be any type of such signal to light conversion device. Figure 1 In hybrid physical virtual reality of the type shown, the display 108 may be any type of device that allows both real-world light 102 and computer-rendered light 106 (generated by the display 108) to simultaneously fall upon the eyes of the user 100, thereby forming a composite physical-virtual image on the retina of the user 100. The display 108 may be a transparent or translucent device ("transparent" as used hereinafter will also refer to "translucent") that both generates the computer-rendered light 106 and allows the real-world light 102 to pass through it (often referred to as a "heads-up" display). Figure 1 The display 108 in the case of a head mounted display can be a small video projector mounted on goggles or glasses that projects its image onto the transparent lenses of the goggles or glasses (head mounted heads up display). The display 108 can be a projector that projects onto a large transparent surface (fixed heads up display). The display 108 can be a small projector that projects directly onto the user's retina without the use of a reflective surface. The display 108 can be a transparent volumetric display or a three-dimensional (3D) projection. Any type of device that can convert the video signal 110 into visible light and that can also allow this light to be composited with physical world light will be suitable for use in the present invention. Figure 1 The type of mixed reality shown.

[0021] Figure 2A mixed reality configuration is shown in which the eyes of the user 100 perceive the mixed reality as primarily computer-rendered light 106. The computer-rendered light 106 comprises a rendered video whose frames include (i) real-world image data of a physical scene 104 captured by a camera 120, and (ii) computer-generated virtual image data. The virtual image data is computer-generated and rendered, for example, from a 3D model 122 approximating the geometry (and possibly other features) of the physical scene 104, a two-dimensional (2D) model (e.g., a windowed desktop), or other virtual space under the interactive control of the user 100. The 3D model 122 may be a reconstruction of the physical scene 104 by applying known image processing algorithms to the signal from the camera 120, possibly in combination with concurrent information about the pose of the camera. Figure 1 The mixed reality system can also reconstruct 3D models from its video stream.

[0022] exist Figure 2 In the type of mixed reality shown, user 100 sees a full computer-rendered image, but the image seen by the viewer includes both artificially generated graphics data and image data provided by camera 120. Note that the video signal from camera 120 can be a pre-recorded signal or a real-time signal. The mixed reality view is presented by display 108, which can be a flat panel display, a touch-sensitive display surface, a projector, a volumetric display, a head-mounted display (e.g., virtual reality (VR) goggles), or any other technology for producing a full-frame rendering of the video generated by the computing device.

[0023] Figure 1 and Figure 2 The mixed reality system shown can be configured such that the camera and display 108 are both part of a rigid body mixed reality device (e.g., a wearable or mobile device). Such a mixed reality device can also have a known hardware system for tracking and reporting changes in the relative physical position and orientation (pose) of the device, implicitly including the camera and / or display. Positioning and orientation can additionally or alternatively be inferred from video analysis. A pose update stream can be used to synchronize the physical scene 104 captured by the camera 120 with a 3D model 122 of the physical scene.

[0024] Figure 3The figure shows how the mixed reality system 150 constructs a 3D model 122 of the physical scene and draws views of the 3D model 122 corresponding to the pose of the camera relative to the physical scene 104. As described above, the mixed reality system 150 can include a camera and a display, as well as a pose detection mechanism (gyroscope, video analysis, radio triangulation, etc.). The mixed reality software 152 running on (or communicating with) the mixed reality system 150 performs two main functions 154 and 156.

[0025] First function 154 receives spatial (pose) data from the camera and / or display at the physical scene 104. It uses this information in a known manner to construct a 3D model of the physical scene. Generally speaking, first function 154 identifies features such as textures, lines, planes, surfaces, and feature points, adds representations of these features to the 3D model, and uses the corresponding spatial pose of the camera to determine where the features belong in the 3D model. In effect, the 3D model is anchored to the physical scene. Furthermore, points or objects added to the 3D model via software are effectively anchored to corresponding fixed positions and orientations in the physical scene 104. Such functionality can be performed using known tools, such as ARKit™ released by Apple Inc., the ARCore platform released by Google Inc., and the toolkit available in Visual Studio™ released by Microsoft Inc. First function 154 also maintains a virtual camera 158 (i.e., a view or viewpoint), whose pose in the 3D model 122 mirrors the pose of the camera / display in the physical scene 104.

[0026] The second function 156 draws graphics based on the 3D model and the pose of the virtual camera 158. The drawn view of the 3D model from the current pose of the virtual camera 158 will mirror the physical view "seen" by the camera / display. In addition, because any 3D objects, points, lines, etc. added to the 3D model by software are effectively anchored to the physical scene through the spatiotemporal synchronization between the camera / display and the virtual camera 158, the drawing of such 3D objects relative to the user's real-time view of the physical scene will continuously have a position, size, orientation, and perspective on the display that is consistent with the real-time view of the physical scene seen on or through the display. The mixed reality system can sometimes maintain the 3D model without displaying any drawing of the 3D model.

[0027] Figure 4A process for defining an animation path in a mixed reality presentation is shown. As described above, it is assumed that the mixed reality device is capturing video of a physical scene, forming a 3D model of the scene, and is capable of drawing views of the 3D model or elements in the 3D model. It is also assumed that the mixed reality system has a user input device that can input points in at least two dimensions relative to the display through which the mixed reality is viewed. The input device can be: a 3D pointer device, such as a laser / sensor or handheld pointer that reports its position and orientation; a 2D pointer device (such as a touch-sensitive layer of a display (mouse, etc.)); a system for detecting the direction of eye gaze, etc. In step 170, the user draws a 2D (at least) input path 171 on the display 108 or in a manner that allows the input path 171 to be relative to the display. The input path 171 can be a set of discrete input points, a stream of tightly sampled points interpolated to a linear path such as a B-spline, etc.

[0028] The input path 171 is input relative to the display space of the display 108. At step 172, the input path 171 is transformed into a 3D path 173 in the 3D model 122. In one embodiment, the input path 171 is input into the display while the video from the camera is being displayed, and at the same time, the pose of the camera is changing and the view of the physical view changes accordingly. The continuously updated mapping / transformation between the camera pose and the 3D model enables the input points of the input path to be consistently mapped to the 3D model. The points of the input path are mapped to the 3D model and projected from the virtual camera to find intersections with the 3D model. For example, if the input path is drawn as covering the surface of a cube or table ( Figure 4 ), the projection of the input path from the virtual camera intersects with the corresponding surface in the 3D model 122 ( Figure 4 ), and the intersection defines a 3D path 173.

[0029] In another embodiment, the display may only display a still frame captured by the camera of the physical scene. The camera pose corresponding to the capture of the frame is then used to project the input path 171 onto the 3D model to define the 3D path 173. In yet another embodiment, a video clip of the camera including a corresponding stream of camera pose data is played back on the display while the input path 171 is input, and the input path is mapped to the 3D model, as described in the above paragraph.

[0030] As can be seen, a variety of techniques can be used to map user input in the display space to a corresponding 3D path or point in the 3D model. Furthermore, the path can be tracked while a frozen, real-time, or playback view of the physical scene is being viewed or displayed. It should be noted that steps 170 and 172 need not be consecutive, discrete steps, but can be repeatedly performed as input path 171 is input. In other words, as input path 171 is input, input path 171 can be mapped to 3D path 173 in real time. Similarly, a graphical representation of input path 171 can be displayed as input path 171 is input.

[0031] At step 174, additional input is received for defining animations associated with the 3D path 173. Such input may include specification of the object to be animated, the object's motion during the animation, changes in the object's state, animation parameters such as speed, and the like. The animation definitions may be stored as part of the 3D model 122 or as a separate software component that interfaces with the 3D model and the mixed reality system. In one embodiment described further below, animation actions may be added to the path by input directed to the path, such as by a drag-and-drop action from a displayed tool palette. After a period of idle interaction with the path, the path may optionally be hidden (not displayed).

[0032] In step 176, the defined animation is executed in response to a triggering event. The triggering event may be an explicit user input, such as a button tap, activation of a user interface element, or a voice command. The triggering event may be the expiration of a timer that starts after the last edit of the defined animation. The triggering event may also be the satisfaction of a condition of the mixed reality system that may also be defined by the user. Conditions may be defined with respect to the 3D model, such as proximity of the camera / display to the 3D path or the surface on which the path lies, a threshold ratio of the 3D path being viewed or displayed, proximity of a physical object to the 3D path, or any other spatiotemporal condition. The triggering condition may also be external to the mixed reality system; for example, the occurrence of a time or date, a remote command, etc.

[0033] When the animation is displayed, a graphical representation of the 3D path may or may not be displayed. In one embodiment, the animation of a 3D object may include the translation of the object and the manipulation of the orientation of the 3D object. If the 3D object has a frontal or forward-facing orientation, the animation process may repeatedly redirect the 3D object as the 3D object translates along the 3D path. The 3D object may be oriented so that the forward direction of the 3D object is aligned with the direction of the path at its current point (or its tangent). Preferably, if the animated object models limb-based motion, the point at which the limb contacts the surface containing the 3D path may be fixed to the surface with a certain possible rotation. In other words, if the feet of the animated object need to stick to the ground, the feet can be manipulated separately according to the path rather than directly connecting the feet to the ground, which can avoid the slipping effect. The manipulation logic can calculate the angle between the forward vector of the animated object and the position of the next segment of the path. Similarly, as the object traverses the path, the shape of the 3D object can be transformed or deformed according to the 3D path.

[0034] Figure 5 A process for defining an animation path in a mixed reality presentation is shown. At step 190, the system receives 2D input points from the current physical view of the physical scene. Any known tool or technique for inputting 2D points or paths can be used. The 2D input points can be processed by the underlying graphics system of the display mixed reality system.

[0035] At step 192, the 2D input points are transformed into corresponding views of the 3D model according to the poses of the cameras corresponding to the points. Because 2D points have only two dimensions, at step 194, rays are cast from the virtual camera pose through the points to find intersections with the 3D model.

[0036] In step 196, the intersection points with the 3D model are used to construct a 3D path. In one embodiment, the 3D path is a sequence of segments connecting corresponding 3D points. In another embodiment, a heuristic method is used to select the surface in the 3D model that best fits the 3D points, and the 3D points are then checked to ensure that they lie on the surface; small discrepancies can be resolved and points far from the surface can be discarded. If the sequence of points crosses the edge of a surface, a gap can be inserted. Known techniques for reconstructing geometry from a point cloud can be used. In one embodiment, if the path is initially defined as a sequence of points that intersect a surface in the 3D model, the segments connecting the points can be constructed to lie on that surface.

[0037] Figure 6A user interface for interactively defining an animation associated with an animation path is shown. As described above, once the path is defined in mixed reality, the properties of the animation can be defined. A tool palette 200 is displayed on the display 108. The drawing of the 3D path 201 is also shown. The 3D path can be drawn for a single still frame or continuously drawn in real time based on the pose of the camera. Input can be directed to the drawing of the 3D path 201 just as the input of the points defining the original 2D path can be; the input points can be mapped back to the 3D model in a similar manner. In other words, the mixed reality system enables user input to specify points on the 3D path in the 3D model.

[0038] In one embodiment, points are specified by dragging and dropping animation nodes 202 from the tool palette 200 onto the drawing of the 3D path. Each animation node 202 represents a different animation action, for example, "run," "jump," "pause," "accelerate," or any other type of animation effect. Script nodes may also be provided. When a script node is added to a path, the user can enter text for script commands that are to be interpreted and executed during the animation. There may be global animation nodes 204 that apply to any animation. There may also be object-specific animation nodes 206 associated with an animation object (or its class) that the user has associated with the path. Animation nodes may specify state changes for an animation object to change the inherent motion or appearance of the animation object, trigger an action for the object, modify the audio of the object, and the like.

[0039] In one embodiment, an animation object is specified by selecting an animation object node 208 or dragging an animation object node onto a path. Figure 6 In the example of , a set 210 of animation object representations is provided, each animation object representation representing a different animation object, such as a character, a creature, a vehicle, etc.

[0040] In another embodiment, a pop-up menu can be used. User input pointing to a point on the path causes the menu to be displayed. When a menu item is selected, the action or object represented by the selected menu item is added to the path at the input point that called the menu.

[0041] Figure 7 The process for executing an animation path is shown. The animation is executed by a control loop. The control loop iterates for small animation steps, such as once per animation frame, once every X milliseconds, etc. The sequence of loop iterations drives the update of the 3D model, which is animated on the mixed reality display 108 as it is drawn. If the animation is executed in real time, the drawing from the virtual camera's view tracks the physical camera and therefore remains aligned with the physical scene seen on or through the display. Alternatively, the animation can be executed in advance and layered onto a corresponding video clip captured by the camera. The animation is then seen during video playback.

[0042] At step 220, the animation iteration begins. At step 222, the length of the path segment is calculated for the current velocity or acceleration of the animated object traversing the path. At step 224, the path segment is tested for any animation commands or nodes. Any animation commands or instructions detected are executed or evaluated in the order in which they appear on the path. At step 226, the animation for the path segment is executed by the rendering engine based on the current animation parameters. Any commands on the path segment are executed by appropriately updating the 3D model (including the movement of the animated object). If the animated object has its own in-place animation, the cycle is repeated as the character is inserted along the length of the path. If the object does not have in-place animation, the object is rotated to align with each segment along the path and inserted along its length. The animation loop continues until the object's traversal along the path is complete.

[0043] Figure 8 An example simulation of animation is shown. The 3D model 122 includes a reconstructed scene geometry 250 and a 3D path 252. A 3D rendering engine 254 renders a view of the 3D model from the perspective of a virtual camera. The 3D rendering engine 254 includes (or has an interface with) a path interpreter 256. The path interpreter 256 performs Figure 7 The animation loop continues. The 3D rendering engine 254 can also update the pose of the virtual camera based on the pose 258 of the camera or mixed reality device that includes the camera. As the animation begins, the path interpreter repeatedly updates the position and orientation of the animated object. The path interpreter iterates to detect and execute the first object action. This continues as commands and actions are encountered along the path until the end of the path is reached. The 3D engine continues to render the object's motion according to the path.

[0044] While a 3D path serves as an anchor point for the movement of an associated 3D object, the 3D object need not strictly follow the path as if attached to a track. The path guides the movement of the 3D object, and the path the 3D object takes through the space of the 3D model can vary. For example, commands, scripts, or object behaviors can control the distance and / or angle of the object relative to the path, allowing the object to float, sink, bend, and so on.

[0045] While a real-time embodiment has been described above, it should be noted that the animation techniques can be used for playback of any video with a suitable 3D model, where 2D input points entered in a 2D display space can be mapped back to points in the 3D model.

[0046] While it may be convenient to heuristically map or project 2D input points defined by a path onto corresponding surfaces in a 3D model, paths can also be input in three dimensions in physical space. For example, if the mixed reality system has a three-dimensional pointer that allows manual control of both the direction and distance of the input point, the user can directly specify a 3D path in the 3D model. Similarly, a pointer input device that projects rays (e.g., light, radio, sound) can specify 3D input points in a physical scene that can be directly converted into a 3D model. Such an input device can allow a user to specify points on a physical surface in a physical scene, which can be converted into corresponding points on corresponding surfaces in a 3D model.

[0047] Although there are mixed reality systems that provide software that facilitates the transition between physical space and modeled 3D space, path definition and mapping can be accomplished with custom-coded plane detection and object alignment. Plane detection in camera video output can be performed using markerless RGB (red, green, blue) tracking. Many algorithms can be used to achieve this. In one embodiment, multiple features are extracted per frame in high-contrast areas of the video. Those high-contrast areas are matched across frames, allowing drift to be detected across frames. The transformation of the surface (e.g., plane) where the object will be placed can then be calculated. Once the target surface is detected, the transformation is placed on a point on the plane and rotated to align with the normal of the surface. The 3D object and path are then parented to the transformation.

[0048] The speed of animation along a path (traversal and / or effects) can be determined in several ways. In one embodiment, the length of the path is synchronized with the length of the input video, which is a segment of real-time video defined by the duration of the input path, or a segment of recorded video. The duration or length of the path is tracked to the video so that the animation appears to be locked to the world. In another embodiment, if such information is available, the speed can be keyed to the real-world measurement of the path. User interface elements can be provided to allow the speed, duration, or timing of the animation to be adjusted. Graphic nodes or commands can be added to the animation path, such as at the beginning and end, to control speed, time, duration, etc.

[0049] Complex animation presentations can be created by allowing multiple animation paths to be defined. Complex networks of possibly intersecting paths can allow the display of multiple animated objects that may interact with each other.

[0050] Figure 9Details of a computing device 302 on which the above-described embodiments may be implemented are shown. The computing device 302 is an example of a client / personal device or a backend physical (or virtual) server device that can perform various (or perhaps most) of the processes described herein. The technical disclosure herein will satisfy the needs of programmers writing software and / or configuring reconfigurable processing hardware (e.g., field programmable gate arrays (FPGAs)) and / or designing application-specific integrated circuits (ASICs) and the like to run on the computing device 302 (possibly via a cloud API) to implement the embodiments described herein.

[0051] Computing device 302 may have one or more displays 108, a camera (not shown), a network interface 324 (or several), storage hardware 326, and processing hardware 328, which may be any one or more of the following: a central processing unit, a graphics processing unit, an analog-to-digital converter, a bus chip, an FPGA, an ASIC, an application-specific standard product (ASSP), or a complex programmable logic device (CPLD). Storage hardware 326 may be any combination of magnetic storage, static memory, volatile memory, non-volatile memory, optical or magnetically readable media, and the like. The term "storage" as used herein does not refer to signals or energy per se, but rather to physical devices and states of matter. The hardware elements of computing device 302 may cooperate in a manner well known in the art of machine computing. In addition, input devices may be integrated with computing device 302 or communicate with computing device 302. Computing device 302 may have any form factor or be used in any type of contained device. Computing device 302 may be in the form of a handheld device such as a smartphone, a tablet computer, a gaming device, a server, a rack-mounted or backplane-mounted computer, a system-on-chip, and the like.

[0052] The embodiments and features discussed above can be implemented in the form of information stored in volatile or non-volatile computer or device readable storage hardware. This is considered to include at least hardware such as optical storage (e.g., compact disk read-only memory (CD-ROM)), magnetic media, flash memory, read-only memory (ROM), or any device in which digital information is stored so that processing hardware 328 can easily access it. The stored information can take the form of machine executable instructions (e.g., compiled executable binary code), source code, bytecode, or any other information that can be used to enable or configure a computing device to perform the various embodiments discussed above. This is also considered to include at least volatile memory (such as random access memory (RAM) and / or virtual memory) that stores information (such as central processing unit (CPU) instructions) during the execution of the program that performs the embodiment, and non-volatile media that stores information that allows programs or executable files to be loaded and executed. The embodiments and features can be executed on any type of computing device, including portable devices, workstations, servers, mobile wireless devices, etc.

Claims

1. A method performed by a computing device comprising processing hardware and storage hardware, the method comprising: receiving a video of a physical scene from a camera, the video being captured while tracking a pose of the camera relative to the physical scene, the pose including a position and orientation of the camera relative to the physical scene; Analyzing the captured video to construct a 3D model of the physical scene, and mapping the camera pose to a virtual pose including a virtual position and orientation within the 3D model, the virtual pose including a first virtual pose and a second virtual pose; displaying the captured first segment of the video on a display; receiving, while displaying the first segment of the captured video, a curve generated by sensing user manipulation of a pointer device corresponding to the curve, the pointer device being used for manual input of an input point, and the curve including a 2D position relative to the display and provided by the pointer device; generating the 3D path by mapping the 2D position of the curved line to a position for a 3D path in the 3D model according to the first virtual pose as the 2D position is input and received, wherein as the 3D path is generated, drawing the 3D path as a graphical path within a portion of the first segment of the captured video such that the graphical path is displayed as and corresponds to the input of the curved line path, wherein as the curved line path is progressively input, the graphical path is progressively displayed; modeling automated movement of a 3D object according to the 3D path in the 3D model during capture of a second segment of the captured video corresponding to the second virtual pose; as well as A composite video is displayed on the display, the composite video comprising the second segment of the captured video composited with a rendering of the 3D model according to the second virtual gesture, the rendering comprising the modeled movement of the 3D object in the 3D model.

2. The method of claim 1 , further comprising associating an animation action with an additional 3D point on the 3D path based on additional 2D positions on the graphical path, wherein the additional 2D positions are mapped to the additional 3D points based on the first virtual gesture, and wherein the additional 2D positions are input by the user using the pointer device before displaying the drawing of the 3D object. 3 . The method of claim 2 , wherein the automatic movement of the 3D object is modeled according to the animation action as the object reaches the additional 3D points on the 3D path. 4 . The method of claim 1 , wherein mapping the 2D position to the position for the 3D path comprises casting a ray from the 2D position to find an intersection with the 3D model.

5. The method according to claim 1, further comprising: Surfaces in the captured video are identified, corresponding surfaces are added to the 3D model, and the positions for the 3D path are placed according to the surfaces in the 3D model.

6. The method of claim 1 , further comprising displaying a graphical user interface including graphical representations of corresponding animation commands, and enabling dragging and dropping of the graphical representations onto the graphical path to specify the animation commands with respect to positions on the 3D path.

7. The method of claim 1, wherein receiving the input points, modeling the automatic movement of the 3D object, and displaying the composite video are performed in real time.

8. A computing device comprising: Processing hardware; monitor; camera; as well as Storage hardware storing instructions configured to cause the processing hardware to perform a process, the process comprising: continuously receiving video of a physical scene from the camera, and constructing a 3D model of the physical scene based on the received video of the physical scene; continuously receiving gesture updates corresponding to changing a physical gesture of the display; receiving, concurrently with the gesture updates and the video, a curved input path of interactive input generated by a first 2D user input point input using a pointer device, and progressively displaying the curved input path on the display corresponding to the progressively input 2D user input point, wherein the curved input path is progressively displayed by progressively mapping the curved input path to a 3D path in the 3D model as the first 2D user input point is progressively input; receiving a second 2D user input point inputted using the pointer device corresponding to the displayed curved input path, the second 2D user input point mapping an animation action to a point on the 3D path in the 3D model; and After mapping the animation actions to the points on the 3D path, animating a 3D object following the 3D path in the 3D model based on the pose updates, and displaying the rendering on the display, wherein the 3D object is rendered to perform the actions as the 3D object reaches the corresponding points on the 3D path.

9. The computing device of claim 8, wherein the gesture updates are obtained from a motion sensor of the computing device and / or from analysis of the video.

10. The computing device of claim 8, wherein the display comprises a transparent material on which the rendering of the animation is displayed, and wherein the animation in the 3D model is updated in real time consistent with the pose updates from a viewpoint from which it was rendered.

11. The computing device of claim 8, wherein the input path is input on the display while the display is displaying the segment of the video.

12. The computing device of claim 11, wherein the segments of the video comprise still frames from the video.

13. The computing device of claim 8, wherein the display displays the video from the camera in real time while displaying the rendering of the animation.

14. The computing device of claim 8, wherein the animation comprises the 3D object in the 3D model following the 3D path in the 3D model.

15. The computing device of claim 14, wherein: In response to determining that the object has reached one of the points on the 3D path during the animation, the animation is modified as specified by the animation action corresponding to the one point.

16. Computer-readable storage hardware storing instructions configured to cause a computing device to perform a process comprising: executing a mixed reality system that constructs a 3D model of a physical scene and maintains a mapping between the 3D model and a physical pose of a display for displaying a rendering of the 3D model; receiving a curved line from a pointer device manipulated by a user, the user tracing a curved path with the pointer device to progressively input the curved line, wherein the curved line corresponds to the curved path; progressively displaying a graphic path coincident with the curve on the display by progressively adding a 3D path corresponding to the curve to the 3D model, wherein the graphic path is displayed as the curve is input by the pointing device; After displaying the graphical path, receiving an input point associated with the graphical path, and associating an animation action with the 3D path based on the input point, wherein the input point is input by the user manipulating the pointer device; as well as While the mixed reality system updates the mapping between the 3D model and the physical pose of the display, and after adding the 3D path to the 3D model, an animation of a 3D object transitioning relative to the 3D model is drawn according to the 3D path, and the drawing is displayed on the display.

17. The computer-readable storage hardware of claim 16, wherein the input points are mapped to the 3D model according to one or more metrics of the physical pose.

18. The computer-readable storage hardware of claim 16, wherein the 3D object transforming relative to the 3D model comprises: The 3D object is repeatedly reoriented relative to the 3D model according to the orientation of the 3D path at the corresponding location of the 3D object.

19. The computer-readable storage hardware of claim 18, wherein as the 3D object transforms, the 3D object is offset from the 3D path.

20. The computer-readable storage hardware of claim 16, wherein the animation is defined and rendered in real time while the mixed reality system executes.

Citation Information

Patent Citations

  • Method and system for generating motion sequence of animation, and computer-readable recording medium

    CN104969263A

  • Method and System for Compositing an Augmented Reality Scene

    US20120069051A1