Content playback and modification in 3D environments

CN116457883BActive Publication Date: 2026-08-18APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180076652.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-14
Filing Date
2021-08-10
Publication Date
2026-08-18
Estimated Expiration
2041-08-10

AI Technical Summary

Technical Problem

开发和调试在3D环境中使用的应用程序内容可能是耗时且困难的

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116457883B_ABST
    Figure CN116457883B_ABST
Patent Text Reader

Abstract

Various implementations disclosed herein include devices, systems, and methods of presenting playback of application content within a three-dimensional (3D) environment. An example process presents, within a 3D environment, a first set of views of a scene that includes application content provided by the application. The first set of views is provided from a first set of viewpoints during execution of the application. The process records execution of the application based on recording program state information and changes to the application content determined from user interactions, and presents a second set of views that includes playback of the application content within the 3D environment based on the recording. The second set of views is provided from a second set of viewpoints that is different from the first set of viewpoints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates in its entirety to systems, methods, and apparatus for presenting views of application content in a three-dimensional (3D) environment on an electronic device, and more particularly to providing views that include playback of application content within a 3D environment. Background Technology

[0002] Electronic devices can present 3D environments that include content provided by applications. For example, an electronic device can provide a view of application content within an environment that includes a three-dimensional (3D) representation of a user's living room. The appearance of the user's living room can be provided based on optical perspective techniques, pass-through video techniques, etc., using 3D reconstruction. Developing and debugging application content used in 3D environments can be time-consuming and difficult. Existing technologies may not easily detect, audit, and correct unwanted application activity and other problems. For example, identifying and resolving unintended application behavior (such as a physics engine causing virtual objects to move unintended) can involve an undesirable workload of application testing and retesting. Summary of the Invention

[0003] The various specific implementations disclosed herein include devices, systems, and methods for recording and playing back the use / testing of an application that provides content in a three-dimensional (3D) environment (e.g., computer-generated reality (CGR)) based on recorded program states, parameter values, changes, etc., for debugging purposes within the 3D environment. For example, a developer can use a device (e.g., a head-mounted display (HMD)) to test an application and, in the test, roll a virtual bowling ball into virtual bowling pins, then rewind the recording of the test to observe why one of the pins did not respond as expected. In an exemplary use case where an application interacts / synchronizes with changes in a system application, recording may involve capturing / reusing those changes and writing them to a video file—that is, recording a snapshot of the scene content and a video file containing any changes that may have occurred within it. Additionally, the system can reconstruct sound (e.g., spatial stereo audio) and other details of the application for playback and review (e.g., a complete snapshot of the application).

[0004] Generally speaking, an innovative aspect of the subject matter described in this specification can be embodied in a method comprising the following actions: during the execution of an application, presenting a first set of views including application content provided by the application in a three-dimensional (3D) environment, wherein the first set of views is provided from a first set of viewpoints in the 3D environment; during the execution of the application, generating a video recording of the application's execution based on recorded program state information and changes in application content determined according to user interaction; and presenting a second set of views including playback of application content in the 3D environment based on the video recording, wherein the second set of views is provided from a second set of viewpoints different from the first set of viewpoints.

[0005] These and other implementation schemes may optionally include one or more of the following features.

[0006] In some respects, communication between the application and system processes during application execution includes changes to the application content, where a second view set is generated based on the communication.

[0007] In some aspects, generating a video recording includes: obtaining recorded program state information corresponding to the state of application content at multiple points in time; obtaining changes in the application content, wherein the changes include changes in the application content that occur between states; and generating a video recording of the execution of the application based on the recorded program state information and the changes in the application content from communications between the application and system processes.

[0008] In some respects, the application content includes objects and the changes include incremental values ​​of changes in the position of the objects.

[0009] In some respects, the method also includes receiving input of a point in time during the execution of the selected application as the starting point for playback.

[0010] In some aspects, presentation playback includes a graphical depiction of the user's head position, gaze direction, or hand position while executing the application.

[0011] In some aspects, playback presentation includes presenting a graphical depiction of the sound source.

[0012] In some aspects, the view of the scene includes a video passthrough or perspective image that presents at least a portion of the physical environment, wherein a 3D reconstruction of at least a portion of the physical environment is dynamically generated during the execution of the application, and the presentation of playback includes presenting the 3D reconstruction.

[0013] In some respects, during application execution, objects within the application content are located based on a physics engine, and during playback, objects within the application content are located based on determining the location of objects according to program state information and relocating objects as needed.

[0014] In some respects, during the execution of the application, a view of the scene is presented on a head-mounted device (HMD).

[0015] According to some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of the apparatus, cause the apparatus to perform or cause to perform any of the methods described herein. According to some embodiments, an apparatus includes: one or more processors, non-transitory memory, and means for performing or causing to perform any of the methods described herein. Attached Figure Description

[0016] Therefore, this disclosure will be understood by those skilled in the art, and a more detailed description can be made with reference to some exemplary embodiments, some of which are shown in the accompanying drawings.

[0017] Figure 1 This is based on some specific implementation examples of operating environments.

[0018] Figure 2 This is a flowchart representation of an exemplary method for recording and playing back application objects within a three-dimensional (3D) environment, based on some specific implementations.

[0019] Figure 3 An exemplary presentation of a view showing a snapshot of the playback of application content within a 3D environment, according to some specific implementation, is shown.

[0020] Figure 4 An exemplary presentation of a view showing a snapshot of the playback of application content within a 3D environment, according to some specific implementation, is shown.

[0021] Figure 5 An exemplary presentation of a view showing a snapshot of application content playback within a 3D environment, based on some specific implementation, is shown.

[0022] Figure 6 These are exemplary devices based on some specific implementations.

[0023] As is customary, the various features shown in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Additionally, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Detailed Implementation

[0024] This document provides numerous specific details to enable those skilled in the art to thoroughly understand the claimed subject matter. However, the claimed subject matter can be practiced without these details. In other instances, methods, apparatus, or systems known to those of ordinary skill in the art have not been described in detail so as not to obscure the claimed subject matter.

[0025] Figure 1 This is a block diagram of an exemplary operating environment 100 according to some specific implementation. In this example, the exemplary operating environment 100 illustrates an exemplary physical environment 105 including a table 130, a chair 132, and application objects 140 (e.g., virtual objects). Although relevant features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure further relevant aspects of the exemplary specific implementations disclosed herein.

[0026] In some embodiments, device 120 is configured to present an environment to user 102. In some embodiments, device 120 is a handheld electronic device (e.g., a smartphone or tablet). In some embodiments, user 102 wears device 120 on their head. Device 120 may include one or more displays provided for displaying content. For example, device 120 may surround user 102's field of vision.

[0027] In some implementations, the functionality of device 120 is provided by more than one device. In some implementations, device 120 communicates with a separate controller or server to manage and coordinate the user experience. Such a controller or server may be local or remote relative to the physical environment 105.

[0028] According to some specific implementations, device 120 can generate and present a computer-generated reality (CGR) environment to its corresponding user. A CGR environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic systems. In a CGR, a subset of a person's physical motion, or a representation thereof, is tracked, and in response, one or more characteristics of one or more virtual objects simulated in the CGR environment are adjusted in a manner consistent with at least one physical law. For example, a CGR system can detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of characteristics of virtual objects in the CGR environment can be done in response to a representation of physical motion (e.g., a voice command).

[0029] Humans can use any of their senses to sense and / or interact with CGR objects, including sight, hearing, touch, taste, and smell. For example, a person can sense and / or interact with audio objects that create (3D) or spatial audio environments, providing the perception of a point audio source in 3D space. As another example, audio objects can enable audio transparency, selectively introducing ambient sound from the physical environment, with or without computer-generated audio. In some CGR environments, a person can sense and / or interact only with audio objects. In some implementations, image data is registered with image pixels of the physical environment 105 (e.g., RGB, depth, etc.) used with imaging processing techniques within the CGR environment described herein.

[0030] Examples of CGR include virtual reality and mixed reality. A virtual reality (VR) environment is a simulated environment designed to provide one or more senses entirely based on computer-generated sensory input. A VR environment includes virtual objects that a person can sense and / or interact with. For example, trees, buildings, and computer-generated images representing human avatars are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through the simulation of a person's presence within the computer-generated environment and / or through the simulation of a subgroup of physical movements of a person within the computer-generated environment.

[0031] Compared to VR environments, which are designed to be entirely based on computer-generated sensory input, mixed reality (MR) environments are simulated environments designed to incorporate sensory input from the physical environment, or representations thereof, in addition to computer-generated sensory input (e.g., virtual objects). On the virtual continuum, a mixed reality environment is any state between a purely physical environment as one end and a virtual reality environment as the other end, but not including either end.

[0032] In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. Additionally, some electronic systems used to present the MR environment can track position and / or orientation relative to the physical environment, enabling virtual objects to interact with real objects (i.e., physical objects or representations of them from the physical environment). For example, the system can cause movement so that virtual trees appear stationary relative to the physical ground.

[0033] Examples of mixed reality include augmented reality and augmented virtual. An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are overlaid on a physical environment 105 or a representation thereof. For example, an electronic system for presenting an AR environment may have a transparent or semi-transparent display through which a person can directly view the physical environment 105. The system may be configured to present virtual objects on a transparent or semi-transparent display, allowing a person to perceive the virtual objects overlaid on the physical environment 105 using the system. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment 105, which are representations of the physical environment 105. The system combines the images or videos with virtual objects and presents the combination on the opaque display. A person uses the system to indirectly view the physical environment 105 via the images or videos of the physical environment 105 and perceives the virtual objects overlaid on the physical environment 105. As used herein, video of the physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment 105 and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects onto the physical environment 105, for example, as a hologram or onto a physical surface, so that a person can use the system to perceive the virtual objects superimposed on the physical environment 105.

[0034] Augmented reality environments also refer to simulated environments in which the representation of the physical environment 105 is transformed by computer-generated sensory information. For example, in providing pass-through video, the system can transform one or more sensor images to apply a selected viewpoint (e.g., viewpoint) different from the viewpoint captured by the imaging sensor. As another example, the representation of the physical environment 105 can be transformed by graphically modifying (e.g., zooming in) portions of it, such that the modified portions can be representative but not realistic versions of the original captured image. Furthermore, the representation of the physical environment 105 can be transformed by graphically removing or blurring portions of it.

[0035] An augmented virtual (AV) environment refers to a simulated environment in which a virtual or computer-generated environment is combined with one or more sensory inputs from a physical environment 105. Sensory inputs can be representations of one or more characteristics of the physical environment 105. For example, an AV park could have virtual trees and virtual buildings, but a person's face could be realistically reproduced from an image taken of a physical person. Similarly, virtual objects could adopt the shape or color of a physical object imaged by one or more imaging sensors. Furthermore, virtual objects could adopt shadows that correspond to the sun's position within the physical environment 105.

[0036] Many different types of electronic systems enable people to sense and / or interact with a variety of CGR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. Head-mounted systems may have one or more speakers and an integrated opaque display. Alternatively, head-mounted systems may be configured to receive an external opaque display (e.g., a smartphone). Head-mounted systems may incorporate one or more imaging sensors for capturing images or video of the physical environment, and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one implementation, a transparent or translucent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection technology, which projects graphic images onto the human retina. Projection systems can also be configured to project virtual objects onto a physical environment, such as as holograms or on a physical surface.

[0037] Figure 2 This is a flowchart representation of an exemplary method 200 for recording and playing back an application object in a specific implementation of a 3D environment (e.g., CGR). In some implementations, method 200 is performed by a device (e.g., Figure 1 The method 200 is performed by a device 120, such as a mobile device, desktop computer, laptop computer, or server device. In some embodiments, the device has a screen for displaying images and / or a screen for viewing stereoscopic images, such as a head-mounted display (HMD). In some embodiments, method 200 is performed by a processing logic component (including hardware, firmware, software, or a combination thereof). In some embodiments, method 200 is performed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory). The content recording and playback process of method 200 is referenced. Figures 3 to 5 illustrate.

[0038] At box 202, method 200 renders a first set of views, including application content provided by the application, within a 3D environment during application execution, wherein the first set of views is provided from a first set of viewpoints in the 3D environment. For example, during testing of the application in an MR environment, the user is executing the application (e.g., as...). Figures 3 to 4 (The virtual drawing application shown). The first viewpoint set may be based on the user device's position during execution. In exemplary embodiments, for example, the virtual drawing application may overlay video passthrough, optical perspective, or a virtual room. In some embodiments, the application may define the weight of a virtual bowling ball, and the virtual bowling ball may roll based on user hand movements and / or interactions with input / output (I / O) devices. In some embodiments, a view of the scene is rendered on the HMD during application execution.

[0039] At box 204, method 200 generates a recording of application execution based on recorded program state information and changes in application content determined by user interactions during application execution. For example, recorded program state information may include recording scene understanding or snapshots, such as the position of objects in the environment. In some implementations, scene understanding may include head pose data, what the user is looking at in the application (e.g., virtual objects), and hand pose information. Additionally, recording scene understanding may include recording a scene understanding 3D mesh generated concurrently during program execution.

[0040] In some implementations, the reality engine architecture can synchronize changes between the application and the system-wide application. In some implementations, changes and / or increments transmitted between the application and the system-wide application can be captured for recording. Capturing changes and / or increments transmitted between the application and the system-wide application allows recording to be performed with minimal impact on application performance. A debugger (e.g., an automated debugger application or a system engineer debugging an application) can receive application-to-system changes and write them to video files (e.g., video files containing scene content / snapshots at the scene level and changes between snapshots).

[0041] At box 206, method 200 presents a second set of views, including playback of application content within a 3D environment based on video recording. The second set of views is provided from a second set of views different from the first set of views. Alternatively, the second set of views may be the same viewpoint as the first set of views (e.g., a debugger has the same viewpoint as the application user). The second set of views may be based on the user device's position during playback.

[0042] In some implementations, the 3D environment (e.g., scenes and other application content) can be rendered continuously / live-stream throughout execution and playback. That is, the rendering engine can run continuously for one period of time to inject the executed application content and for another period of time to inject the recorded application content. In some implementations, playback may differ from simply reconstructing the content in the same way as it was initially generated. For example, playback may involve using recorded values ​​of the ball's position instead of having the ball use a physics system.

[0043] Users (e.g., programmers / debuggers) can pause testing and use the cleanup tool to return to see the desired point in time. In some implementations, users can replay from the same viewpoint, or alternatively, users can change the viewpoint and see where the HMD is depicted, such as based on gaze direction. In some implementations, users can change the viewpoint to observe scene understanding (e.g., head position, hand position, 3D reconstructed mesh, etc.). In some implementations, developers can return to enable the display of representations of sound sources (e.g., spatial stereo audio) and other invisible items. In some implementations, developers can add data tracking.

[0044] In some implementations, method 200 involves focusing on reusing changes / increments (e.g., scene understanding) of snapshots already used in the application-to-system reality architecture, enabling efficient and accurate recording with minimal performance impact. In one exemplary implementation, communication between the application and a system process (e.g., a debugger / system application) during application execution includes changes to application content, with a second view set generated based on the communication. For example, a system application (e.g., a debugger application) is responsible for the system process. The system process may include rendering, presenting different viewpoints, etc. This exemplary implementation presents a reality engine architecture where changes between the application and the system application are synchronized. For example, the system application may display content / changes provided by multiple applications. In some implementations, the system application may provide environmental content (e.g., walls, floors, system functionality, etc.) and application content within an environment that a user (e.g., a debugger or programmer) can interact with, both system application components and application components.

[0045] In some specific implementations, generating a recording includes: obtaining recorded program state information corresponding to the state of application content at multiple points in time (e.g., the debugger application taking periodic snapshots); obtaining changes in the application content; and generating a recording of the application's execution based on the recorded program state information and changes in the application content from communications between the application and system processes. Changes may include changes in application content that occur between states. For example, changes may include parameter changes, increases / decreases, position changes, new values, new positions (such as user head poses, hand poses, etc.) that change between each snapshot.

[0046] In some implementations, method 200 also involves receiving input of a point in time during the execution of the selected application as the starting point for playback. For example, a cleanup tool could be used to pause and return to an earlier point in time. In some implementations, the application content includes objects and the changes include incremental values ​​of changes in the position of the objects.

[0047] In some implementations, presentation playback includes a graphical depiction of the user executing the application's head position, gaze direction, and / or hand position. In some implementations, presentation playback includes a graphical depiction of a sound source. For example, for spatial stereo audio, a sound icon can be used to show the programmer the location of the spatial stereo sound source at a specific time during playback.

[0048] In some implementations, the view of the scene includes a video passthrough or perspective image presenting at least a portion of the physical environment, and the playback includes presenting a 3D reconstruction. In some implementations, a 3D reconstruction of at least a portion of the physical environment is dynamically generated during application execution. For example, this allows developers to see errors in the application caused by the 3D reconstruction not being completed at that point in time.

[0049] In some implementations, method 200 involves playback that differs based on usage (e.g., execution). That is, it focuses on using the recorded location rather than a physics engine to determine the user's location at a specific point during playback. In one exemplary implementation, during the execution of the application, objects of the application content are located based on a physics engine; and during playback, objects of the application content are located based on determining the object's location according to program state information and relocating the object as needed.

[0050] Reference here Figures 3 to 5 Further description includes the presentation of a second view set based on the playback of application content recorded within a 3D environment, wherein the second view set is provided from a second view set different from the first view set (e.g., based on the user's device position during playback). Specifically, Figure 3 and Figure 4An example is shown of a view presented to a programmer (e.g., a debugger) who is watching a playback / recording of an application (e.g., a virtual application overlaid on a physical environment, i.e., pass-through video) interacting with the application. Figure 5 A system flowchart is shown, illustrating the playback of application content in a 3D environment based on recordings using the techniques described herein.

[0051] Figure 3 An exemplary environment 300 is shown, presenting a snapshot representation of the physical environment (e.g., pass-through video) and a user viewpoint of application 320 (e.g., a virtual drawing application) from different viewpoints (e.g., a debugger's viewpoint). Application 320 includes drawing tools 322. The different viewpoints originate from the debugger's viewpoint, and, depending on some specific implementation, a debugger's cleanup tool 330 (e.g., a virtual playback controller) is presented to the debugger at a certain point in time.

[0052] As used herein, a "user" is a person who uses application 320 (e.g., a virtual drawing application), and a "debugger" is a person who uses a system application and the techniques described herein to rewind application 320 (from another viewpoint shown, or alternatively from the same viewpoint) to a snapshot at a specific point in time. The user and the debugger can be the same person or different people. Thus, for example, Figure 3 A debugger's perspective is shown on a device 310 (e.g., HMD) that views the user's perspective via a virtual application (e.g., application 320), which overlays or places within a 3D representation of real-world content (e.g., a passthrough video of the user's bedroom) in the physical environment 305.

[0053] like Figure 3 As shown, device 310 (e.g., Figure 1 Device 310 displays the user's head posture and gaze direction. For example, left-eye gaze 312 and right-eye gaze 314 are detected by device 310 (e.g., using an inward-facing camera, computer vision, etc.) and displayed within the debugger application, allowing the debugger to determine what the user's gaze is focused on during a playback snapshot. For example, some virtual applications utilize a person's gaze to provide additional functionality to a specific program being used (e.g., application 320). Centerline gaze 316 indicates the user's head posture. In some implementations, centerline gaze 316 may include the user's average gaze direction based on left-eye gaze 312 and right-eye gaze 314. Alternatively or additionally, centerline gaze 316 may be based on sensors in device 310, such as an inertial measurement unit (IMU), accelerometer, magnetometer, gyroscope, etc.

[0054] In some implementations, the debugger's viewpoint may be shown as the same as the user's viewpoint, or it may be illustrated from a different viewpoint. Therefore, the view presented to the debugger illustrates the user's situation during the execution of application 320. During the execution of application 320, the system application (e.g., the debugger application) records program state information and changes in application content determined based on user interactions during application execution. For example, recording program state information may include recording scene understanding or snapshots, such as the location of objects in the environment (e.g., virtual objects within application 320). In some implementations, scene understanding may include head pose data, what the user is looking at in the application (e.g., virtual objects), and hand pose information. Additionally, recording the scene understanding mesh may include recording a scene understanding 3D mesh generated concurrently during program execution.

[0055] Additionally, scene understanding may include recording data other than visual data. For example, spatial stereo audio may be part of application 320. Thus, the system application can play back the spatial stereo audio generated by application 320. In some implementations, visual elements (e.g., virtual icons) may be presented at the debugger's viewpoint to indicate the location (e.g., 3D coordinates) from which the spatial stereo audio originated at that moment during playback.

[0056] In some implementations, the reality engine architecture can synchronize changes between the application (e.g., application 320) and the system-wide application (e.g., a debugger application). In some implementations, changes and / or increments transmitted between the application and the system-wide application can be captured for recording. Capturing changes and / or increments transmitted between the application and the system-wide application allows recording to be performed with minimal impact on the application's performance. The debugger (e.g., an automated debugger application or a system engineer debugging an application) can receive application-to-system changes and write them to video files (e.g., video files containing scene content / snapshots at the scene level and changes between snapshots).

[0057] In some implementations, the 3D environment (e.g., scenes and other application content) can be rendered continuously / live-stream via the Cleaner tool 330 throughout execution and playback. That is, the rendering engine can run continuously for one time period to inject the executed application content and for another time period to inject the recorded application content. In some implementations, playback may differ from simply reconstructing the content in the same way it was initially generated. For example, playback might involve using recorded values ​​of the ball's position (e.g., 3D coordinates) instead of having the ball use a physics system (e.g., in a virtual bowling application). That is, the user (e.g., a programmer / debugger) can pause the test and use the Cleaner tool 330 to return to see the desired point in time (e.g., to see what debug events / bugs occurred, as referenced herein). Figure 4 (Further discussion).

[0058] In some specific implementations, generating a recording includes: obtaining recorded program state information corresponding to the state of application content at multiple points in time (e.g., the debugger application taking periodic snapshots); obtaining changes in the application content; and generating a recording of the application's execution based on the recorded program state information and changes in the application content from communications between the application and system processes. Changes may include changes in application content that occur between states. For example, changes may include parameter changes, increases / decreases, position changes, new values, new positions (such as user head poses, hand poses, etc.) that change between each snapshot.

[0059] Figure 4 An exemplary environment 400 is shown, illustrating a snapshot representation of the physical environment according to some specific implementation (e.g., pass-through video), a user viewpoint presenting an application (e.g., a virtual drawing application) from different viewpoints (e.g., a debugger's viewpoint), and a cleanup tool (e.g., a virtual playback controller) for the debugger at a specific point in time (e.g., during debugging an event such as an error). Thus, for example, Figure 4 This illustrates a debugger's perspective of viewing a virtual application (e.g., application 320) from the user's point of view, which is overlaid or placed within a representation of real-world content (e.g., a paved video of the user's bedroom). However, Figure 4This illustrates a replay where the debugger pauses application execution at a specific time (e.g., when a debug event occurs). For example, at a specific time, debug event 420 (e.g., a "short-duration pulse interference") occurs during the execution of application 320. The debugger can then use the cleaner tool 330 via a system application (e.g., the debugger application) to replay the recording of application 320's execution and stop the playback at any specific time. As the debugger interacts with the cleaner tool 330 (e.g., clicking the pause button when debug event 420 occurs), the hand icon 410 is a 3D representation of the debugger's hand in the system environment.

[0060] In some implementations, the debugger has the opportunity to view the playback from different viewpoints. For example, the debugger can use the system application and hand icon 410 to capture the current snapshot and drag the perspective view relative to the debugger. For example, environment 400 displays a specific viewpoint relative to the debugger from specific 3D coordinates (e.g., x1, y1, z1). In some implementations, the debugger can use hand icon 410 (e.g., a selectable icon) or other means, and change the viewpoint at the same specific snapshot in a timely manner (e.g., during debug event 420) to view the snapshot from different 3D coordinates (e.g., x2, y2, z2). In some implementations, the debugger can move device 310 to an exemplary location x2, y2, z2, and in response to this movement, the device updates the view of the snapshot.

[0061] Figure 5 The diagram illustrates a system flowchart of an exemplary environment 500 according to some specific implementation, in which the system can present a view including snapshots of application content playback within a 3D environment. In some specific implementations, the system flow of exemplary environment 500 is at the device (e.g., Figure 1 The system processes of exemplary environment 500 execute on a device (such as a mobile device, desktop computer, laptop computer, or server device) 120. Images of exemplary environment 500 may be displayed on a device such as an HMD having a screen for displaying images and / or a screen for viewing stereoscopic images. In some embodiments, the system processes of exemplary environment 500 execute on processing logic components (including hardware, firmware, software, or combinations thereof). In some embodiments, the system processes of exemplary environment 500 execute on a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).

[0062] The system flow of the exemplary environment 500 is from the physical environment (e.g., Figure 1 The physical environment 105) sensors acquire environmental data 502 (e.g., image data), and from the application (e.g., Figures 3 to 4Application 320 collects application data 506, integrates and records environment data 502 and application data 506, and generates interactive playback data for debuggers to view the execution of the application (e.g., to identify the occurrence of errors, if any). For example, the virtual application debugger technology described herein can replay the execution of the application from different viewpoints by recording program state information corresponding to the state of the application content at multiple points in time (e.g., the debugger application takes periodic snapshots), obtain changes in the application content, and generate a video recording of the application's execution based on the recorded program state information and changes in the application content from communication between the application and system processes.

[0063] In an exemplary specific implementation, environment 500 includes a device (e.g., Figure 1 An image synthesis pipeline that uses sensors on device 120 to acquire or obtain data about the physical environment (e.g., image data from an image source). Exemplary environment 500 is an example of acquiring image sensor data (e.g., light intensity data, depth data, and location information) from multiple image frames. For example, image 504 represents a user's physical environment (e.g., light intensity data, depth data, and location information) due to the user's position in the physical environment. Figure 1 The image data is collected by the user within a room in the physical environment (105). Image sources may include a depth camera that acquires depth data of the physical environment, a light intensity camera (e.g., an RGB camera) that acquires light intensity image data (e.g., a sequence of RGB image frames), and a position sensor for acquiring positioning information. For positioning information, some embodiments include a visual inertial ranging (VIO) system to estimate the distance traveled by determining equivalent ranging information using a sequence of camera images (e.g., light intensity data). Alternatively, some embodiments of this disclosure may include a SLAM system (e.g., a position sensor). The SLAM system may include a GPS-independent, multi-dimensional (e.g., 3D) laser scanning and range measurement system that provides real-time simultaneous localization and mapping. The SLAM system can generate and manage highly accurate point cloud data produced by reflections from laser scans of objects in the environment. Accurately tracking the movement of any point in the point cloud over time allows the SLAM system to use points in the point cloud as reference points for its position, maintaining an accurate understanding of its position and orientation as it travels through the environment. The SLAM system can also be a visual SLAM system that relies on light intensity image data to estimate the position and orientation of the camera and / or device.

[0064] In an exemplary implementation, environment 500 includes an application data pipeline that collects or obtains application data (e.g., application data from an application source). For example, the application data may include a virtual drawing application 508 (e.g., Figures 3 to 4(320) Virtual application. Application data may include 3D content (e.g., virtual objects) and user interaction data (e.g., haptic feedback of user interaction with the application).

[0065] In one exemplary embodiment, environment 500 includes an integrated environment recording instruction set 510 configured with instructions executable by a processor to generate playback data 515. For example, the integrated environment recording instruction set 510 obtains environment data 502 (e.g., such as...). Figure 1 The system obtains image data of the physical environment 105, acquires application data (e.g., a virtual application), integrates environmental data and application data (e.g., overlays a virtual application onto a 3D representation of the physical environment), records state changes and scene understanding during application execution, and generates playback data 515.

[0066] In one exemplary embodiment, the integrated environment recording instruction set 510 includes an integration instruction set 520 and a recording instruction set 530. The integration instruction set 520 is configured with instructions executable by a processor to integrate image data of the physical environment and application data from a virtual application to overlay the virtual application onto a 3D representation of the physical environment. For example, the integration instruction set 520 analyzes environmental data 502 to generate a 3D representation of the physical environment (video passthrough, optical perspective, or a reconstructed virtual room) and integrates the application data with the 3D representation, allowing the user to view the application as an overlay on the 3D representation during application execution, as referenced herein. Figures 3 to 4 As shown.

[0067] The recording instruction set 520 is configured with instructions executable by the processor to acquire integrated data 522 from the integrated instruction set 520 and generate recording data 532. For example, the recording instruction set 520 generates a recording of the application's execution based on recording program state information and changes in application content determined by user interactions during application execution. For example, the recording program state information may include recording scene understanding or snapshots, such as the position of objects in the environment. In some implementations, scene understanding may include head pose data, what the user is looking at in the application (e.g., virtual objects), hand pose information, etc. Additionally, the recording program state information may include a recording scene understanding mesh. The scene understanding mesh may include a 3D mesh generated concurrently during program execution.

[0068] In some implementations, environment 500 includes a debugger instruction set 540 configured with instructions executable by a processor to evaluate playback data 422 from integrated environment recording instruction set 510, presenting a set of views (e.g., from the same perspective as the application data, or from a different perspective) that replays application content within a 3D environment based on the recording data 532. In some implementations, in the device (e.g., Figure 1 The device 120 displays a second view set (e.g., from the device display 550) on the device monitor 550. Figures 3 to 4 (The different user perspectives / viewpoints shown). In some specific implementations, the debugger instruction set 540 generates interactive display data 542, such as cleaner tools 546 (e.g., Figures 3 to 4 The Cleaner Tool 330 allows debuggers to interact with playback data (e.g., rewind, change view, etc.).

[0069] In some implementations, the generated 3D environment 544 (e.g., scenes and other application content) can be rendered continuously / live-stream throughout execution and playback. That is, the rendering engine can run continuously for one time period to inject the executed application content and for another time period to inject the recorded application content. In some implementations, playback may differ from simply reconstructing the content in the same way as it was initially generated. For example, playback may involve using recorded values ​​of a ball's position instead of having the ball use a physics system. That is, a user (e.g., a programmer / debugger) can pause the test and use a cleanup tool to return to see the desired point in time. In some implementations, the user can replay from the same viewpoint, or alternatively, the user can change the viewpoint and see where the HMD is depicted, such as based on the gaze direction. In some implementations, the user can change the viewpoint to observe scene understanding (e.g., head position, hand position, 3D reconstructed mesh, etc.).

[0070] In some implementations, developers can return to enable the display of representations of sound sources (e.g., spatial stereo audio) and other invisible items. For example, when application data is paused via the cleaner tool 546, icon 547 can be presented to the debugger. For instance, icon 547 represents the 3D location of the sound source presented to the user when application data is "paused." In cases where the debugging event also includes sound located in an incorrect 3D position within the application data, the visual representation via icon 547 can help the debugger better pinpoint the location of the spatial stereo audio.

[0071] In some implementations, the integrated environment recording instruction set 510 and debugger instruction set 540 involve focusing on reusing changes / increments (e.g., scene understanding) of snapshots already used in the application-to-system reality architecture, enabling efficient and accurate recording with minimal performance impact. In one exemplary implementation, communication between the application and system processes (e.g., debugger / system applications) during application execution includes changes to application content, with a second view set generated based on the communication. For example, the system application (e.g., the debugger application) is responsible for the system processes. System processes may include rendering, presenting different viewpoints, etc. This exemplary implementation presents a reality engine architecture where changes between the application and system application are synchronized. For example, the system application may display content / changes provided by multiple applications. In some implementations, the system application may provide environmental content (e.g., walls, floors, system functionality, etc.) and application content within an environment that a user (e.g., a debugger or programmer) can interact with both system application components and application components.

[0072] Figure 6 This is a block diagram of an exemplary device 600. Device 600 shows... Figure 1 An exemplary device configuration of device 120. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure further relevant aspects of the specific implementations disclosed herein. Therefore, as a non-limiting example, in some specific implementations, device 600 includes one or more processing units 602 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 606, one or more communication interfaces 608 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BlueTooth, ZigBee, SPI, I2C and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 610, one or more displays 612, one or more internal and / or external image sensor systems 614, memory 620, and one or more communication buses 604 for interconnecting these components and various other components.

[0073] In some embodiments, one or more communication buses 604 include circuitry for interconnecting and communicating between system components. In some embodiments, the one or more I / O devices and sensors 606 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, or one or more depth sensors (e.g., structured light, time-of-flight, etc.), etc.

[0074] In some embodiments, one or more displays 612 are configured to present a view of a physical or graphical environment to a user. In some embodiments, one or more displays 612 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more displays 612 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. In one example, device 600 includes a single display. As another example, device 600 includes displays for each of the user's eyes.

[0075] In some embodiments, the one or more image sensor systems 614 are configured to acquire image data corresponding to at least a portion of the physical environment 105. For example, the one or more image sensor systems 614 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, depth cameras, event-based cameras, etc. In various embodiments, the one or more image sensor systems 614 may also include an illumination source emitting light, such as a flash. In various embodiments, the one or more image sensor systems 614 may also include an on-camera image signal processor (ISP) configured to perform multiple processing operations on the image data.

[0076] In some embodiments, device 120 includes an eye-tracking system for detecting eye position and eye movement (e.g., eye gaze detection). For example, the eye-tracking system may include one or more infrared (IR) light-emitting diodes (LEDs), an eye-tracking camera (e.g., a near-infrared (NIR) camera), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) towards the user's eyes. Furthermore, the illumination source of device 120 may emit NIR light to illuminate the user's eyes, and the NIR camera may capture images of the user's eyes. In some embodiments, the images captured by the eye-tracking system may be analyzed to detect the position and movement of the user's eyes, or to detect other information about the eyes such as pupil dilation or pupil diameter. Furthermore, the gaze point estimated from the eye-tracking images enables gaze-based interaction with content displayed on a near-eye display of device 120.

[0077] Memory 620 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 620 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 620 optionally includes one or more storage devices remotely located to one or more processing units 602. Memory 620 includes a non-transitory computer-readable storage medium.

[0078] In some embodiments, memory 620 or a non-transitory computer-readable storage medium of memory 620 stores an optional operating system 330 and one or more instruction sets 640. Operating system 630 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, instruction set 640 includes executable software defined by binary information stored in charge. In some embodiments, instruction set 640 is software executable by one or more processing units 602 to implement one or more of the techniques described herein.

[0079] Instruction set 640 includes integrated environment instruction set 642 and debugger instruction set 644. Instruction set 640 can be represented as a single software executable file or multiple software executable files.

[0080] Integrated Environment Instruction Set 642 (e.g., Figure 3 The point cloud registration instruction set 320 can be executed by the processing unit 602 to generate playback data 515. For example, the integrated environment recording instruction set 642 obtains environmental data (e.g., such as...). Figure 1The system records image data of the physical environment 105, obtains application data (e.g., a virtual application), integrates the environment data and application data (e.g., overlaying the virtual application onto a 3D representation of the physical environment), records state changes and scene understanding during application execution, and generates playback data. In one exemplary embodiment, the integrated environment recording instruction set includes an integration instruction set and a recording instruction set. The integration instruction set is configured with instructions executable by a processor to integrate image data of the physical environment and application data from the virtual application to overlay the virtual application onto a 3D representation of the physical environment. For example, the integration instruction set analyzes the environment data to generate a 3D representation of the physical environment (video passthrough, optical perspective, or a reconstructed virtual room) and integrates the application data with the 3D representation, allowing the user to view the application as an overlay on the 3D representation during application execution, as referenced herein. Figures 3 to 4 As shown.

[0081] The debugger instruction set 644 is configured with instructions executable by the processor to evaluate playback data from the integrated environment instruction set 642, and presents a view set including playback of application content within a 3D environment based on the playback data. In some specific implementations, in the device (e.g., Figure 1 The device 120 displays a second view set (e.g., from the display of the device 120). Figures 3 to 4 (The different user perspectives / viewpoints shown). In some specific implementations, the debugger instruction set 540 generates interactive display data, such as cleaner tools (e.g., Figures 3 to 4 The Cleaner Tool 330 allows debuggers to interact with playback data (e.g., rewind, change view, etc.).

[0082] Although instruction set 640 is shown as residing on a single device, it should be understood that in other specific implementations, any combination of elements may reside in separate computing devices. Furthermore, Figure 6 More often, this serves as a functional description of various features present in a particular implementation, which differ from the structural schematics of the specific implementation described herein. As will be appreciated by those skilled in the art, items shown individually may be combined, and some items may be separate. The actual number of instruction sets and how features are allocated therein will vary depending on the specific implementation and may depend in part on the specific combination of hardware, software, and / or firmware chosen for that particular implementation.

[0083] Some specific embodiments disclosed herein provide techniques for integrating information (e.g., partial point clouds) from any number of images of a scene captured from arbitrary views (e.g., viewpoints 1102a and 2102b). In some specific embodiments, the techniques disclosed herein are used to estimate transformation parameters of partial point clouds using two or more images (e.g., images captured from two different viewpoints by a mobile device, HMD, laptop computer, or other device (e.g., device 120)). These techniques can leverage machine learning models that include deep learning models, which take two partial point clouds as input and directly predict the point-by-point position of one point cloud in the coordinate system of the other without explicit matching. Deep learning can be applied to each partial point cloud to generate a latent representation of the estimated transformation parameters. In some specific embodiments, the deep learning model generates an associated confidence level for each latent value. Using the predictions and confidence values ​​from different images, these techniques can combine the results to generate a single estimate of each transformation parameter of the partial point cloud.

[0084] Several details have been described to provide a thorough understanding of the exemplary aspects illustrated in the accompanying drawings. Furthermore, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will recognize that well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein. Moreover, other effective aspects and / or variations do not include all the specific details described herein.

[0085] The aspects of the methods disclosed herein can be performed in the operation of a computing device. The order of the boxes presented in the above examples can be varied; for example, the boxes can be reordered, combined, and / or divided into sub-blocks. Some boxes or processes can be executed in parallel. The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0086] The embodiments of the subject matter and operations described in this specification may be implemented in digital electronic circuits or in computer software, firmware, or hardware (including the structures disclosed in this specification and their equivalents) or in a combination thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by or control of the operation of a data processing device. Alternatively or otherwise, the program instructions may be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof, or may be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device. Furthermore, although the computer storage medium is not a propagating signal, it may be a source or destination of computer program instructions encoded in artificially generated propagating signals. Computer storage media can also be one or more separate physical components or media (e.g., multiple CDs, disks or other storage devices), or included in one or more separate physical components or media.

[0087] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines for processing data, including programmable processors, computers, systems-on-a-chip, or multiple or combinations thereof. The apparatus may include special-purpose logic circuitry (e.g., FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits)). In addition to hardware, the apparatus may include code that creates an execution environment for the computer program under consideration, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The apparatus and execution environment can implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures. Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as "processing," "computing," "calculating," "determining," and "identifying" refer to the actions or processes of computing devices, such as one or more computers or similar electronic computing devices, that manipulate or convert data represented as physical electronic or magnetic quantities within the memory, registers, or other information storage devices, transmission devices, or display devices of a computing platform.

[0088] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on multi-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more specific embodiments of the subject matter of this invention. The teachings contained herein can be implemented in the software used for programming or configuring the computing device using any suitable programming, scripting, or other type of language or combination of languages.

[0089] Specific implementations of the methods disclosed herein can be performed in the operation of such a computing device. The order of the boxes presented in the above examples can be varied; for example, the boxes can be reordered, combined, and / or divided into sub-blocks. Some boxes or processes can be executed in parallel.

[0090] The use of “applies to” or “configured to” in this document implies open and inclusive language, which does not exclude applicability to or configuration to devices performing additional tasks or steps. Similarly, the use of “based on” implies openness and inclusivity, as processes, steps, calculations, or other actions “based on” one or more of the stated conditions or values ​​may in practice be based on additional conditions or values ​​beyond those stated. The headings, lists, and numbering included in this document are for illustrative purposes only and are not intended to be restrictive.

[0091] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first node can be called a second node, and similarly, a second node can be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.

[0092] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and the appended claims, the singular forms “a” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the term “comprising” as used in this specification specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0093] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" are interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when the prerequisite is detected" or "in response to detection" that the prerequisite is true, depending on the context.

[0094] While this specification contains numerous specific implementation details, these details should not be construed as limiting the scope of any invention or potentially claimed content, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in the context of different embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, while certain features may be described above as functioning in certain combinations and even initially claimed in this manner, one or more features of a claimed combination may be removed from that combination in certain circumstances, and the claimed combination may involve sub-combinations or variations thereof.

[0095] Similarly, although operations are shown in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in a sequential order or the specific order shown, or requiring all shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the division of various system components in the above embodiments should not be construed as requiring such division in all embodiments, and it should be understood that the program components and the system may generally be integrated together in a single software product or packaged into multiple software products.

[0096] Therefore, specific embodiments of the subject matter have been described. Other embodiments are also within the scope of the following claims. In some cases, the actions described in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes shown in the accompanying drawings do not necessarily require a specific order or sequence to achieve the desired result. In some specific embodiments, multitasking and parallel processing may be advantageous.

Claims

1. A non-transitory computer-readable storage medium storing program instructions executable by one or more processors to perform operations, the operations including: During the execution of the application, A first set of views, comprising application content provided by the application, is presented within a three-dimensional (3D) environment, wherein the first set of views is provided from a first set of viewpoints within the 3D environment; Record program state information corresponding to the execution of the application content at multiple points in time, wherein recording the program state information includes recording scene understanding associated with the 3D environment, the scene understanding including the position and attributes of objects in the 3D environment and posture data associated with the user's head or hand; Changes to the application content are recorded, the changes being determined based on user interactions including user actions and responses to those actions, the responses being specified by the configuration of the executing application; Receive debugger / system process communication between the application and a system process that includes the changes to the application's content, wherein the debugger / system process communication includes information associated with execution time changes related to the application being executed in the 3D environment; as well as Based on the recording of the program state information, the recording of the changes in the application content, and the debugger / system process communication between the application and the system process, a recording of the execution of the application is generated; as well as After the application is executed, a second view set is presented in the 3D environment based on the recording of the program state information, the recording of the changes in the application content, and the debugger / system process communication between the application and the system process. The second view set is generated based on the debugger / system process communication between the application and the system process at the plurality of time points, and the second view set is provided from a second view set different from the first view set.

2. The non-transitory computer-readable storage medium according to claim 1, wherein the recording of the execution of the application based on the program state information and the changes in the application content comprises: Obtain recorded program status information corresponding to the state of the application content at multiple points in time; Obtain the changes in the application content, wherein the changes include changes in the application content that occur between the states; as well as The recording of the application's execution is generated based on the recorded program state information and the changes in the application content from the communication between the application and the system process at the multiple time points.

3. The non-transitory computer-readable storage medium of claim 1, wherein the application content includes objects, and the change includes an incremental value of a change in the position of the objects.

4. The non-transitory computer-readable storage medium of claim 1, wherein the operation further includes receiving input, the input selecting a point in time during the execution of the application as the start point of the playback.

5. The non-transitory computer-readable storage medium of claim 1, wherein presenting the playback includes presenting a graphical depiction of the head position, gaze direction, or hand position of the user executing the application.

6. The non-transitory computer-readable storage medium of claim 1, wherein presenting the playback includes presenting a graphical depiction of the sound source.

7. The non-transitory computer-readable storage medium according to claim 1, wherein: Presenting the first view set or the second view set includes presenting video passthrough or perspective images of at least a portion of a physical environment, wherein a 3D reconstruction of at least said portion of the physical environment is dynamically generated during the execution of the application; and Presenting the playback includes presenting the 3D reconstruction.

8. The non-transitory computer-readable storage medium according to claim 1, wherein, During the execution of the application, objects within the application content are located based on a physics engine; and During playback, the objects in the application content are positioned based on determining the position of the objects according to the program state information and repositioning the objects according to the changes.

9. The non-transitory computer-readable storage medium according to claim 1, wherein, During the execution of the application, the first view set or the second view set is presented on the head-mounted device (HMD).

10. An apparatus, the apparatus comprising: Non-transitory computer-readable storage medium; as well as One or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the one or more processors to perform operations, the operations including: During the execution of the application, A first set of views, comprising application content provided by the application, is presented within a three-dimensional (3D) environment, wherein the first set of views is provided from a first set of viewpoints within the 3D environment; Record program state information corresponding to the execution of the application content at multiple points in time. The program state information includes the position and attributes of objects in the 3D environment and information related to user interaction in the 3D environment. Changes to the application content are recorded, the changes being determined based on user interactions including user actions and responses to those actions, the responses being specified by the configuration of the executing application; Debugger / system process communication between the application and a system process that receives the changes to the application's content, wherein the debugger / system process communication includes information associated with execution time changes related to the application running in the 3D environment; and Based on the recording of the program state information, the recording of the changes in the application content, and the debugger / system process communication between the application and the system process, a recording of the application's execution is generated; and After the application is executed, a second view set is presented in the 3D environment based on the recording of the program state information, the recording of the changes in the application content, and the debugger / system process communication between the application and the system process. The second view set is generated based on the debugger / system process communication between the application and the system process at the plurality of time points, and the second view set is provided from a second view set different from the first view set.

11. The device of claim 10, wherein the recording of the execution of the application based on the recording of the changes in the program state information and the application content comprises: Obtain recorded program status information corresponding to the state of the application content at multiple points in time; Obtain the changes in the application content, wherein the changes include changes in the application content that occur between the states; as well as The recording of the application's execution is generated based on the recorded program state information and the changes in the application content from the communication between the application and the system process at the multiple time points.

12. The device of claim 10, wherein the operation further comprises receiving input, the input selecting a point in time during the execution of the application as the start point of the playback.

13. The device of claim 10, wherein presenting the playback includes presenting a graphical depiction of the head position, gaze direction, or hand position of the user performing the execution of the application.

14. The device of claim 10, wherein the application content includes objects, and the change includes an incremental value of a change in the position of the objects.

15. The device of claim 10, wherein presenting the playback includes presenting a graphical depiction of the sound source.

16. The device according to claim 10, wherein: The view presenting the scene includes a video passthrough or perspective image presenting at least a portion of the physical environment, wherein a 3D reconstruction of at least said portion of the physical environment is dynamically generated during the execution of the application; and Presenting the playback includes presenting the 3D reconstruction.

17. The device of claim 10, wherein during the execution of the application, objects of the application content are located based on a physics engine; and During playback, the objects in the application content are positioned based on determining the position of the objects according to the program state information and repositioning the objects according to the changes.

18. A method, the method comprising: In electronic devices with processors: During the execution of the application, A first set of views, comprising application content provided by the application, is presented within a three-dimensional (3D) environment, wherein the first set of views is provided from a first set of viewpoints within the 3D environment; Record program state information corresponding to the execution of the application content at multiple points in time. The program state information includes the position and attributes of objects in the 3D environment and information related to user interaction in the 3D environment. Changes to the application content are recorded, the changes being determined based on user interactions including user actions and responses to those actions, the responses being specified by the configuration of the executing application; Receive debugger / system process communication between the application and a system process that includes the changes to the application's content, wherein the debugger / system process communication includes information associated with execution time changes related to the application being executed in the 3D environment; as well as Based on the recording of the program state information, the recording of the changes in the application content, and the debugger / system process communication between the application and the system process, a recording of the execution of the application is generated; as well as After the application is executed, a second view set is presented in the 3D environment based on the recording of the program state information, the recording of the changes in the application content, and the debugger / system process communication between the application and the system process. The second view set is generated based on the debugger / system process communication between the application and the system process at the plurality of time points, and the second view set is provided from a second view set different from the first view set.

Citation Information

Patent Citations

  • Media Compositor For Computer-Generated Reality

    CN110809149A

  • Visual history for content state changes

    US20180276004A1