Methods of streaming and / or recording from an extended reality headset
By rendering and streaming video/audio from virtual cameras/microphones to a remote server, the headset streamlines content contribution, overcoming setup complexities and lag, making it accessible to a broader user base.
Patent Information
- Application Number
- PCT/GB2025/050665
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2025-03-27
- Publication Date
- 2025-10-02
AI Technical Summary
Existing extended reality headsets cannot stream video directly to social platforms without requiring a separate computer, which increases processing demands and limits user participation due to setup complexities and potential lag, making it inaccessible to most users.
The headset renders video and audio from virtual cameras and microphones, sending it to a remote server for compilation, recording, and streaming, reducing processing burden on the headset and eliminating the need for a separate PC.
Enables direct streaming and recording from standalone headsets without lag, lowering entry barriers and allowing users to contribute content efficiently, even on low-powered devices.
Smart Images

Figure 00000032_0000 
Figure 00000033_0000 
Figure 00000034_0000
Abstract
Description
[0001] Methods of streaming and / or recording from an extended reality headset The present invention relates to extended reality headsets and in particular improved methods of streaming and / or recording from extended reality headsets which do not require tethering or casting to a separate computer. In embodiments of the invention a standalone extended reality headset can stream directly to the internet, and optionally store the video remotely, without significantly interrupting or otherwise affecting the extended reality experience of a user of the headset. The invention lowers the bar to entry for users who might wish to contribute content to meet the ever-increasing demand for same. Background to the invention
[0002] The term extended reality (also known by the shorthand XR) is used herein as an umbrella term covering virtual reality (VR), augmented reality (AR) and mixed reality (MR), spatial computing or spatial reality devices and the like, and refers to any such immersive experience in which a user is presented and / or interacts with a virtual or augmented environment, and in particular where a headset (or other personal visual display device) is used to present said environment to the user. Any references to virtual reality (VR) can therefore be extended to any other form of extended reality (XR) experience, and vice versa, unless they are clearly incompatible. Examples of such headsets include the Meta Quest line of all-in-one VR headsets, and the recently-launched Apple Vision Pro.
[0003] At the filing date of the present application, there is no way to stream video of gameplay or the like directly from standalone virtual reality headsets (or other extended reality devices - see above) to social live streaming platforms such as Twitch or YouTube. This limitation also applies to private streaming platforms such as may be used in business applications. One way of indirectly streaming video to such platforms is to cast video from a standalone headset to a local PC (i.e. a PC on the same local network), and then to stream the cast video as it is reproduced on the PC display to the streaming platform. This has a number of disadvantages including that the PC must be sufficiently powerful to both display the cast video and then to record, encode and stream the displayed cast video to the internet (and / or to store it). Another obvious disadvantage is the requirement to actually have a PC which negates the main advantage of having a standalone headset. Note PC is used generally herein to refer to a separate computer, such as a Windows, Linux or Mac desktop or laptop, which will typically have a high performance graphics card (GPU) to render the incoming headset data.
[0004] A mobile phone or tablet may take the place of the PC, and most users will have access to at least one of these devices if they have a standalone headset. However, all that does is shift the burden of displaying the cast video and recording, encoding and streaming the displayed cast video to the internet onto the mobile phone or tablet. As capable as such devices have become in recent times, recording, encoding and streaming high resolution video (and high bit-rate audio) is extremely intensive and at the levels of quality demanded by viewers of such streams it is beyond all but the most advanced devices. Again, this negates the main advantage of having a standalone headset, and furthermore requires additional setup outside of the headset environment. Often headsets require some “getting comfortable” prior to use and requiring a user to put on a headset for initial setup and then take off to configure the PC-or-other-device based stream, and put the headset back on again can be an additional source of frustration and barrier to participation.
[0005] In the case of PC-powered extended reality experiences where the experience is generated (i.e. visuals rendered and audio produced) on the PC and sent to the headset, streaming to the internet must also be done from the PC itself, in a similar manner. As mentioned above the rendering of high-resolution video and production of high bit-rate and high-definition audio are memory, processor and GPU intensive, and requiring the PC to also record, encode and stream the video and audio to the internet significantly increases the memory and processor requirements and puts this out of the reach of most users.
[0006] As standalone headsets are now available at relatively modest price points the extended reality experience is now within the reach of more users than ever before. There is also an ever-increasing demand for gameplay streaming content; to meet this demand it will be very beneficial to enable direct streaming from such headsets so that users without powerful PCs can contribute content.
[0007] It is also critical that there be no slowdown or lag in delivery of the extended reality experience to the headset user. All in one headsets tend to run on relatively low powered computing devices, typically based on an ARM64 processor architecture, and developers tend to utilise the power of such devices to near maximum to deliver the best possible experience without lag or slowdown. If there is lag or slowdown it can induce sickness in a user akin to travel sickness, limiting enjoyment and ultimately curtailing use. It is therefore an object of aspects of the invention to enable direct streaming from standalone extended reality headsets and to obviate and / or mitigate one or more disadvantages of prior art arrangements for streaming such experiences. Further aims and objects of the invention, and likewise benefits, will become apparent from reading the following description.
[0008] Summary of the invention
[0009] According to a first aspect of the invention there is provided a method of streaming from an extended reality headset comprising the steps of: in the headset, rendering video in accordance with a first person extended reality experience of a user; in the headset, simultaneously rendering video from the perspective of at least one virtual camera and sending it to a remote server; at the remote server, producing a composite video comprising the video from the virtual camera; and streaming the composite video from the remote server to one or more viewers.
[0010] Most preferably, rendering video in accordance with a first person extended reality experience of a user comprises rendering video from the perspective of a virtual camera corresponding to the left eye of a user and a virtual camera corresponding to the right eye of the user.
[0011] Most preferably, the method also comprises generating audio in accordance with the first person extended reality experience of the user, simultaneously generating audio from the perspective of at least one virtual audio listener and / or physical microphone and sending it to the remote server, wherein the composite video further comprises the audio from the virtual audio listener or physical microphone.
[0012] In this process, video is rendered and audio generated for presentation to a user in the usual way (in accordance with an extended reality experience of a user) or at least in a manner which is relatively uninterrupted or unaffected in comparison to the state of the art despite the additional steps. Simultaneously, video from the perspective of one or more virtual cameras (and optionally audio from one or more virtual audio listeners and / or physical microphones) is sent to a remote server which handles the compilation, recording, encoding and streaming of the composite video. This means that the processing burden is for the most part not placed on the headset (or other personal visual display device) and does away with the need for a separate / tethered PC.
[0013] Optionally, the virtual camera and the optional virtual audio listener and / or physical microphone correspond to a user’s first person perspective. In this case, the video from the perspective of the virtual camera and the optional audio from the perspective of the virtual audio listener and / or physical microphone may be the same video and audio of the user’s extended reality experience, or at least correspond to it. Preferably, the first person perspective corresponds to one of a left or right eye channel of the first person extended reality experience of the user.
[0014] Alternatively, or additionally (in the case of more than one), the virtual camera and the optional virtual audio listener and / or physical microphone may correspond to a third party perspective.
[0015] Further alternatively, or further additionally (in the case of more than one), the virtual camera and the optional virtual audio listener and / or physical microphone may be selected to follow an object such as a virtual tool within the environment. Optionally, the position and / or focus of the virtual camera and the optional virtual audio listener and / or physical microphone may be controlled by the user and / or by a viewer.
[0016] Optionally, the method comprises establishing a plurality of virtual cameras and optional virtual audio listeners and / or physical microphones, and selecting which of the virtual cameras and optional virtual audio listeners and / or physical microphones to render video and optionally generate audio whereby the composite video is from the perspective of the selected virtual camera and optional virtual audio listener and / or physical microphone.
[0017] Optionally, the virtual camera and optional virtual audio listener and / or physical microphone are selected by the user and / or a viewer. Alternatively, the virtual camera and optional virtual audio listener and / or physical microphone are selected responsive to an event occurring within the extended reality experience of the user. Optionally, the method comprises authenticating the user and / or the headset. Optionally, the user and / or the headset are authenticated at the remote server. Alternatively, the user and / or the headset are authenticated at a separate server.
[0018] Optionally, the method comprises selecting the remote server from a plurality of remote servers. Preferably, the remote server is selected based on a predicted or measured connection performance.
[0019] Most preferably, the method comprises establishing a peer-to-peer connection between the headset and the remote server. Optionally, the method comprises determining an optimal route to the remote server.
[0020] Optionally, the rendered video and generated audio are sent from the headset to the remote server via a videoconference session. The rendered video may be sent via a WebRTC video channel and the generated audio may be sent via a WebRTC audio channel.
[0021] Most preferably, the composite video rendering is carried out at an adaptive bitrate, wherein the bitrate may be selected and / or adjusted responsive to network parameters.
[0022] Most preferably, the method comprises capturing one or more events within the extended reality experience of the user, one or more real world parameters, and / or one or more network parameters. These might include for example the position and rotation of the user’s head and the user’s hands, controller positions and rotations, and button presses. Optionally, the composite video may comprise the events and / or parameters, or representations thereof. Alternatively, or additionally, the method comprises streaming the events and / or parameters. Optionally, the method further comprises storing the events and / or parameters separately from the composite video.
[0023] Optionally, streaming the video comprises presenting the video to one or more viewers via a web browser. Alternatively, streaming the video comprises starting a session at a streaming platform and sending the video to the streaming platform where the video may be rendered to one or more viewers, preferably in real-time.
[0024] Optionally, the method further comprises, at the remote server, storing the composite video for later retrieval.
[0025] It may be desirable not to stream the composite video (and optionally the events and parameters) but just to store it (or them) for later retrieval, for example to view on demand, or if the video requires editing prior to distribution.
[0026] Accordingly, in a second aspect of the invention there is provided a method of recording video from an extended reality headset comprising the steps of: in the headset, rendering video and generating audio in accordance with a first person extended reality experience of a user; in the headset, simultaneously rendering video from the perspective of at least one virtual camera and optionally generating audio from the perspective of at least one optional virtual audio listener and / or physical microphone; sending the rendered video and optionally generated audio from the headset to a remote server; at the remote server, producing a composite video comprising the video and optionally the audio from the virtual camera and the optional virtual audio listener and / or physical microphone and storing the composite video for later retrieval.
[0027] Embodiments of the second aspect of the invention may comprise features of or corresponding to the preferred or optional features of the first aspect of the invention, which for brevity are not repeated above.
[0028] It is also foreseen in the description that follows that the video to be streamed and / or recorded is taken from the video already being rendered for the user’s extended reality experience. Accordingly a variant of the foregoing is provided by a third aspect of the invention, which is a method of streaming or recording video from an extended reality headset comprising the steps of: in the headset, rendering video from the perspective of a virtual camera corresponding to the left eye of a user and rendering video from the perspective of a virtual camera corresponding to the right eye of the user, displaying both to the user, and simultaneously sending one or both to a remote server; at the remote server, producing a composite video comprising the video received from the virtual camera; and storing the composite video and / or streaming the composite video from the remote server to one or more viewers.
[0029] Embodiments of the third aspect of the invention may comprise features of or corresponding to the preferred or optional features of the first or second aspects of the invention, which for brevity are not all repeated above.
[0030] According to a fourth aspect of the invention there is provided a computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method of the first, second or third aspects.
[0031] The computer program may comprise a software development kit comprising a package to be imported into a game engine. The package may comprise a plurality of modules or assets including virtual cameras, virtual audio listeners and / or physical microphone interfaces, event trackers, controls, virtual objects and which may provide interactive features such as chat and commenting. There may be provided a management module to track key events, start and stop recording on the remote server, and control parameters such as resolution and bitrate of the rendered and / or composite video. There may also be provided user interface elements
[0032] According to a fifth aspect of the invention there is provided a computer readable medium or data carrier comprising the computer program of the fourth aspect. According to a sixth aspect of the invention there is provided a data carrier signal carrying the computer program of the fourth aspect.
[0033] Embodiments of the fourth to sixth aspects of the invention may comprise features of or corresponding to the preferred or optional features of any other aspect of the invention or vice versa.
[0034] According to a seventh aspect of the invention, there is provided an extended reality headset configured, modified, adapted or arranged to perform the method (or at least the steps performed by the headset) of the first, second or third aspects of the invention.
[0035] According to an eighth aspect of the invention, there is provided a server configured modified, adapted or arranged to perform the method (or at least the steps performed by the remote server) of the first, second or third aspects of the invention.
[0036] According to a ninth aspect of the invention, there is provided a system comprising the extended reality headset of the seventh aspect and the server of the eight aspect.
[0037] Embodiments of the seventh to ninth aspects of the invention may comprise features of or corresponding to the preferred or optional features of any other aspect of the invention or vice versa.
[0038] Brief description of the drawings
[0039] Aspects and advantages of the present invention will become apparent upon reading the following detailed description and upon reference to the following drawings (like reference numerals referring to like features) in which:
[0040] Figure 1 is a flow diagram (split across pages 1 / 8-3 / 8) illustrating a method of streaming video from an extended reality headset in accordance with an embodiment of the invention;
[0041] Figure 2 is a flow diagram illustrating the development steps in implementing the invention in a cross-platform game engine in accordance with an embodiment of the invention;
[0042] Figure 3A is a flow diagram illustrating the conventional approach to streaming extended reality games and Figure 3B is a flow diagram illustrating the inventive approach to streaming extended reality games;
[0043] Figure 4A is a flow diagram illustrating the conventional approach to streaming extended reality line-of-business (LOB) applications and Figure 4B is a flow diagram illustrating the inventive approach to streaming extended reality line-of-business (LOB) applications;
[0044] Figure 5A is a flow diagram illustrating the conventional approach to recording extended reality games and Figure 5B is a flow diagram illustrating the inventive approach to recording extended reality games; and
[0045] Figure 6A is a flow diagram illustrating the conventional approach to recording extended reality line-of-business (LOB) applications and Figure 6B is a flow diagram illustrating the inventive approach to recording extended reality line-of-business (LOB) applications. Detailed description of the drawings
[0046] Figure 1 is a flow diagram illustrating a method of streaming video (and / or recording video as the case may be) from an extended reality headset in accordance with an embodiment of the invention (embodied in a system labelled “Beam”). To demonstrate how the method can be distributed (in a non-limiting way) between the headset, remote or ancillary services and servers, and viewers (e.g. via browser or social media streaming platforms) making up a corresponding streaming system or platform, various steps are shown in corresponding vertical locations. In the examples herein the system is implemented on a headset running a game created and operated on the Unity Engine which will be familiar to the person skilled in the art at the filing date. It will be trivial to replicate the invention described herein in that context in other development engines and the like, for example the Unreal Engine or any other current or future integrated development environment for 3D. Note that while the term headset is used extensively herein this term will be understood to encompass any personal visual display device.
[0047] Steps are shown as taking place on or in specific parts of the system or platform (to which the different vertical locations relate) but it will be understood that with a few exceptions, steps which occur or are implemented in one part (such as within “Beam Services”) might instead occur or be implemented in another part (such as on the “Beam Video server”). For example, user authentication may take place on the same server that renders the composite video.
[0048] At 101 the process begins with a user actively selecting an option to stream video from the user’s headset. This corresponds to the step of pressing a button in the VR / MR environment shown in later figures. It will however be appreciated that streaming may occur automatically upon opening an application configured to stream video from a headset in accordance with an aspect of the invention, in which case this step 101 may instead correspond to the step of launching an application configured in accordance with an aspect of the invention. At 103 package authentication occurs, by which it is determined whether the headset, the application and / or the user is allowed to access the remote services. If authentication is successful a unique token is generated at 105 and returned to the headset. Alternatives to token based authentication may be implemented but this is the preferred approach for user authentication at least. Authentication of the headset and / or the application might make use of API keys instead for example. Thereafter the process of recording may begin at 107.
[0049] An optimal server is chosen at 109 by the remote service. The server is chosen from amongst a plurality of servers comprised in the system or platform and may, but not necessarily, be based on proximity to the user and / or viewer(s) of the video stream. Server selection is preferably performed on the basis of predicted or measured best connection performance, and will take account of distance, load and consistency for example.
[0050] This server can interchangeably be referred to as a remote server (acknowledging its physical separation from the headset) or a video server (acknowledging its purpose to generate and / or store the video stream).
[0051] Having selected the server, a videoconference session is initiated at 111 , and a corresponding session monitor started at 113. At the headset an optimal route to the video server (selected at 109) from the headset is found by interactive connectivity establishment (ICE) at 115, following which a peer-to-peer connection is established by initiating a call 117 and attaching a camera 119 and an audio listener 121 as explained in more detail below.
[0052] This is of course just one way in which a remote server might be selected and / or a session might be initiated or a contact established between a headset and a remote server; what is important in this respect is that there is a remote server which has the primary purpose of compiling, recording, encoding and / or streaming generated content (and / or storing it), and that the headset is capable of sending video and audio to the server so that it can do these things rather than requiring the headset to. On the headset itself, a normal virtual reality or mixed reality experience is rendered 123 in the usual way but in contrast to current approaches a Render Texture (being a visual reproduction of a virtual in game camera view in the Unity engine) is simultaneously generated and sent via the peer to peer connection 125 and the output of an audio listener (being an audio reproduction of a virtual in game microphone) is also simultaneously generated and sent via the peer to peer connection 127 to the selected server. Optionally, audio can also be captured from a physical microphone, for example to transmit and / or record what the user says in game.
[0053] In a specific embodiment of the invention, the Render Texture is sent to the remote server via a WebRTC video channel and the audio listener is sent to the remote server via a WebRTC audio channel. WebRTC has a number of advantages over the HTTP Live Streaming (HLS) protocol typically used to stream a VR / MR experience, including that WebRTC runs with significantly lower latency than HLS because WebRTC is peer-to-peer whereas HLS uses a client-server model requiring an intermediary server. Though the arrangement described herein requires a remote server, HLS would require another server still as an intermediary.
[0054] In this example, only one virtual camera and only one virtual audio listener is described, but in other embodiments of the invention there may be provided multiple of either or both, and also optionally one or more physical microphones, and a user, developer or viewer may select from amongst one or more of these. It is also foreseen that the / a virtual camera may be from the point of view of the user (e.g. left or right eye view), and the inventive method might indeed simply take such a left or right eye view as the video channel to be sent to the remote server. It is also foreseen that the / a virtual camera might be controlled (e.g. position, focus, etc. by a user, developer, or viewer) and / or that the system may switch automatically between different virtual cameras responsive to events occurring in the extended reality experience of the user (e.g. switching to show an otherwise unseen non-player- character affected by in-game actions of a user, or to show alternative views of a sporting occasion such as a “crossbar cam” when a user-controlled character takes a shot at goal). For the avoidance of doubt, it will be understood that multiple video channels or a single video channel comprising content from multiple virtual cameras can be sent to the remote server. In this way, multiple camera angles can be combined, in a variety of ways, in a single composite video stream.
[0055] It is also envisaged to allow real-time viewer interaction whereby, in addition to camera-switching and object tracking, viewers may be able to modify gameplay (for example) directly on the headset. This is possible because the viewer can provide input to a streaming interface that is directly communicated to the headset, thereby allowing modification to game mechanics without requiring server-side game state processing. In addition to switching camera angles and object tracking, viewers might be able to control or affect NPC behaviour or the behaviour of other objects in the game environment.
[0056] In addition to allowing viewer interaction, two-way communication can also be provided for (e.g. via WebRTC) to allow for tightly integrated features from video forwarding platforms like Twitch or Youtube. Live chat and polling from onward streaming parties means that a streamer (wearing the headset) can change their ingame behaviour live, according to viewer feedback, without taking several steps to view the feedback in the headset. A further, tighter integration between, say, an SDK and the game can allow polls and other feedback, received via incoming communication channels to influence gameplay live. Furthermore, video can be enhances with bespoke and / or interactive overlays, allowing for in-stream marketing as well as other communication to viewers (e.g. statistical or event data).
[0057] In some applications the virtual audio listener(s) and / or physical microphone(s) might be dispensed with altogether, e.g. if audio does not form part of the extended reality experience or is not of importance to the viewer.
[0058] At this stage the best video codec is selected at 129, which for example may be Vp8, H.264, H.265 or VP9, based on parameters such as network conditions (e.g. bandwidth and latency). A composite video is then rendered at 131 from the render texture and the audio listener output. Importantly, this rendering is carried out at an adaptive bitrate, selected and adjusted responsive to network parameters. A viewer is then able to watch the composite video 133 via an internet browser or the like. Note the term composite video here suggests video and audio but as intimated above audio might be dispensed with. It is foreseen however that the composite video might also comprise data and / or captured events (see below).
[0059] Again, what is important in this respect is not what is sent to the remote server in terms of video content (in the example above it is Render Texture but it could be any other form of virtual camera output showing a view of the extended reality world being experienced by the user) but simply that the video content is sent to the remote server. As mentioned above, it is key that it is the remote server that does the compiling, recording, encoding and streaming, not the headset.
[0060] It is also very beneficial that the process is made easy for the user; e.g. in the simplest case providing a simple button or the like in the extended reality environment (or indeed a simple hardware button on the headset) with no need for pre-setting up or forward planning by the user.
[0061] Another key aspect of the solution which finds particular application in streaming of gaming is the capture of events within the extended reality experience of the user 135. This data can take any form which the user and / or the developer desire (e.g. game related data such as framerate, resolution, in-game ammo count etc.), and may also include real world parameters (e.g. user’s heart rate measured by a heart rate monitor). This data can be provided in an endless stream (aka “firehose”) to the service 137 as well as being fed forward for the purposes of inclusion in or alongside live video broadcast. As mentioned elsewhere herein, some or all of the data, or representations thereof, might be included in the composite video.
[0062] If the user chooses to stream video to a social media platform 139, typically by selecting an in-game menu option, a session is started at (say) Twitch or YouTube 141 , following which a stream can be initiated by a viewer on said platform 143. At the service the existing video server session is attached 145 and the composite video from 131 is rendered to the social media session 147 thus generating a social livestream of the composite video 149.
[0063] In other examples, such as those described further below, the composite video can be stored and / or streamed on a private platform or service.
[0064] In any case it is again important to note that the streaming is done by the remote server, not by the headset, which involvement is limited to generating the video and audio to be streamed by the remote server thus limiting the additional processing burden (and in some limited cases, for example where the left or right eye view is utilised, no additional burden is placed on the headset).
[0065] The user may terminate the session 151 at any time, following which the server closes the session 153 and the remote server terminates the social streams 155 thus bringing the social media sessions to an end 157. Recording is stopped 159 and the recording may be stored 161 for playback at a later date. At the same time, and preferably with the video and audio data (and optionally within a composite video), the data can also be stored. One way of doing this is by creating a digital twin of the data source in the AR / MR experience. The recorded video can optionally be uploaded to Youtube or TikTok directly from the headset browser, from the headset via a portal in a web browser, or automatically via API. This brings the streaming process to an end 163.
[0066] At a later date, the user and / or viewers can watch back the recorded video of the streaming session. If data has been recorded with the video, the data can also be played back with the video of the streaming session, and / or observed via the abovedescribed digital twin. In one example, where the user’s heart rate is measured, this could be represented by a virtual ECG in the video stream.
[0067] It will be understood that in the foregoing many steps may be performed out of the order described and also by different parts of the system or platform. For example, the selection of an optimal server might be performed on the headset itself, or the optimal route to a videoconference server might be performed by the remote service. It is expected that the invention will be most easily implemented via a software development kit (SDK). In one embodiment of the invention there can be provided an SDK for the Unity cross-platform game engine that will allow developers to incorporate the above-described method of streaming into their own projects.
[0068] An SDK according to or embodying the invention might include a number of assets or modules which enable developers to add virtual cameras and / or virtual audio listeners and / or physical microphones to a extended reality environment generated by an application. In Unity there are already provided virtual cameras which capture and display the world within the extended reality experience to the user. In the present invention, there may in addition (via an SDK) be provided virtual cameras which, instead of displaying to the user, generate video which is sent to a remote server for recording, encoding and streaming etc. Likewise virtual audio listeners and / or physical microphones may do the same for capturing and sending audio from within the experience.
[0069] An example of an SDK according to an embodiment of the invention may include a “package” which the developer will import to Unity or Unreal or other platform. Within the package the developer will be able to “drop” a management module and any number of virtual cameras into their virtual scenes.
[0070] A management module may handle all aspects of data, video and audio streaming whilst feeding this from the virtual cameras. The developer can also add any number of object trackers to virtual objects within their virtual scenes.
[0071] The developer can call upon the management module to allow them (amongst other things) to:
[0072] - track key events
[0073] - start / stop video recording on the remote server.
[0074] - change / control aspects of the stream itself such as video bitrate, resolution etc. The SDK may also include assets known as prefabs to do things such as provide user interface elements, physical virtual cameras which the user will be aware of and be able to interact with in the extended reality experience “to take selfies etc” and other 3D rendered objects such as a virtual watch or heart rate monitor.
[0075] The developer can choose to use these prefabs to aid the end users experience of streaming or not. The SDK may also include features to interact with third party platforms such as the ability for a user to see their viewers and comments on platforms such as Twitch.
[0076] Figure 2 is a flow diagram representing the development steps in implementing the invention, in the specific example of the Unity cross-platform game engine referred to above (though it will be appreciated that the skilled person will be able to adapt this to other platforms such that the invention is not limited to such an engine).
[0077] Firstly, the developer downloads a package comprising the SDK required to implement the invention in their Unity project 267. The package is then added to their Unity project 269. A corresponding module can then be added to a particular scene 271 . Within Unity a scene is an asset which contains parts of a game or application; a simple game might comprise one scene within which all action takes place, and a more complex game might comprise multiple scenes. Any number of scenes may be created in a Unity project, and a module according to the invention may be added to any of those scenes to allow video from it to be streamed.
[0078] Within the scene, the developer can choose objects to track (thereby generating data to be streamed as per 135) and add one or more virtual cameras to stream 273. Likewise one or more virtual audio listeners and / or physical microphones can be added. It is foreseen that a virtual camera within a project might be capable of recording video and audio (and indeed other data) but for the purposes of these descriptions they are presented as separate virtual components. Following this step, the developer can choose to combine events with camera presets. The developer can also select which authentication mechanism to use 275; in the example above a token-based authentication is preferred, and parameters of the authentication process can be defined at this time. Following this step, the developer can choose to integrate polling and chat functionalities, and other interactive aspects as described below. Once these steps are complete the project can be built 277, whereby the built Unity project comprises one or more virtual cameras (and microphones) which view and record the scene from the perspective of the virtual components and send it to a remote server which handles the streaming of same to the internet via browsers and / or social media platforms as described above with reference to Figure 1 .
[0079] There is now described, with reference to Figures 3 to 6, four examples showing the difference between the conventional approach and the inventive approach to: streaming extended reality games (Figure 3); streaming extended reality line-of- business (LOB) applications (Figure 4); recording extended reality games (Figure 5); and recording extended reality line-of-business (LOB) applications.
[0080] Figure 3A shows the conventional approach to streaming extended reality games and Figure 3B shows the inventive approach to streaming extended reality games.
[0081] As shown in Figure 3A, the conventional approach to streaming extended reality games (in the specific use case of an oculus / meta quest headset) begins with the user first ensuring that the PC (necessary for streaming as discussed above) and the headset are on the same local area network 379. The user then opens a browser and navigates to a predetermined address 381 , signs into their profile 383, and then opens the camera app on the headset 385 before choosing the “cast” option and selecting the PC to cast to 387. The headset then casts video to the browser on the PC. The user then opens a screencasting and streaming app 389, in this case OBS Studio, and sets up a scene which includes the browser window to which the video from the headset is being cast 391 . The user also configures a microphone in OBS to capture audio that is being cast to the browser 393. Once they have configured their streaming profile in OBS 395 (for example entering their Twitch credentials) the user can then finally open the game on the headset 397 and start streaming from OBS on the PC to their streaming profile 399.
[0082] In comparison, Figure 3B demonstrates that the procedure is significantly simplified in that the user opens a game on the headset 397B that has been built either according to the procedure described above with reference to Figure 2, or a similar / equivalent procedure, whereby virtual camera(s) and microphone(s) generate corresponding video and audio to be sent via a peer-to-peer link to a remote server which handles the adaptive bitrate streaming of composite video to the social media streaming service(s). The user then configures their streaming profile 395B (for example entering their Twitch credentials) and then presses a button in the VR environment to begin streaming 387B. Note that as an alternative the user can preconfigure their streaming profile beforehand. As mentioned above, polling and chat functionalities and other interactive aspects can be integrated, as described below, and these can be started and stopped at any time during streaming. These appear in the headset, for example in overlays, so that the user does not need to remove their headset to interact.
[0083] Figure 4A shows the conventional approach to streaming extended reality line-of- business (LOB) applications and Figure 4B shows the inventive approach to streaming extended reality line-of-business (LOB) applications.
[0084] Similarly, as shown in Figure 4A, the conventional approach to streaming extended reality line of business applications (again in the specific use case of an oculus / meta quest headset) requires the user to ensure that the PC (necessary for streaming as discussed above) and the headset are on the same local area network 479.
[0085] However, it is a prerequisite that the business in question must build a custom RTMP platform to underpin the LOB system 478.
[0086] As with the previous example, the user opens a browser on the PC and navigates to a predetermined address 481 , signs into their profile 483, and then opens the camera app on the headset 485 before choosing the “cast” option and selecting the PC to cast to 487. The headset then casts video to the browser on the PC. The user then opens a screencasting and streaming app 489, also OBS Studio in this example, and sets up a scene which captures the browser window and hence the video from the headset that is being cast 491 . The user configures a separate microphone in OBS to capture audio that is being cast to the browser 493. Once they have configured their RTMP streaming profile in OBS 495 the user can start streaming from OBS on the PC to their streaming profile 499, and remote viewers can log in to the RTMP server to view the stream 498.
[0087] In comparison, Figure 4B demonstrates that the procedure is significantly simplified in that the user opens an application on the headset 497B that has been built according to a similar / equivalent procedure to the procedure described above with reference to Figure 2, whereby virtual camera(s) and microphone(s) generate corresponding video and audio to be sent via a peer-to-peer link to a remote server which handles the adaptive bitrate streaming of composite video. The user then signs in to a streaming profile 495B (which may be provided on a proprietary portal) and then presses a button in the VR environment to begin streaming 487B. Viewers may then log into the proprietary portal to view the video stream 498B.
[0088] Figure 5A shows the conventional approach to recording extended reality games and Figure 5B shows the inventive approach to recording extended reality games.
[0089] Taking the specific use case of an oculus / meta quest headset again, the user opens the camera app on the headset 585, chooses the option to record in the headset OS 587, opens the game on the headset 597 while recording such that the ensuing gameplay session is recorded until the user decides to stop recording (typically when finished playing the game) 559. The recorded gameplay is then saved to the cloud (in this case the meta quest cloud service) 561 , where it can then (or later) be downloaded 562, saved on the user’s PC 564, where it can then be optionally edited and then uploaded to YouTube or similar for viewers to watch at any time 599.
[0090] In comparison, Figure 5B demonstrates that the procedure is significantly simplified in that the user opens an application on the headset 597B that has been built according to a similar / equivalent procedure to the procedure described above with reference to Figure 2, whereby virtual camera(s) and microphone(s) generate corresponding video and audio 587B to be sent via a peer-to-peer link to a remote server where it is recorded and is available to download immediately. In other words, the headset does not need to record or store the video and audio and then upload it when the session is over; the video and audio are effectively recorded on the fly. A user can thereby upload this directly and immediately to YouTube, TikTok or another connected video hosting service. The video can then (or later) be downloaded 562B, saved on the user’s PC 564B, where it can then be optionally edited and then uploaded to YouTube or similar for viewers to watch at any time 599B.
[0091] Figure 6A shows the conventional approach to recording extended reality line-of- business (LOB) applications and Figure 6B shows the inventive approach to recording extended reality line-of-business (LOB) applications.
[0092] This is similar to the conventional approaches described above in the context of streaming LOB content and recording gaming content, but for completeness the user opens the camera app on the headset 685 (an oculus / meta quest headset again in this example), chooses the option to record in the headset OS 687, opens the LOB application on the headset 697 while recording such that the ensuing application session is recorded until the user decides to stop recording (typically when the procedure or other activity is complete) 659. The recorded video is then saved to the cloud (in this case the meta quest cloud service) 661 , where it can then (or later) be downloaded 662, and saved on the user’s PC where it can then be optionally edited and then uploaded to a business platform for viewers to watch at any time 699.
[0093] In comparison, and similarly to previous examples as intimated above, Figure 6B demonstrates that the procedure is significantly simplified in that the user opens an application on the headset 697B that has been built according to a similar / equivalent procedure to the procedure described above with reference to Figure 2, whereby virtual camera(s) and microphone(s) generate corresponding video and audio to be sent via a peer-to-peer link to a remote server where it is recorded and is available to download immediately 662B. Again, the headset does not need to record or store the video and audio and then upload it when the session is over; the video and audio are recorded as they are received by the remote server. In this example, rather than downloading the video it can be downloaded to a business platform 662B by means of an API. Thereafter viewers can watch it at any time via the company platform.
[0094] In each of the above examples, the benefits of the inventive approach are clear; processing burdens are reduced, the cost and technical burden of access to the streaming community is significantly lowered, and hardware requirements and costs are lessened if not completely removed.
[0095] There are other significant benefits that follow from the manner in which the video to be streamed (or recorded) is generated, i.e. from virtual cameras within the extended reality experience on the headset, as compared to conventional approaches which are expressly and inevitably restricted to replicating the first person perspective of the user (because that is what is cast to the PC or, as below, rendered by remote server). The inventive approach allows developers and users to choose different perspectives which can include the conventional first person perspective but cameras can be selected to provide a third person perspective or to track objects such as virtual tools thus providing a headset streaming experience that cannot be realised currently. Multiple virtual cameras in the XR environment can be activated to contribute to a single video stream depending on what is happening in the game. This means that video streams are pre-curated according to specified potential actions that can occur within a game environment. Furthermore, developers may define logic for switching cameras or perspectives dynamically, for example based on in game events or, say, during cut scenes to provide bespoke cinematic experiences. As made clear earlier in the description, multiple video channels or channels comprising video from multiple virtual cameras can be streamed from a headset, so such logic can include activating one or more virtual cameras simultaneously as may be required.
[0096] Other known approaches to streaming XR content have relied on a local copy of the same game or software (or at least a version of it) on a streaming server which itself performs video encoding, composition and / or event processing. This increases bandwidth and latency. In comparison to those approaches, the inventive approach in which rendered data can be streamed directly from the headset to the server, provides for a more lightweight and scalable solution and avoids licensing and / or compatibility issues that can arise when game content or other software needs to be replicated on a server. The server is therefore only required to process and transmit, effectively functioning as a lightweight relay, though additional functionality can be added (without undue burden because of the already lightweight requirements).
[0097] In embodiments of the invention a user can use their voice to activate clipped recordings of a set time-period of the previous action. For example, during a football game a user may wish to record a short duration video of a goal or other meaningful in-game event. This recording can be stored on the headset and / or on a remote server, and may be sent automatically to the user’s TikTok, YouTube or other videohosting service(s) for rapid viewing. Voice-activated recording during gameplay minimises steps needed to be taken by the user.
[0098] Likewise, the user may be able to interact with other features enabled by the invention, such as creating Twitch polls or switching camera by voice. This allows the user a more continuous gameplay experience while retaining capacity to manage in-game capture and interaction. Voice recognition software, as may be provided within a voice-activation module, can activate and modify inward and outward communication from the headset according to user needs. Accordingly, any aspect of control described in the foregoing may advantageously (but not necessarily) be achieved via voice control.
[0099] The invention provides a method of streaming and / or recording video from an extended reality headset. The headset renders video in accordance with a first person extended reality experience of a user as usual, but in addition the headset also renders video from the perspective of at least one virtual camera. The video is sent to a remote server, where a composite video is produced which comprises the video from the at least one virtual camera. The composite video is stored and / or streamed from the remote server to one or more viewers. The headset can also capture and send one or more events from within the extended reality experience of the user, and the captured events (or representations of the captured events) can be included in the composite video. Likewise, the headset can render audio from one or more virtual microphones or listeners, also for inclusion in the composite video.
[0100] Throughout the specification, unless the context demands otherwise, the terms 'comprise' or 'include', or variations such as 'comprises' or 'comprising', 'includes' or 'including' will be understood to imply the inclusion of a stated integer or group of integers, but not the exclusion of any other integer or group of integers.
[0101] The foregoing description of the invention has been presented for the purposes of illustration and description and is not intended to be exhaustive or to limit the invention to the precise form disclosed. The described embodiments were chosen and described in order to best explain the principles of the invention and its practical application to thereby enable others skilled in the art to best utilise the invention in various embodiments and with various modifications as are suited to the particular use contemplated. Therefore, further modifications or improvements may be incorporated without departing from the scope of the invention as defined by the appended claims.
[0102] For example, the inventive concept is described with reference to gaming where streaming is widely popular, but it will be understood that it is equally applicable to streaming of content such as training programs and other interactive activities where it may be desirable to watch a live feed and observe data representative of or generated during the extended reality experience.
Claims
Claims1 . A method of streaming and / or recording from an extended reality headset comprising the steps of: in the headset, rendering video in accordance with a first person extended reality experience of a user; in the headset, simultaneously rendering video from the perspective of at least one virtual camera and sending it to a remote server; at the remote server, producing a composite video comprising the video from the at least one virtual camera; and storing the composite video and / or streaming the composite video from the remote server to one or more viewers.
2. The method of claim 1 , wherein the method comprises capturing one or more events within the extended reality experience of the user, one or more real world parameters, and / or one or more network parameters.
3. The method of claim 2, wherein the composite video comprises the events and / or parameters, or representations thereof.
4. The method of claim 1 or claim 2, wherein the method comprises streaming the events and / or parameters, and / or wherein the method further comprises storing the events and / or parameters separately from the composite video.
5. The method of any preceding claim, wherein rendering video in accordance with a first person extended reality experience of a user comprises rendering video from the perspective of a virtual camera corresponding to the left eye of a user and a virtual camera corresponding to the right eye of the user.
6. The method of any preceding claim, wherein the method also comprises generating audio in accordance with the first person extended reality experience of the user, simultaneously generating audio from the perspective of at least one virtual audio listener and / or physical microphone and sending itto the remote server, wherein the composite video further comprises the audio from the at least one virtual audio listener and / or at least one physical microphone.
7. The method of any preceding claim, wherein the virtual camera and the optional virtual audio listener and / or physical microphone correspond to a user’s first person perspective.
8. The method of claim 7, wherein the first person perspective corresponds to one of a left or right eye channel of the first person extended reality experience of the user.
9. The method of any preceding claim, wherein the or a virtual camera and the or an optional virtual audio listener and / or physical microphone correspond to a third party perspective.
10. The method of any preceding claim, wherein the or a virtual camera and the or an optional virtual audio listener and / or physical microphone is selected to follow an object such as a virtual tool within the environment.11 . The method of any preceding claim, wherein the position and / or focus of the or a virtual camera and the or an optional virtual audio listener and / or physical microphone are controlled by the user and / or by a viewer.
12. The method of any preceding claim, wherein the method comprises establishing a plurality of virtual cameras and optional virtual audio listeners and / or physical microphones, and selecting which of the virtual cameras and optional virtual audio listeners and / or physical microphones to render video and optionally generate audio, whereby the composite video is from the perspective of the selected virtual camera or cameras and optional virtual audio listener or listeners and / or physical microphone or microphones.
13. The method of claim 12, wherein the virtual camera or cameras and optional virtual audio listener or listeners and / or physical microphone or microphones are selected by the user and / or a viewer.
14. The method of claim 12, wherein the virtual camera or cameras and optional virtual audio listener or listeners and / or physical microphone or microphones are selected responsive to an event occurring within the extended reality experience of the user.
15. The method of any preceding claim, wherein the method comprises authenticating the user and / or the headset.
16. The method of claim 15, wherein the user and / or the headset are authenticated at the remote server.
17. The method of any preceding claim, wherein the method comprises selecting the remote server from a plurality of remote servers.
18. The method of claim 17, wherein the remote server is selected based on a predicted or measured connection performance.
19. The method of any preceding claim, wherein the method comprises establishing a peer-to-peer connection between the headset and the remote server.
20. The method of any preceding claim, wherein the method comprises determining an optimal route to the remote server.21 . The method of any preceding claim, wherein the rendered video and generated audio are sent from the headset to the remote server via a videoconference session.
22. The method of claim 21 , wherein the rendered video is sent via one or more WebRTC video channels and the generated audio is sent via one or more WebRTC audio channels.
23. The method of any preceding claim, wherein the composite video rendering is carried out at an adaptive bitrate, wherein the bitrate is selected and / or adjusted responsive to network parameters.
24. The method of any preceding claim, wherein streaming the video comprises presenting the video to one or more viewers via a web browser.
25. The method of any preceding claim, wherein streaming the video comprises starting a session at a streaming platform and sending the video to the streaming platform where the video may be rendered to one or more viewers, preferably in real-time.
26. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method of any of the preceding claims.
27. The computer program of claim 26, comprising a software development kit comprising a package to be imported into a game engine, wherein the package comprises a plurality of modules or assets including virtual cameras, virtual audio listener and / or physical microphones, event trackers, controls, virtual objects and which may provide interactive features such as chat and commenting, a management module to track key events, start and stop recording on the remote server, and control parameters such as resolution and bitrate of the rendered and / or composite video, and / or user interface elements28. A computer readable medium or data carrier comprising the computer program of claim 26 or claim 27.
29. A data carrier signal carrying the computer program of claim 26 or claim 27.
30. A system comprising an extended reality headset and a server, the extended reality headset and server configured, modified, adapted or arranged to perform the method of any of claims 1 to 25.
Citation Information
Patent Citations
Methods and systems for computer video game streaming, highlight, and replay
US20170157512A1
Streaming of composite alpha-blended ar / mr video to others
US20240064387A1
Methods and systems for game video recording and virtual reality replay
US9473758B1