Picture stream-pushing method and apparatus, and electronic device, computer-readable storage medium and computer program product
By selecting the target streaming method and encoding and displaying texture data, the dizziness problem of Windows-side VR applications pushing streaming to openxr head-mounted display devices is solved, and compatibility across hardware manufacturers and smooth virtual scene display is achieved.
Patent Information
- Application Number
- PCT/CN2024/137785
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-18
- Filing Date
- 2024-12-09
- Publication Date
- 2025-07-24
AI Technical Summary
In the prior art, VR applications on the Windows side cannot effectively push and stream to head-mounted display devices based on the openxr protocol, resulting in poor user experience, especially inaccurate information transmission of head postures, resulting in dizziness.
A picture streaming method has been developed. By selecting the target streaming method, the texture data of the virtual scene screen is obtained and encoded, and sent to the head-mounted display device for decoding and displaying. It supports two streaming methods: OpenXR and SteamVR, ensuring that the head position information is initialized to 0 to reduce dizziness.
It realizes effective streaming of virtual scenes on Windows, and is compatible with head-mounted display devices from various hardware manufacturers, improving user experience, reducing dizziness, and improving the smoothness and immersion of screen display.
Smart Images

Figure CN2024137785_24072025_PF_FP_ABST
Abstract
Description
Screen streaming method, device, electronic device, computer-readable storage medium, and computer program product
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is based on the Chinese patent application with application number 202410072006.3 and application date of January 18, 2024, and claims the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into this application as a reference. Technical Field
[0003] The present application relates to the field of virtual reality technology, and in particular to a method, device, electronic device, computer-readable storage medium, and computer program product for pushing images. Background Art
[0004] With the continuous development of virtual scene-related applications, more and more hardware manufacturers are beginning to use OpenXR's rendering and interaction protocols to realize the display and interaction of virtual images. For example, several hardware manufacturers have already embedded head-mounted display (HMD) software devices developed based on OpenXR in their hardware devices to adapt to their own hardware rendering and controller interaction.
[0005] In the related art, many VR applications on personal computers (PCs) push images and audio to the HMD through streaming. However, there is a lack of solutions for streaming VR application images from PCs to head-mounted displays based on the OpenXR protocol. Summary of the Invention
[0006] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium and computer program product for pushing images, which can push virtual scene images on a PC to a head-mounted display device.
[0007] The technical solution of the embodiment of the present application is implemented as follows:
[0008] The present application provides a method for pushing a video stream, which is performed by an electronic device. The method includes:
[0009] In response to a selection operation for multiple streaming methods, the selected streaming method is used as the target streaming method; based on the target streaming method, texture data of the virtual scene screen of the target application is obtained, wherein the target application is used to render the virtual scene screen; the texture data is encoded to obtain a texture encoding result; the texture encoding result is sent to a head-mounted display device, wherein the texture encoding result is used to trigger the head-mounted display device to perform the following processing: the texture encoding result is decoded to obtain the texture data, and the virtual scene screen is displayed based on the texture data.
[0010] The present invention provides a device for streaming video, including:
[0011] A streaming mode determination module is used to respond to a selection operation for multiple streaming modes and use the selected streaming mode as the target streaming mode; a texture data acquisition module is used to acquire texture data of a virtual scene screen of a target application based on the target streaming mode, wherein the target application is used to render the virtual scene screen; an encoding module is used to encode the texture data to obtain a texture encoding result; an encoding result sending module is used to send the texture encoding result to a head-mounted display device, wherein the texture encoding result is used to trigger the head-mounted display device to perform the following processing: decoding the texture encoding result to obtain the texture data, and displaying the virtual scene screen based on the texture data.
[0012] An embodiment of the present application provides an electronic device, comprising:
[0013] a memory for storing computer-executable instructions or computer programs;
[0014] The processor is used to implement the screen streaming method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0015] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the image streaming method provided in the embodiment of the present application when executed by a processor.
[0016] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the screen streaming method provided in the embodiment of the present application is implemented.
[0017] The embodiments of the present application have the following beneficial effects:
[0018] When streaming the virtual scene image of the target application to the head-mounted display device, the target streaming method is first determined from a plurality of streaming methods, and then based on the target streaming method, the texture data of the virtual scene image is obtained from the target application, and then the texture data is encoded to obtain a texture encoding result, and then the texture encoding result is sent to the head-mounted display device. The head-mounted display device can decode the texture encoding result to obtain the texture data, and display the virtual scene image based on the texture data. Therefore, when the embodiment of the present application is applied to a PC-side scenario, any one of the plurality of streaming methods can be selected to adapt to the head-mounted display devices developed by different hardware manufacturers, thereby enabling users to experience not only the applications carried by the head-mounted display device itself, but also the virtual scene images of the target application on the PC side when using a head-mounted display device developed by any hardware manufacturer. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] FIG1 is a schematic diagram showing the position and orientation of a head-mounted display according to an embodiment of the present application;
[0020] FIG2 is a flowchart of VR rendering and mobile phone Android rendering in related technologies;
[0021] FIG3 is a structural diagram of the picture streaming system architecture provided in an embodiment of the present application;
[0022] FIG4 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;
[0023] FIG5 is a flow chart of a method for pushing a video stream according to an embodiment of the present application;
[0024] FIG6 is a schematic diagram of cross-fusion of left and right eye images involved in this application;
[0025] FIG7 is a schematic diagram of a client interface of a DPT provided in an embodiment of the present application;
[0026] FIG8 is a schematic diagram of the client interface after DPT selects steamvr streaming according to an embodiment of the present application;
[0027] FIG9 is a schematic diagram of the client interface after DPT selects openvr streaming according to an embodiment of the present application;
[0028] FIG10 is a schematic diagram of a client interface of a game editor provided in an embodiment of the present application;
[0029] FIG11 is a schematic diagram of a level preview on a game editor provided in an embodiment of the present application;
[0030] FIG12 is a schematic diagram of the streaming process of openxr and steamvr provided in an embodiment of the present application;
[0031] FIG13 is a schematic diagram of the interaction process between the game side and the VRClient provided in an embodiment of the present application;
[0032] FIG14 is a schematic diagram of a rendering process of a game provided in an embodiment of the present application;
[0033] FIG15 is a schematic diagram of multiple textures created by a VRClient according to an embodiment of the present application;
[0034] FIG16 is a schematic diagram of a selection interface of a DPT graphics rendering device provided in an embodiment of the present application;
[0035] FIG17 is a schematic diagram of the rendering process within SteamVR provided in an embodiment of the present application;
[0036] FIG18 is a schematic diagram of the process of steamvr streaming provided by an embodiment of the present application;
[0037] FIG19 is a schematic diagram of left and right eye textures after game rendering provided by an embodiment of the present application;
[0038] FIG20 is a schematic diagram of left and right eye textures obtained after decoding by a head-mounted display according to an embodiment of the present application;
[0039] FIG21 is a schematic diagram of left and right eye textures displayed on the HMD provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0041] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0042] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0043] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0044] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0045] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0046] 1) OpenXR: A software development kit (SDK) for image rendering and device access in gaming and virtual reality (VR) applications. Its goal is to provide a unified interface that allows VR and augmented reality (AR) applications to run on multiple devices and platforms without having to write specific code for each system.
[0047] 2) SteamVR: Also known as OpenVR, it is the predecessor of the OpenXR standard. It is designed specifically to support virtual reality head-mounted display devices. It supports a variety of VR hardware and allows developers to write unified application logic for different brands of devices.
[0048] 3) D3D11 and D3D12: Graphics rendering SDKs for Windows, also known as DirectX or Direct3D (D3D). Currently, most PC games (personal computer games) use D3D11, while a few newer games use D3D12.
[0049] 4) Image acquisition: refers to copying and transmitting the texture drawn by the game, which is generally used for secondary rendering and video encoding.
[0050] 5) Push streaming: In the embodiments of this application, push streaming refers to the process of encoding the virtual scene image and sending it to the display end for display. For example, in VR games, push streaming refers to the audio and video encoding of the left and right eye images and sounds of the game and sending them to the head-mounted display (HMD) side for playback; the HMD side sends the tracking data back to the game side (such as the Windows side) for game control.
[0051] 6) Tracking data: A term used in VR, mainly including the posture information of the HMD and the two controllers.
[0052] 7) Encoding: Generally, there are two encoding methods: hardware encoding and software encoding. Hardware encoding refers to the technology that uses the computing power of the graphics processing unit (GPU) to achieve fast video encoding. Its encoding speed is faster than that of the central processing unit (CPU). Software encoding refers to the use of the computing power of the CPU for video encoding. Because the CPU has limited cache and computing power, it is not suitable for scenarios with high resolution and high real-time requirements.
[0053] 8) VRClient: In the embodiment of this application, it refers to a dynamic link library, which can be openvr_client.dll or openxr_client.dll, which are called by steamvr or games that implement the openxr protocol respectively. The main purpose is to connect with the game editor. For example, controlling the rendering frame rate of the game drawing, creating textures for the game and handing them over to the game for drawing are all performed in VRClient.
[0054] 9) Head-mounted display (HMD): This generally refers to VR glasses, sometimes also called helmets or head-mounted displays (HMDs). The monado SDK creates a software device for each manufacturer to connect to hardware input (controller buttons, head tracking data, etc.). By writing this hardware input to the game editor, game control is achieved.
[0055] 10) Quaternion: Quaternion is often used to represent rotation transformations in three-dimensional space. It can also be considered as a three-dimensional complex number area, where x, y, and z are imaginary parts and w is the real part (fixed value is 1).
[0056] 11) Pose information: This includes position and orientation. As shown in Figure 1, position represents the up-down, left-right, and forward-backward displacement of the headset. Orientation is usually associated with quaternions and is used to represent the headset's orientation in space: the x-axis (pitch), the y-axis (yaw), and the z-axis (roll).
[0057] 12) Monado: An open-source SDK based on the OpenXR protocol, Monado is a collaborative effort between several vendors to develop an Extended Reality (XR) SDK that has become a standard in the VR and XR industries. The Monado SDK can be compiled to run on HMDs and Windows.
[0058] VR games that are directly installed in the HMD are still a minority. Most VR games are installed on the PC side (such as the Windows side. This application uses the Windows side as an example, but it can also be applied to other operating systems on the PC side). The images and sounds of the VR games are pushed to the HMD for user experience through streaming. Therefore, the streaming software on the Windows side is indispensable in the development and operation of VR games. In the related art, the streaming software of the Windows system cannot be applied to the monado open source project based on the openxr protocol, resulting in the HMD using the monado SDK being unable to experience VR applications on the Windows side.
[0059] Additionally, the HMD can obtain the head position and orientation through the get_tracked_pose interface. The get_tracked_pose interface is primarily used for headset movement prediction, reducing the delay between actual head movement and image movement. Part of the nausea and vomiting experienced by users in VR comes from the lack of synchronization between the body and the image, so the get_tracked_pose interface is designed to ensure that the game and headset are as synchronized as possible. The HMD can also obtain the position and orientation of both eyeballs and the field of view (FOV) through the get_view_poses interface, where the FOV is determined by the HMD's hardware parameters. However, this method of obtaining pose information through multiple interfaces is implemented on the HMD side. No related solution exists for writing data during streaming (pushing). In other words, there is no Monado SDK implementation suitable for Windows-based streaming scenarios. If, during streaming, the HMD's head and eye pose information is written to the Windows game editor according to the above logic, severe nausea and vomiting can occur. When the user turns their head, the surrounding scene also rotates, creating a sensation similar to standing on a turntable. In a normal scenario, when the user turns his head, the surrounding scene should not rotate with him. Instead, the objects seen by the eyes move with the eyeballs and meet the brain's expectations. Motion sickness occurs when the movement of objects seen by the eyes is inconsistent with the rotation of the head muscles, which means it does not meet the brain's expectations.
[0060] There are significant differences between running games on OpenXR and on mobile devices. For example, the OpenXR runtime (XR runtime environment) on a specific hardware device, as shown in Figure 2, assumes that when the phone is not connected to an HMD, the game runs in Android mode. During the Android rendering process, a window is first created to display the game screen. The game screen is then rendered using the display engine and sent to the display (mobile interface) for display. When the phone is connected to an HMD, the SDK switches the mode, and the OpenXR game runs in VR mode. During VR rendering, both left-eye and right-eye textures are rendered, and the rendered left and right eye textures are sent to the display (HMD). OpenXR renders two textures for the VR display. During the rendering process, time prediction and pose prediction are performed. Time prediction predicts the rendering time and on-screen display time of the next VR frame. Pose prediction relies on the rendering time and on-screen display time provided by the time prediction to predict the displacement coordinates of the user's next pose. For example, if a user is showing an image appearing on top of an object, when the user has rotated their head (VR glasses are worn on their head), the object may still remain there, or move faster or slower than the user expected. However, by using the above-mentioned time prediction and pose prediction, a better correspondence between the image and the object can be achieved.
[0061] Similarly, when streaming a VR game, the Windows-based VR game also runs an OpenXR runtime in real time. It predicts the rendering time of the next frame of the VR game based on the graphics card rendering and the set rendering frame rate. Because there are two OpenXR runtimes (HMD and Windows), the head pose (position and orientation) predicted by the headset cannot be accurately combined with the Windows OpenXR runtime (because the rendering prediction time on the Windows side is calculated based on the graphics card and game logic). Therefore, after passing the head and eye pose information to the get_tracked_pose and get_view_poses interfaces of the Windows OpenXR runtime, the image seen by the eyes will be slightly fast-forwarded or delayed, causing dizziness.
[0062] As more and more manufacturers begin to use monado as the implementation SDK of openxr, it is imperative to develop a virtual HMD for its own hardware streaming based on the monado SDK. Therefore, based on at least one of the above problems existing in the relevant technology, the embodiment of the present application develops a virtual HMD device (XR_DirectPreview_Tool, referred to as DPT) based on the monado SDK for VR game streaming, which can support both openxr streaming and steamvr streaming, so as to fill the gap of monado in the PC-side VR game streaming function. The picture streaming method provided in the embodiment of the present application can be applied to DPT, which implements the openxr streaming method based on the monado SDK source code, receives trajectory data from the HMD, and encodes the texture data after the VR game rendering is completed and sends it to the HMD for display.
[0063] Among them, in the screen streaming method provided by the embodiment of the present application, first, in response to the selection operation for multiple streaming methods, the selected streaming method is used as the target streaming method; then, based on the target streaming method, the texture data of the virtual scene screen of the target application is obtained, wherein the target application is used to render the virtual scene screen; then, the texture data is encoded to obtain a texture encoding result; finally, the texture encoding result is sent to the head-mounted display device, so that the head-mounted display device decodes the texture encoding result to obtain the texture data, and displays the virtual scene screen based on the texture data. In this way, when the screen streaming method provided by the embodiment of the present application is applied to the Windows-side scenario, the streaming of the virtual scene screen on the Windows side can be achieved.
[0064] Here, first, an exemplary application of the screen push stream device of the embodiment of the present application is described, and the screen push stream device is an electronic device for implementing the screen push stream method. In one implementation, the screen push stream device (i.e., electronic device) provided by the embodiment of the present application can be implemented as a terminal or as a server. In one implementation, the screen push stream device provided by the embodiment of the present application can be implemented as a laptop, tablet computer, desktop computer, mobile phone, portable music player, personal digital assistant, dedicated messaging device, portable gaming device, intelligent robot, smart home appliance and smart car-mounted device, etc., any terminal with data processing and screen push stream functions; in another implementation, the screen push stream device provided by the embodiment of the present application can also be implemented as a server, wherein the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected by wired or wireless communication, which is not limited in the embodiment of the present application. Next, an exemplary application in which the screen streaming device is implemented as a server will be described.
[0065] Refer to Figure 3, which is a schematic diagram of the architecture of the picture streaming system provided in an embodiment of the present application. In order to support a picture streaming application, the virtual scene picture of the target application is streamed to the head-mounted display device through the picture streaming application. The terminal of the embodiment of the present application is at least installed with the picture streaming application and the target application (for example, various types of VR games). The picture streaming system 100 includes at least a head-mounted display device 500, a terminal 400, a network 300 and a server 200, wherein the server 200 is a server for the picture streaming application. The server 200 can constitute the picture streaming device of the embodiment of the present application, that is, the picture streaming method of the embodiment of the present application is implemented through the server 200. The terminal 400 is connected to the server 200 via the network 300, and the network 300 can be a wide area network or a local area network, or a combination of the two.
[0066] When streaming the virtual scene screen of the target application, the user can input a target streaming mode selection operation through the screen streaming application running on the terminal 400. The terminal 400 generates a selection instruction for the target streaming mode in response to the target streaming mode selection operation, and sends the selection instruction for the target streaming mode to the server 200 through the network 300. After receiving the selection instruction for the target streaming mode, the server 200 determines the target streaming mode from at least the first streaming mode and the second streaming mode in response to the selection instruction for the target streaming mode; then, when the target streaming mode is the first streaming mode, the server 200 obtains texture data of the virtual scene screen from the target application, wherein the target application is used to render the virtual scene screen; then, the texture data is encoded to obtain a texture encoding result; finally, the texture encoding result is sent to the head-mounted display device 500, so that the head-mounted display device 500 decodes the texture encoding result to obtain texture data, and displays the virtual scene screen based on the texture data.
[0067] In some embodiments, the terminal 400 can also execute the screen streaming method of the embodiment of the present application, that is, the user can input a selection instruction for the target streaming mode through the screen streaming application running on the terminal 400, and the terminal 400 responds to the selection instruction for the target streaming mode, and determines the target streaming mode from at least the first streaming mode and the second streaming mode; then, when the target streaming mode is the first streaming mode, the texture data of the virtual scene screen is obtained from the target application, wherein the target application is used to render the virtual scene screen; then, the texture data is encoded to obtain a texture encoding result; finally, the texture encoding result is sent to the head-mounted display device 500, so that the head-mounted display device 500 decodes the texture encoding result to obtain texture data, and displays the virtual scene screen based on the texture data.
[0068] The image streaming method provided in the embodiment of the present application can also be based on a cloud platform and implemented through cloud technology. For example, the above-mentioned server 200 can be a cloud server. The cloud server responds to the selection instruction for the target streaming mode and determines the target streaming mode from at least the first streaming mode and the second streaming mode; then, when the target streaming mode is the first streaming mode, the texture data of the virtual scene image is obtained from the target application, wherein the target application is used to render the virtual scene image; then, the texture data is encoded to obtain a texture encoding result; finally, the texture encoding result is sent to the head-mounted display device 500, so that the head-mounted display device 500 decodes the texture encoding result to obtain texture data, and displays the virtual scene image based on the texture data.
[0069] Referring to FIG. 4 , FIG. 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device shown in FIG. 4 includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the terminal 400 are coupled together via a bus system 440. It will be understood that the bus system 440 is used to implement connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in FIG. 4 , all various buses are labeled as the bus system 440.
[0070] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0071] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0072] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0073] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0074] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0075] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0076] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB).
[0077] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0078] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.
[0079] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in software. FIG4 shows a video streaming device 455 stored in a memory 450. The device 455 can be software in the form of a program or plug-in, and includes the following software modules: a streaming mode determination module 4551, a texture data acquisition module 4552, an encoding module 4553, and an encoding result transmission module 4554. These modules are logical and can be arbitrarily combined or further separated according to the functions implemented. The functions of each module will be described below.
[0080] In other embodiments, the apparatus provided in the embodiments of the present application may be implemented in hardware. As an example, the apparatus provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the image streaming method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0081] The screen streaming method provided in each embodiment of the present application can be executed by an electronic device, wherein the electronic device can be a server or a terminal, that is, the screen streaming method provided in each embodiment of the present application can be executed by a server, or by a terminal, or by interaction between a server and a terminal.
[0082] FIG5 is a schematic diagram of an optional flow chart of the method for pushing a picture stream provided in an embodiment of the present application. The following will be described in conjunction with the steps shown in FIG5 . As shown in FIG5 , the method for pushing a picture stream is described by taking the server as an example. The method includes the following steps S101 to S104:
[0083] Step S101 : In response to a selection operation on multiple streaming modes, a selected streaming mode is used as a target streaming mode.
[0084] Here, a plurality of streaming methods may include a first streaming method and a second streaming method. For example, a user can input a selection instruction for a target streaming method through a screen streaming application (DPT) on a terminal. DPT can support at least two streaming methods: a first streaming method and a second streaming method. Among them, the first streaming method can be an openxr streaming method, which is implemented based on the monado software development kit SDK. The second streaming method can be a steamvr streaming method. Users can choose between these two streaming methods, for example, they can choose to use the openxr streaming method or the steamvr streaming method as the target streaming method for streaming.
[0085] Step S102: Acquire texture data of the virtual scene screen of the target application based on the target streaming mode.
[0086] Here, the target application can be used to render a virtual scene image, wherein the target application can be various types of VR games.
[0087] In some embodiments, step S102 can be implemented in the following manner: when the target streaming mode is the first streaming mode (such as the openxr streaming mode), texture data of the virtual scene image is obtained from the target application.
[0088] For example, the target application can be a virtual reality related application, which can support HMD to display the virtual scene screen of the target application. For example, in the scenario where the user experiences the target application, DPT can be deployed on the user's personal PC (such as Windows), the target application can be a VR game, and the virtual scene screen of the target application is the game screen of the VR game. In the development scenario of the target application, DPT can be deployed on the cloud server, the target application can be a game editor of the VR game, and the virtual scene screen of the target application is the game screen of the VR game. Among them, the game editor of the VR game is a tool for making and developing VR games, such as game engines (such as Unity) and Unreal Engine (UE, Unreal Engine), etc. Game developers can edit a level scene of the VR game in the game editor, or edit the game scene of each frame in the level.
[0089] It should be noted that in the embodiments of this application, there is no specific order for launching the target application and selecting the target streaming method. That is, the user can first launch the target application and then launch the DPT to select the target streaming method; alternatively, the user can first launch the DPT to select the target streaming method and then launch the target application. This embodiment of the application does not specifically limit this.
[0090] In an embodiment of the present application, the target application will render the virtual scene screen during operation and obtain the texture data of the virtual scene screen. Among them, texture is an image format, which is mainly used for rendering the screen, such as the rendering of the virtual scene screen in the VR game. The texture data of the virtual scene screen may include the texture coordinates of each pixel on the image corresponding to the virtual scene screen, and the red, green, and blue (RGB, Red, Green, Blue) values. Exemplarily, when the target streaming mode is the openxr streaming mode, the texture data of the virtual scene screen of the VR game can be obtained directly from the game editor.
[0091] In some embodiments, the target application may be loaded with a first dynamic link library corresponding to the first streaming method. Then the above-mentioned acquisition of the texture data of the virtual scene picture from the target application can be achieved in the following ways: determine the rendering start time of the virtual scene picture through the first dynamic link library; when the current moment reaches the rendering start time, obtain the posture information from the head-mounted display device; send the posture information to the target application through the first dynamic link library, so that the target application renders the virtual scene picture based on the posture information and obtains the texture data of the virtual scene picture. For example, a monado SDK can be run on the head-mounted display device, and then the posture information can be obtained through the monado SDK, and the obtained posture information can be sent to the target application, so that the target application renders the virtual scene picture based on the posture information. In this way, when the virtual scene picture rendering is completed, the texture data of the virtual scene picture can be obtained from the target application through the first dynamic link library.
[0092] Here, the first dynamic link library can be openxr_client.dll, which is a VRClient that can be called by VR games that implement the OpenXR protocol. For example, after determining that the target streaming method is the OpenXR streaming method (i.e., the first streaming method), the video streaming application DPT can register the path of the openxr_client.dll for OpenXR integration in the Windows registry. The DPT can then write the path to the registry path. The path can point to a JSON file that specifies the runtime library that the target application can integrate with the OpenXR plug-in. This openxr_client.dll is the runtime library. The user can pre-set the rendering frame rate in the first dynamic link library. The first dynamic link library can then determine the rendering start time for each frame of the virtual scene based on the rendering frame rate. The frame rate is the number of frames per second (FPS), which refers to the number of times the target application draws the virtual scene per second, such as the number of times a game draws an image per second. The rendering frame rate can typically be set to 30FPS, 60FPS, 72FPS, 90FPS, etc.
[0093] In other embodiments, before streaming the virtual scene image of the target application, a plurality of textures to be rendered may be created through the first dynamic link library. For example, three textures to be rendered may be created through the first dynamic link library, assuming that they are textures 1 to 3, so that these three textures can be rendered synchronously, thereby improving rendering efficiency. Before sending the pose information to the target application through the first dynamic link library, the texture to be rendered corresponding to the virtual scene image may also be sent to the target application through the first dynamic link library, so that the target application renders the texture to be rendered based on the pose information to obtain texture data of the virtual scene image.
[0094] Here, multiple shared textures can be pre-created in the first dynamic link library. In this case, the shared texture can be a blank canvas. Shared textures can include various texture types, such as depth textures and color textures. Typically, virtual scene rendering only requires a color texture, so the first dynamic link library can determine the color texture from the multiple shared textures as the texture to be rendered. Each frame of the virtual scene corresponds to a texture to be rendered. For any frame of the virtual scene, when the rendering start time for that frame of the virtual scene is reached at the current moment, the texture to be rendered required for that frame of the virtual scene can be sent to the target application via the first dynamic link library. The first dynamic link library can then obtain the current pose information of the head-mounted display device and send this pose information to the target application. The target application renders the texture to be rendered based on this pose information, obtaining the texture data for that frame of the virtual scene. When the rendering of the virtual scene frame is completed, the texture data for that frame of the virtual scene can be obtained from the target application via the first dynamic link library.
[0095] The DPT can obtain the pose information of the head-mounted display device from the head-mounted display device and then write it into the first dynamic link library. Exemplarily, the head-mounted display device can be an HMD. The head-mounted display device is also implemented based on the monado SDK. The monado SDK running on the head-mounted display device can directly obtain head pose information, eye pose information, and information about the controller connected to the head-mounted display device, and send this head pose information, eye pose information, and controller information to the DPT. The eye pose information may include left eye pose information and right eye pose information. The texture data may include left eye texture data and right eye texture data. The left eye texture data is the texture data of the virtual scene image rendered by the target application using the rendered texture based on the left eye pose information. After the HMD decodes the left eye texture data, the left eye image can be provided to the user's left eye. The right eye texture data is the texture data of the virtual scene image rendered by the target application using the rendered texture based on the right eye pose information. After the HMD decodes the right eye texture data, the right eye image can be provided to the user's right eye. As shown in Figure 6, the HMD can obtain the left-eye texture data and right-eye texture data of the VR game through the get_view_poses interface in the monado SDK, draw the left-eye texture data to obtain the left-eye image, and draw the right-eye texture data to obtain the right-eye image. The left-eye image and the right-eye image are cross-fused to obtain the binocular image seen by the user on the HMD.
[0096] The embodiment of the present application obtains the posture information of the head-mounted display device through the first dynamic link library, and sends the posture information corresponding to each frame of the virtual scene image and the texture to be rendered to the target application, so that the target application renders the texture to be rendered based on the posture information to obtain texture data, obtains the texture data through the first dynamic link library, and encodes the texture data and sends it to the HMD for decoding and display, realizing the rendering and streaming process based on the monado SDK, so that when the user uses the head-mounted display device developed by the monado SDK, he can not only experience the applications carried by the head-mounted display device itself, but also experience the virtual scene images of the target application on the Windows side.
[0097] In some embodiments, determining the rendering start time of the virtual scene picture through the first dynamic link library can be achieved based on the following methods: determining the rendering frame rate of the target application through the first dynamic link library; determining the rendering start time of the virtual scene picture based on the rendering frame rate through the first dynamic link library.
[0098] Here, the user can write a preset rendering frame rate in the first dynamic link library in advance. For example, the user can receive the expected rendering frame rate written by the user in the first dynamic link library through a user interface or a configuration file. However, the actual rendering frame rate of the target application can be determined by the first dynamic link library based on various factors such as the performance of the graphics card, and the embodiments of the present application are not specifically limited here. For example, after receiving the rendering frame rate written by the user, the first dynamic link library can call the graphics card API (such as DirectX or OpenGL) to obtain the model, memory size and processing power of the current graphics card, and then perform a short performance benchmark test to estimate the maximum rendering frame rate that can be achieved under the current system configuration. For example, when the rendering frame rate written by the user is higher than the rendering frame rate that the graphics card can stably provide, the first dynamic link library can select a slightly lower rendering frame rate as the rendering frame rate of the target application to ensure stability and smoothness; when the graphics card can reach the rendering frame rate written by the user, the rendering frame rate written by the user can be used as the rendering frame rate of the target application. After determining the rendering frame rate of the target application, for any frame of the virtual scene picture, the rendering start time of the virtual scene picture can be determined through the first dynamic link library based on the rendering frame rate and the rendering end time of the previous frame of the virtual scene picture.
[0099] In some embodiments, the rendering start time of the virtual scene picture is determined based on the rendering frame rate through the first dynamic link library, which can be implemented in the following way: when the rendering of the i-th frame virtual scene picture is completed, the rendering end time of the i-th frame virtual scene picture is recorded through the first dynamic link library, where i is an integer greater than 0; for the i+1-th frame virtual scene picture, the rendering start time of the i+1-th frame virtual scene picture is determined based on the rendering frame rate and the rendering end time of the i-th frame virtual scene picture through the first dynamic link library.
[0100] For example, when the rendering of the first frame of the virtual scene picture is completed, the rendering end time of the first frame of the virtual scene picture can be recorded by the first dynamic link library (such as openxr_client.dll), and then for the second frame of the virtual scene picture, the rendering start time of the second frame of the virtual scene picture can be determined by openxr_client.dll based on the rendering frame rate and the rendering end time of the first frame of the virtual scene picture. When the rendering of the second frame of the virtual scene picture is completed, the rendering end time of the second frame of the virtual scene picture can be recorded by openxr_client.dll, and then for the third frame of the virtual scene picture, the rendering start time of the third frame of the virtual scene picture can be determined by openxr_client.dll based on the rendering frame rate and the rendering end time of the second frame of the virtual scene picture. And so on, the rendering start time of the last frame of the virtual scene picture can be determined.
[0101] In an embodiment of the present application, for the i-th frame of the virtual scene image, when the rendering of the i-th frame of the virtual scene image is completed, the rendering end time of the i-th frame of the virtual scene image can be recorded by the first dynamic link library. At the same time, the target application notifies the first dynamic link library that the rendering of the i-th frame of the virtual scene image is completed, and puts the texture data of the i-th frame of the virtual scene image into a cache queue, so that the first dynamic link library obtains the texture data of the i-th frame of the virtual scene image from the cache queue. Then, based on the rendering frame rate and the rendering end time of the i-th frame of the virtual scene image, the first dynamic link library determines the rendering start time of the i+1-th frame of the virtual scene image.
[0102] It should be noted that the embodiments of the present application do not specifically limit the specific method for determining the start time of the current frame rendering based on the rendering frame rate and the rendering end time of the previous frame of the virtual scene image. For example, taking 1s as a time unit and the rendering frame rate as the number of renderings within the time unit, if the rendering frame rate is 60fps (i.e., rendering 60 times per second), after obtaining the rendering end time of the i-th frame of the virtual scene image, the remaining time within 1s can be calculated based on the rendering end time. For example, assuming that the i-th frame is rendered from 0s, the subtraction result of 1s and the rendering end time can be used as the remaining time within 1s, and then the remaining number of renderings within 1s can be calculated based on 60fps. For example, the remaining time within 1s can be divided by the time required to render each frame at a frame rate of 60fps (i.e., 1 / 60 = The result of dividing the remaining time in 1s and the number of renders available in 1s is used as the remaining number of renders available in 1s. Subsequently, the time required for each rendering can be evenly distributed based on the remaining time in 1s and the number of renders available. For example, the result of dividing the remaining time in 1s and the number of renders available in 1s can be used as the evenly distributed time required for each rendering. Then, the rendering start time of the i+1th frame can be determined according to the rendering end time of the i-th frame and the time required for each rendering. For example, the sum of the rendering end time of the i-th frame and the evenly distributed time required for each rendering can be used as the rendering start time of the i+1th frame, thereby achieving smooth rendering.
[0103] That is to say, when the rendering start time of the i+1th frame virtual scene picture is reached at the current moment, the texture to be rendered of the i+1th frame virtual scene picture can be sent to the target application through the first dynamic link library, thereby starting the rendering of the i+1th frame virtual scene picture.
[0104] In the embodiment of the present application, the first dynamic link library is used to determine the rendering start time of the current frame of the virtual scene screen based on the rendering frame rate and the rendering end time of the previous frame of the virtual scene screen, so as to achieve smooth rendering of each frame of the virtual scene screen, improve the screen rendering effect, and further improve the display effect of the virtual scene screen after being pushed to the HMD, thereby improving the playback smoothness of the virtual scene screen.
[0105] In some embodiments, the posture information of the head-mounted display device may include head posture information and eye posture information. Sending the posture information to the target application via the first dynamic link library can be implemented in the following manner: sending the eye posture information to the target application via the first dynamic link library; initializing the head posture information to a set value (e.g., 0) via the first dynamic link library, and sending the head posture information initialized to 0 to the target application.
[0106] In an embodiment of the present application, the posture information of a head-mounted display device may include head posture information and eye posture information. The head posture information may include the position and orientation of the user's head, and the eye posture information may include the position, orientation, and field of view of the user's eyes. When the current moment reaches the rendering start time for the i-th frame of the virtual scene, the texture to be rendered for the i-th frame of the virtual scene is sent to the target application via a first dynamic link library. Then, the eye posture information returned by the head-mounted display device (HMD) is sent to the target application via the first dynamic link library. The position and orientation values in the head posture information returned by the HMD are filled with 0 via the first dynamic link library, and the head posture information, initialized to 0, is then sent to the target application via the first dynamic link library. Since the value of the head posture information received by the target application is 0, the target application will not use the head posture information to render the texture to be rendered. That is, the target application will render the texture to be rendered for the i-th frame of the virtual scene based solely on the eye posture information, obtaining texture data for the i-th frame of the virtual scene.
[0107] In the related art, when the monado SDK requires the HMD to obtain posture information based on the openxr protocol, it is necessary to pass in the head posture information and the eye posture information at the same time, but this method is only suitable for VR games running on the HMD, and is not suitable for streaming scenarios. Therefore, the embodiment of the present application initializes the head posture information to 0 when streaming, so that the VR game on the Windows side does not refer to the head posture prediction when rendering, but only refers to the existing eye posture. The head posture information is filled with 0, so that the VR game will not use the head posture information for prediction and rendering, which can ensure that the user will not feel dizzy when streaming, and the immersive feeling can achieve the same effect as steamvr.
[0108] In other embodiments, the above-mentioned step S102 can also be implemented in the following manner: when the target streaming mode is the second streaming mode, the texture data of the virtual scene screen of the target application is obtained through the streaming process corresponding to the second streaming mode.
[0109] Here, the second streaming method can be the steamvr streaming method. The streaming process corresponding to the second streaming method is the steamvr streaming application. After determining that the target streaming method is the second streaming method, the DPT can start the steamvr streaming application and obtain the texture data of each frame of the virtual scene picture of the target application through the steamvr streaming application. Then, the DPT encodes the texture data of each frame of the virtual scene picture to obtain the texture encoding result corresponding to each frame of the virtual scene picture. At the same time, the DPT can also encode the audio used for each frame of the virtual scene picture to obtain the audio encoding result. The texture encoding result and the audio encoding result are sent together to the head-mounted display device, so that the head-mounted display device decodes the texture encoding result and the audio encoding result to obtain texture data and audio, and plays the decoded audio while displaying the virtual scene picture to the user based on the texture data.
[0110] The embodiment of the present application connects to the SteamVR streaming application, and while implementing the OpenXR streaming method, it can also support the SteamVR streaming method to be compatible with HMD devices or VR games that are not developed using the OpenXR protocol, thereby improving the universality of the streaming method.
[0111] In some embodiments, obtaining texture data of the virtual scene screen of the target application through the streaming process corresponding to the second streaming method can be achieved based on the following methods: registering the second dynamic link library corresponding to the second streaming method to the streaming process corresponding to the second streaming method; obtaining head posture information from the head-mounted display device; writing the head posture information into the streaming process through the second dynamic link library, wherein the streaming process is used to render the texture data of the virtual scene screen based on the head posture information; and obtaining texture data from the streaming process through the second dynamic link library.
[0112] Here, the second dynamic link library corresponding to the second streaming mode can be openvr_client.dll, for example, a VRClient that can be called by SteamVR. For example, after determining that the target streaming mode is SteamVR, openvr_client.dll can be registered with the SteamVR streaming application. The SteamVR streaming application can directly connect to the target application. After the DPT obtains head pose information from the HMD, it sends the head pose information to the SteamVR streaming application via openvr_client.dll. The SteamVR streaming application can calculate eye pose information based on the head pose information. For example, when the head-mounted display device supports eye detection, the head pose information and the user's eye detection data can be combined to estimate the user's eye pose information (e.g., including the position and orientation of the user's left and right eyes). If the head-mounted display device does not support eye detection, the eye pose information can be approximated using a fixed eye offset (typically offset outward from the center of the user's head) and rendered to obtain texture data for the target application's virtual scene based on the head pose information and eye pose information. Finally, the texture data is obtained from the streaming process through the second dynamic link library.
[0113] Step S103: Encode the texture data to obtain a texture encoding result.
[0114] Here, after obtaining the texture data for each frame of the virtual scene image sent by the first dynamic link library, the DPT can encode the texture data to obtain a texture encoding result corresponding to each frame of the virtual scene image. The embodiment of the present application does not specifically limit the encoding method of the texture data. For example, the texture data can be encoded using software encoding or hardware encoding. At the same time, the audio used in each frame of the virtual scene image can also be encoded to obtain an audio encoding result.
[0115] Step S104: Send the texture encoding result to the head-mounted display device, wherein the texture encoding result is used to trigger the head-mounted display device to perform the following processing: decode the texture encoding result to obtain texture data, and display a virtual scene image based on the texture data.
[0116] Here, the DPT can send both the texture encoding result and the audio encoding result to the HMD. After receiving the texture encoding result and the audio encoding result, the HMD decodes the texture encoding result and the audio encoding result to obtain the texture data and audio of the virtual scene image, and displays the virtual scene image to the user based on the texture data while playing the decoded audio.
[0117] When the embodiment of the present application pushes the virtual scene screen of the target application to the head-mounted display device, the target push streaming method is determined from at least the first push streaming method and the second push streaming method. When the first push streaming method is used, the texture data of the virtual scene screen is first obtained from the target application, and then the texture data is encoded to obtain a texture encoding result. The texture encoding result is then sent to the head-mounted display device. The head-mounted display device can decode the texture encoding result to obtain the texture data and display the virtual scene screen based on the texture data. Since the first push streaming method is implemented based on the monado software development kit SDK, when the embodiment of the present application is applied to the Windows side scenario, the push streaming of the virtual scene screen on the Windows side can be implemented based on the monado SDK, so that the head-mounted display device developed based on the monado SDK can also adapt to the target application on the Windows side, thereby enabling the user to experience not only the application carried by the head-mounted display device itself, but also the virtual scene screen of the target application on the Windows side when using the head-mounted display device developed based on the monado SDK. In addition, the image streaming method provided by the embodiment of the present application, in the first streaming mode, when the target application renders the to-be-rendered texture to obtain the texture data of the virtual scene image, the head pose information of the HMD is initialized to 0, and only the eye pose information is used to render the to-be-rendered texture. This ensures that when the user turns their head during the streaming process, the surrounding scene will not rotate with it, preventing dizziness. The image streaming method provided by the embodiment of the present application can also support the second streaming mode to be compatible with HMD devices or VR games that are not developed using the OpenXR protocol, thereby improving the universality of the image streaming method.
[0118] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0119] The present application provides a method for pushing a video stream. This method can be applied to the DPT (XR_DirectPreview_Tool) video streaming application provided in the present application. The DPT is installed on the Windows side and can support both OpenXR streaming (the first streaming method) and SteamVR streaming (the second streaming method). The first streaming method in the video streaming method provided in the present application is implemented based on the Monado software development kit (SDK). Therefore, the present application can implement streaming of virtual scene images on the Windows side based on the Monado SDK.
[0120] The solution of the embodiment of the present application can be applied in the development scenario of VR games, and can also be applied in the scenario where users experience VR games. The following will take the development scenario of VR games as an example to illustrate the screen streaming method and screen streaming application DPT provided by the embodiment of the present application.
[0121] In the development scenario of VR games, DPT can be deployed on a cloud server for common use by multiple users, and the users can be game developers. Figure 7 is a schematic diagram of the client interface of DPT provided in an embodiment of the present application. Referring to Figure 7, the user can perform interactive operations through the selection box 701 in the DPT client interface. The interactive operation can be a click operation on "1.openxr push streaming" or "2.steamvr push streaming" under "Push streaming platform selection" in the selection box 701. Referring to Figure 8, when the user selects steamvr for push streaming, DPT will pull up the steamvr process 801 (the push streaming process corresponding to the second push streaming method) and register the VRClient (openvr_client.dll, the second dynamic link library) of the push streaming tool to the steamvr environment. Referring to Figure 9, openvr_client.dll can obtain the left and right eye textures (i.e., texture data of the virtual scene screen) of the VR game (i.e., the target application) from steamvr, and send the left and right eye textures to DPT for video encoding, and finally send the encoded data (i.e., texture encoding results) to the head display HMD for decoding and displaying the screen. DPT can also preview texture operations on the left and right eye textures. Referring to Figure 10, when the user selects openxr for streaming, the cloud server can open the game editor of the VR game (such as UE). Referring to Figure 11, the current level screen can be streamed directly to the head display HMD in the game editor. Users do not need to deploy the PC host again and wait for UE to package, they can directly experience the current level content.
[0122] In an embodiment of the present application, DPT can be compatible with both openxr and steamvr functional modules, among which openvr_client.dll is responsible for implementing steamvr's rendering and image acquisition; openxr_client.dll can put aside the steamvr platform to connect to the game editor and directly implement rendering and image acquisition from the game. Steamvr relies on the steam platform and has a large amount of game content, so the streaming tools of VR hardware manufacturers will give priority to accessing steamvr as its streaming platform. Compared with steamvr's openvr_client.dll, openxr_client.dll is more complex to implement. It connects directly to the game editor (or game) and is responsible for completing the rendering of the game, predicting the time for rendering the next frame, and writing events to the input device. There is no mature openxr rendering and streaming tool in the related art. See Figure 12, which is a schematic diagram of the streaming process of openxr and steamvr provided in an embodiment of the present application. DPT1201 is responsible for connecting to the VR device (HMD) 1202 via a Universal Serial Bus (USB) or wireless network (WIFI) to obtain the HMD's pose information and controller information. During OpenXR streaming, DPT1201 sends this pose and controller information to openxr_client.dll, which implements the OpenXR protocol and interfaces with the rendering logic of the game editor (UE, Unity, or a native app). openxr_client.dll sends this pose and controller information to the game editor, allowing it to render the left and right eye textures based on these pose and controller information. openxr_client.dll obtains the left and right eye textures from the game editor and sends them to DPT1201. DPT1201 encodes the left and right eye textures and sends them to HMD1202 along with the audio encoding. During SteamVR streaming, DPT1201 sends this pose and controller information to openvr_client.dll. openvr_client.dll implements the OpenVR protocol and connects to SteamVR. It sends pose and controller information to SteamVR, enabling SteamVR to render left and right eye textures based on these information. openvr_client.dll retrieves the left and right eye textures from SteamVR and sends them to DPT1201. DPT1201 encodes the left and right eye textures and sends them, along with the audio encoding, to HMD1202.
[0123] The openxr streaming method provided in the embodiment of the present application is described in detail below.
[0124] Refer to Figure 13, which is a schematic diagram of the interaction process between the game side and the VRClient provided in an embodiment of the present application. The steps shown in Figure 13 will be explained in conjunction with them. Step S1301: Request a GPU model. First, the game can request the GPU model of the terminal where the game is located from the VRClient through the xrGetD3D11GraphicsRequirementsKHR interface of OpenXR. Step S1302: Return the GPU model. The VRClient can return the GPU model to the game. Step S1303: Create a D3D device and request a texture. The game side can create a D3D device based on the GPU model returned by the VRClient and request a texture from the VRClient through the xrCreateSwapchain interface. Step S1304: Create a shared texture and return it to the game. The VRClient creates a shared texture and returns the shared texture used for left and right eye texture drawing to the game. Step S1305: Texture drawing and notification of drawing status. The game can draw on the shared texture returned by the VRClient to obtain the left and right eye textures. After the game has rendered each frame of the left and right eye texture images, it notifies the VRClient of the drawing status as complete through the xrEndFrame interface. Step S1306: Copy the shared texture. At this point, the VRClient can copy a texture from the game and pass it to the external process for use, that is, pass it to the DPT for encoding.
[0125] Referring to Figure 14, Figure 14 is a schematic diagram of the rendering process of the game provided in an embodiment of the present application, which will be explained in conjunction with the steps shown in Figure 14. Step S1401, record the rendering start time and frame number. First, the game can call the xrWaitFrame interface and the xrBenginFrame interface to enable the dynamic link library (VRClient) to record the rendering time and frame number of each frame image. Step S1402, obtain the texture that enters the waiting drawing mode. The game can obtain the renderable texture number from the VRClient through the xrAcquireSwapchainImage interface. Referring to Figure 15, multiple textures can be created in the VRClient (for example, including 3 textures, namely texture 1, texture 2 and texture 3), and the game can render them in turn, thereby realizing asynchronous copying of textures. Then, the game obtains the texture that enters the waiting drawing mode through the xrWaitSwapchainImage interface. The xrWaitSwapchainImage interface is used to wait for the selected texture to enter the waiting drawing mode. The specific implementation method can refer to the different implementation modes in OpenGL and Vulkan to allow the texture to enter the waiting drawing mode. Step S1403, obtain the pose information. The game calls the xrLocateViews and xrLocateSpace interfaces to obtain the pose information of the HMD, where the pose information is mainly used to simulate the changes in the head's up, down, left, right, and front and back positions. Step S1404, perform rendering. Based on the pose information obtained above, the game renders the texture in the waiting drawing mode (Execute Graphics Work). Step S1405, release the texture. After the rendering is completed, the game will call the xrReleaseSawapchain interface to make the VRClient release the texture. Step S1406, record the rendering completion time. The game calls the xrEndFrame interface to indicate that the current frame image has been drawn, notifies the VRClient that the texture has been rendered, and records the rendering completion time. In addition, users can control the rendering frame rate through the three XR interfaces: the xrWaitFrame interface, the xrBenginFrame interface, and the xrEndFrame interface. For example, game developers can configure a rendering frame rate of 60FPS or 90FPS in the VRClient. The above three interfaces can control the rendering cycle of the game according to the specified rendering frame rate. After calling the xrEndFrame interface, the game can place the rendered texture into a cache queue. The DPT then retrieves the rendered texture from the cache queue, encodes it, and sends it to the headset for decoding and playback. Referring to Figure 16 , the DPT provided in this embodiment of the application can render on both D3D11 and D3D12 graphics devices. Users can switch between different graphics devices as needed in the graphics device selection box 1601 in Figure 16 .
[0126] The following is a detailed description of the steamvr streaming method provided in the embodiment of the present application.
[0127] See Figure 17, which is a schematic diagram of the rendering process inside SteamVR provided by an embodiment of the present application. As shown in Figure 17, SteamVR's rendering is different from OpenXR. SteamVR implements two rendering modes of OpenVR and OpenXR internally, but ultimately provides external calls in the form of OpenVR to facilitate compatibility with old games and streaming tools. That is, SteamVR includes an OpenXR plugin (OpenXR plugin) and an OpenVR plugin (OpenVR plugin), but only uses the openvr_client.dll dynamic link library for streaming. See Figure 18, which is a schematic diagram of the flow of SteamVR streaming provided by an embodiment of the present application. As shown in Figure 18, openvr_client.dll is loaded into the SteamVR process by SteamVR, and then the DPT sends the Tracking data obtained from the head display to openvr_client.dll, and openvr_client.dll writes the Tracking data to the SteamVR interface. When the game finishes rendering a frame of image, it will notify SteamVR, and then SteamVR will send the left and right eye textures of the game to openvr_client.dll. Finally, openvr_client.dll will send the left and right eye textures of the game to DPT for video encoding and sending.
[0128] The image streaming method provided in the embodiment of the present application is implemented based on the monado SDK. Correspondingly, the external HMD end also runs the monado SDK to realize the rendering and operation of the VR game installed on the HMD end. The rendering method of the HMD end is described in detail below. The head display HMD also runs a set of openxr runtime (implementation of the monado SDK). Each hardware manufacturer will implement its own set of openxr runtime according to its own hardware environment and parameters. The openxr runtime will first calculate the pose information and predict the pose information through the hardware environment of the openxr runtime itself (screen refresh, rendering frame rate); then expose it to the user through the openxr interface (xrLocateViews and xrLocateSpace) for external calls. When pushing the stream, the embodiment of the present application can obtain the position and direction of the head and eyes through the two interfaces xrLocateViews and xrLocateSpace. After sending the pose information to the DPT on the Windows side, the game will draw the left and right eye textures based on the pose information. The drawn left and right eye textures can be seen in Figure 19. After the DPT encodes the left and right eye textures and sends them to the headset, the decoded textures are also left and right eye textures. However, after undergoing optical distortion processing by the headset hardware, they are rendered again and displayed on the screen, resulting in two images with elliptical and curved effects, as shown in Figure 20. See Figure 21. These left and right eye textures appear distorted on a flat surface, but appear more natural when projected to the human eye through the VR lens.
[0129] The image streaming method provided in this embodiment of the application can achieve a streaming experience consistent with the basic SteamVR experience. Before the Monado SDK is open-sourced and a HMD for streaming is released, the image streaming method provided in this embodiment of the application can be used to achieve the effects of posture movement and streaming. It should be noted that the image streaming method provided in this embodiment of the application can also transmit the predicted time information of the headset to the Windows end, allowing the game to refer to the headset's predicted time information and posture information when rendering.
[0130] It is understandable that in the embodiments of the present application, when user information and other related data are involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0131] The following further describes an exemplary structure of the image streaming device 455 provided in an embodiment of the present application implemented as a software module. In some embodiments, as shown in FIG4 , the software modules stored in the image streaming device 455 in the memory 450 may include:
[0132] The streaming mode determination module 4551 is configured to respond to a selection operation on multiple streaming modes and use the selected streaming mode as a target streaming mode.
[0133] The texture data acquisition module 4552 is configured to acquire texture data of a virtual scene screen of a target application based on a target streaming mode, wherein the target application is used to render the virtual scene screen.
[0134] The encoding module 4553 is configured to encode the texture data to obtain a texture encoding result.
[0135] The encoding result sending module 4554 is configured to send the texture encoding result to the head-mounted display device, wherein the texture encoding result is used to trigger the head-mounted display device to perform the following processing: decoding the texture encoding result to obtain texture data, and displaying a virtual scene picture based on the texture data.
[0136] In some embodiments, the multiple streaming methods include a first streaming method and a second streaming method, and the texture data acquisition module 4552 is further configured to obtain texture data of the virtual scene image from the target application when the target streaming method is the first streaming method.
[0137] In some embodiments, the target application is loaded with a first dynamic link library corresponding to the first streaming method; the texture data acquisition module 4552 is also configured to determine the rendering start time of the virtual scene picture through the first dynamic link library; when the rendering start time is reached at the current moment, the posture information is obtained from the head-mounted display device; the posture information is sent to the target application through the first dynamic link library, so that the target application renders the virtual scene picture based on the posture information and obtains the texture data of the virtual scene picture; when the rendering of the virtual scene picture is completed, the texture data of the virtual scene picture is obtained from the target application through the first dynamic link library.
[0138] In some embodiments, the posture information of the head-mounted display device includes head posture information and eye posture information; the texture data acquisition module 4552 is also configured to send the eye posture information to the target application through the first dynamic link library; initialize the head posture information to a set value through the first dynamic link library, and send the head posture information initialized to the set value to the target application.
[0139] In some embodiments, the screen streaming device 455 also includes a texture creation module, which is configured to create multiple textures to be rendered through a first dynamic link library; the texture data acquisition module 4552 is also configured to send the textures to be rendered corresponding to the virtual scene screen to the target application through the first dynamic link library, so that the target application renders the textures to be rendered based on the posture information to obtain the texture data of the virtual scene screen.
[0140] In some embodiments, the texture data acquisition module 4552 is further configured to determine the rendering frame rate of the target application through the first dynamic link library; and determine the rendering start time of the virtual scene image based on the rendering frame rate through the first dynamic link library.
[0141] In some embodiments, the texture data acquisition module 4552 is further configured to record the rendering end time of the i-th frame virtual scene picture through the first dynamic link library when the rendering of the i-th frame virtual scene picture is completed; i is an integer greater than 0; for the i+1-th frame virtual scene picture, the rendering start time of the i+1-th frame virtual scene picture is determined through the first dynamic link library based on the rendering frame rate and the rendering end time of the i-th frame virtual scene picture.
[0142] In some embodiments, the multiple streaming methods include a first streaming method and a second streaming method, and the picture streaming device also includes a second streaming module, which is configured to obtain the texture data of the virtual scene picture of the target application through the streaming process corresponding to the second streaming method when the target streaming method is the second streaming method.
[0143] In some embodiments, the second streaming module is further configured to register the second dynamic link library corresponding to the second streaming mode to the streaming process corresponding to the second streaming mode; obtain head posture information from the head-mounted display device; write the head posture information into the streaming process through the second dynamic link library; the streaming process is used to render texture data of the virtual scene image based on the head posture information; and obtain texture data from the streaming process through the second dynamic link library.
[0144] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions or computer program from the computer-readable storage medium and executes the computer-executable instructions or computer program, causing the electronic device to perform the above-described method for pushing images to a video stream according to the present invention.
[0145] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which stores computer-executable instructions or computer programs. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the screen streaming method provided by an embodiment of the present application, for example, the screen streaming method shown in Figure 5.
[0146] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0147] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0148] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0149] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0150] In summary, through the embodiments of the present application, the streaming of virtual scene images on the Windows side can be achieved based on the monado SDK.
[0151] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A method for pushing a video stream, which is executed by an electronic device. The method includes: In response to a selection operation for multiple video stream pushing methods, taking the selected video stream pushing method as the target video stream pushing method; Based on the target video stream pushing method, obtaining texture data of a virtual scene screen of a target application, where the target application is used to render the virtual scene screen; Encoding the texture data to obtain a texture encoding result; Sending the texture encoding result to a head-mounted display device, where the texture encoding result is used to trigger the head-mounted display device to perform the following processing: decoding the texture encoding result to obtain the texture data, and displaying the virtual scene screen based on the texture data.
2. The method according to claim 1, wherein The multiple video stream pushing methods include a first video stream pushing method and a second video stream pushing method; The obtaining texture data of a virtual scene screen of a target application based on the target video stream pushing method includes: When the target video stream pushing method is the first video stream pushing method, obtaining texture data of a virtual scene screen from the target application.
3. The method according to claim 2, wherein The first dynamic link library corresponding to the first video stream pushing method is loaded in the target application; The obtaining texture data of a virtual scene screen from the target application includes: Determining the start time of rendering of the virtual scene screen through the first dynamic link library; When the current time reaches the start time of rendering, obtaining pose information from the head-mounted display device; Sending the pose information to the target application through the first dynamic link library, where the pose information is used to trigger the target application to perform the following processing: rendering the virtual scene screen based on the pose information to obtain the texture data of the virtual scene screen; When the rendering of the virtual scene screen ends, obtaining the texture data of the virtual scene screen from the target application through the first dynamic link library.
4. The method according to claim 3, wherein, The pose information of the head-mounted display device includes head pose information and eye pose information; The sending the pose information to the target application through the first dynamic link library includes: Sending the eye pose information to the target application through the first dynamic link library; Initializing the head pose information to a set value through the first dynamic link library, and sending the head pose information initialized to the set value to the target application.
5. The method according to claim 3 or 4, wherein The method further includes: Creating a plurality of textures to be rendered through the first dynamic link library; Before sending the pose information to the target application through the first dynamic link library, the method further includes: Sending the texture to be rendered corresponding to the virtual scene screen to the target application through the first dynamic link library, where the texture to be rendered is used to trigger the target application to perform the following processing: rendering the texture to be rendered based on the pose information to obtain the texture data of the virtual scene screen.
6. The method according to claim 3 or 4, wherein The determining the start time of rendering of the virtual scene screen through the first dynamic link library includes: Determining the rendering frame rate of the target application through the first dynamic link library; Based on the first dynamic link library, determine the rendering start time of the virtual scene image according to the rendering frame rate.
7. The method according to claim 6, wherein The step of determining the rendering start time of the virtual scene image based on the first dynamic link library and the rendering frame rate includes: When the rendering of the i-th virtual scene image ends, record the rendering end time of the i-th virtual scene image through the first dynamic link library, where i is an integer greater than 0; For the (i + 1)-th virtual scene image, determine the rendering start time of the (i + 1)-th virtual scene image through the first dynamic link library based on the rendering frame rate and the rendering end time of the i-th virtual scene image.
8. The method according to any one of claims 1 to 7, wherein The multiple streaming methods include a first streaming method and a second streaming method; The step of obtaining the texture data of the virtual scene image of the target application based on the target streaming method includes: When the target streaming method is the second streaming method, obtain the texture data of the virtual scene image of the target application through the streaming process corresponding to the second streaming method.
9. The method according to claim 8, wherein, The step of obtaining the texture data of the virtual scene image of the target application through the streaming process corresponding to the second streaming method includes: Register the second dynamic link library corresponding to the second streaming method into the streaming process corresponding to the second streaming method; Obtain the head pose information from the head-mounted display device; Write the head pose information into the streaming process through the second dynamic link library, where the streaming process is used to render the texture data of the virtual scene image of the target application based on the head pose information; Obtain the texture data from the streaming process through the second dynamic link library.
10. The method according to claim 9, wherein The method further includes: Determine the eye pose information based on the head pose information; The step of writing the head pose information into the streaming process through the second dynamic link library, where the streaming process is used to render the texture data of the virtual scene image of the target application based on the head pose information, includes: Write the head pose information and the eye pose information into the streaming process through the second dynamic link library, where the streaming process is used to render the texture data of the virtual scene image of the target application based on the head pose information and the eye pose information.
11. The method according to any one of claims 1 to 10, wherein The method further includes: Encode the audio used in the virtual scene image to obtain an audio encoding result; The step of sending the texture encoding result to the head-mounted display device, where the texture encoding result is used to trigger the head-mounted display device to perform the following processing: decode the texture encoding result to obtain the texture data, and display the virtual scene image based on the texture data, includes: Send the texture encoding result and the audio encoding result to the head-mounted display device, where the texture encoding result and the audio encoding result are used to trigger the head-mounted display device to perform the following processing: decode the texture encoding result and the audio encoding result to obtain the texture data and the audio, and play the audio when displaying the virtual scene picture based on the texture data.
12. A video streaming device, the device comprising: A video streaming mode determination module configured to, in response to a selection operation for multiple video streaming modes, use the selected video streaming mode as the target video streaming mode; A texture data acquisition module configured to, based on the target video streaming mode, acquire texture data of a virtual scene picture of a target application, where the target application is used to render the virtual scene picture; An encoding module configured to encode the texture data to obtain a texture encoding result; An encoding result sending module configured to send the texture encoding result to the head-mounted display device, where the texture encoding result is used to trigger the head-mounted display device to perform the following processing: decode the texture encoding result to obtain the texture data, and display the virtual scene picture based on the texture data.
13. An electronic device, the electronic device comprising: A memory for storing computer-executable instructions or a computer program; A processor for implementing the video streaming method according to any one of claims 1 to 11 when executing the computer-executable instructions or the computer program stored in the memory.
14. A computer-readable storage medium storing computer-executable instructions or a computer program, the computer-executable instructions or the computer program implementing the video streaming method according to any one of claims 1 to 11 when executed by a processor.
15. A computer program product comprising computer-executable instructions or a computer program, the computer-executable instructions or the computer program implementing the video streaming method according to any one of claims 1 to 11 when executed by a processor.
Citation Information
Patent Citations
Method and device for controlling virtual reality scene to rotate, storage medium and electric device
CN108905202A
Stream pushing method and device for virtual scene, equipment and storage medium
CN116546228A
Picture stream pushing method and device, electronic equipment, storage medium and program product
CN117596377A
System-adaptive augmented reality
US20200364937A1