Virtual remote live streaming method, apparatus, device, and storage medium
By synchronously displaying virtual models on both the local and remote ends, and adjusting poses using pose monitoring and interactive data, the problem of low information transmission efficiency in existing virtual remote live streaming systems is solved. This enables remote linkage and collaboration between the local and remote ends, improving the interactivity and accuracy of virtual remote live streaming.
Patent Information
- Application Number
- CN202411029003.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2044-07-30
AI Technical Summary
Existing virtual remote live streaming systems suffer from low information transmission efficiency between local and remote ends, fail to provide immersive guidance, resulting in low efficiency for multi-person collaboration and poor virtual remote live streaming effects.
By synchronously displaying virtual models in a virtual scene on both local and remote ends, and adjusting the pose of the virtual models using pose monitoring and interactive data, remote linkage and collaboration between local and remote ends can be achieved. Feature tags are used to register virtual and real models, unifying the spatial coordinate systems of virtual and real scenes.
It improves the interactivity and effectiveness of virtual remote live streaming, enabling both local and remote ends to receive immersive operation and guidance, ensuring information synchronization, and enhancing the accuracy and interactivity of virtual remote live streaming.
Smart Images

Figure CN119052516B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of virtual reality technology, and more specifically, to a virtual remote live streaming method, apparatus, electronic device, and storage medium. Background Technology
[0002] In surgical procedures, Virtual Reality (VR) technology uses virtual models generated through 3D reconstruction or modeling techniques to provide doctors with an immersive experience of observing anatomical models. Doctors can interact with the models in a virtual environment and receive simulated feedback through remote virtual live streaming systems, thereby enhancing the real environment and creating virtual content.
[0003] Existing remote virtual live streaming systems in the medical field fall into two categories: one involves deploying multiple MR (Mixed Reality) devices and monocular / binocular cameras at the local hospital, while a display screen is deployed at the remote expert's location. A communication system then transmits the local feed to the remote display screen for presentation. The remote expert joins the MR device via video conferencing, guiding multiple doctors wearing MR devices through voice communication. However, this type of remote virtual live streaming system, due to its reliance on local collaboration and the remote expert's video conferencing, fails to provide immersive guidance to local doctors, resulting in low efficiency for multi-person collaboration and poor virtual remote live streaming effects.
[0004] Another approach involves deploying MR equipment and depth cameras at the local hospital, while using AR glasses at a remote expert's device. The depth camera captures, reconstructs, and transmits the local environment and the doctor's actions in real-time via a communication system to the remote expert's AR device for display. The expert then uses VR controllers to overlay virtual annotations and actions onto the local doctor's MR device, achieving remote-to-local interaction. However, this method only involves remote interaction between the expert and the local hospital. Furthermore, the local depth camera's acquisition and reconstruction presents a conflict between real-time display and detailed modeling at the remote end. This results in the local doctor's mixed reality device only viewing the remote expert's annotations and other information, leading to low information transmission efficiency and poor virtual remote live streaming effects.
[0005] As can be seen from the above, the problem of poor virtual remote live streaming effects in existing technologies still needs to be solved. Summary of the Invention
[0006] This application provides a virtual remote live streaming method, apparatus, electronic device, and storage medium, which can solve the problem of poor virtual remote live streaming effect in related technologies. The technical solutions are as follows: According to one aspect of this application, a virtual remote live streaming method is provided, characterized in that the method includes: constructing a virtual scene based on a real model in a real scene, wherein the virtual scene includes a virtual model corresponding to the real model; synchronously displaying the virtual model in the virtual scene on a local end and a remote end based on the pose of the real model in the real scene; continuously acquiring pose data generated by the local end for pose monitoring of the real model, and / or, interaction data generated by the remote end for remote interaction with the virtual model to adjust the pose of the virtual model.
[0007] According to one aspect of this application, a virtual remote live streaming method apparatus is provided, characterized in that the apparatus comprises: The scene construction module is used to construct virtual scenes based on real models in real scenes. The virtual scenes include virtual models that correspond to the real models. A synchronous display module is used to synchronously display the virtual model in the virtual scene on both the local and remote ends based on the pose of the real model in the real scene. The pose adjustment module is used to continuously acquire pose data generated by local terminal for pose monitoring of the real model, and / or interaction data generated by remote terminal for remote interaction of the virtual model, in order to adjust the pose of the virtual model.
[0008] In one exemplary embodiment, the scene construction module includes: A tag generation unit is used to generate feature tags for the virtual model; A registration unit is used to register the virtual model to a real model in a real scene based on the feature tags; The scene generation unit is used to generate virtual scenes based on the registered virtual models.
[0009] In one exemplary embodiment, the registration unit includes: The tag setting subunit is used to set feature tags for the corresponding virtual model for the real model; A detection subunit is used to detect the feature identifier in a real scene to determine the spatial coordinates of the real model carrying the feature identifier in the real scene; The registration subunit is used to register the virtual model corresponding to the feature marker to the real model at the spatial coordinates.
[0010] In an exemplary embodiment, the real model includes a camera model, and the virtual model includes a character model; the synchronized display module includes: A binding unit is used to bind the camera model in the real model to the character model in the virtual model; A coordinate system unit is used to construct a spatial coordinate system with the camera model as the base point to determine the relative spatial relationship between each of the real models, wherein the relative spatial relationship is used to indicate the spatial coordinates of each real model relative to the base point; The display unit is used to display the corresponding virtual models on the local end and the remote end based on the relative spatial relationship between the real models, with the character model as the base point. In one exemplary embodiment, the display unit includes: The coordinate system sub-unit is used to construct a spatial coordinate system with the character model in the virtual space as the base point, and / or to obtain the camera spatial coordinates mapped by the camera model in the virtual space, and to construct a spatial coordinate system with the camera spatial coordinates as the base point; The display subunit is used to display the corresponding virtual model at the spatial coordinates corresponding to each of the real models in the virtual scene, based on the relative spatial relationship between each of the real models.
[0011] In one exemplary embodiment, the pose adjustment module includes: The detection unit is used to perform pose monitoring on the real model locally and obtain the pose data of the real model. The update unit is used to update the pose of the corresponding virtual model based on the pose data of the real model.
[0012] In one exemplary embodiment, the pose adjustment module includes: An interaction unit is used to acquire interaction data based on remote interaction between a remote terminal and the virtual model; The prediction unit is used to predict the pose of the virtual model based on the interaction data and update the pose of the virtual model accordingly.
[0013] According to one aspect of this application, an electronic device includes at least one processor and at least one memory, wherein program instructions or code are stored in the memory; the program instructions or code are loaded and executed by the processor, causing the electronic device to implement the virtual remote live streaming method as described above.
[0014] According to one aspect of this application, a storage medium stores program instructions or code thereon, which are loaded and executed by a processor to implement a virtual remote live streaming method as described above.
[0015] The beneficial effects of the technical solution provided in this application are: In the above technical solution, by constructing a virtual model and displaying it synchronously on the local and remote ends, the interaction and collaboration of the virtual model are realized through the joint efforts of the local and remote ends, thereby achieving remote linkage and collaboration, which improves the interactivity and effect of virtual remote live streaming. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0017] Figure 1 This is a schematic diagram based on the implementation environment involved in this application; Figure 2 This is a flowchart illustrating a virtual remote live streaming method according to an exemplary embodiment; Figure 3 yes Figure 2 A flowchart of step 230 in one embodiment corresponds to the following example; Figure 4 This is a flowchart illustrating a generation optimization algorithm according to an exemplary embodiment; Figure 5 This is a flowchart illustrating a generation optimization algorithm according to an exemplary embodiment; Figure 6 This is a flowchart illustrating the calculation of line loss parameters according to an exemplary embodiment; Figure 7 yes Figure 4 In the corresponding embodiment, step 430 is a generative optimization algorithm; Figure 8 yes Figure 7 In the corresponding embodiment, step 730 is in a generative optimization algorithm; Figure 9 is a schematic diagram of the specific implementation of the network structure of an image reconstruction network in an application scenario. Figure 10 This is a structural block diagram of a virtual remote live streaming device according to an exemplary embodiment; Figure 11 This is a hardware structure diagram of a server according to an exemplary embodiment; Figure 12 This is a structural block diagram of an electronic device according to an exemplary embodiment. Detailed Implementation
[0018] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0019] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0020] Existing technologies generally employ a main grid-distributed power operation mode where users purchase electricity. Power is centrally controlled within the virtual reality environment, and interaction between multiple microgrids occurs through the main grid. However, due to the complex grid structure and challenging load scheduling, centralized control is ill-suited for information exchange and control between multiple microgrids, resulting in slow interaction speeds and low efficiency. Therefore, it is evident that related technologies still suffer from poor virtual remote live streaming effects between multiple microgrids.
[0021] Therefore, the virtual remote live streaming method provided in this application can improve the interaction efficiency between multiple micronets. Accordingly, the virtual remote live streaming method is applicable to a virtual remote live streaming device, which can be deployed on an electronic device. The electronic device can be a computer device configured with a von Neumann architecture, such as, but not limited to, desktop computers, laptops, servers, etc.
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0023] Figure 1 This is a schematic diagram of the implementation environment involved in a virtual remote live streaming method. The implementation environment includes a server 110, a local terminal 130, and a remote terminal 150.
[0024] Specifically, the local terminal 130 may include an MR device or a depth camera device, and the remote terminal 150 may include a VR device or a VR controller. The local terminal 130 and the remote terminal 150 are interconnected. This connection can be a network connection or other wired or wireless connection, and is not limited here.
[0025] Server 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Server 200 is an electronic device used to provide backend services; for example, in this implementation environment, server 200 provides virtual remote live streaming services to terminal 100.
[0026] Server 110 establishes communication connections in advance with each device in local terminal 130 and each device in remote terminal 130 through wired or wireless means, and realizes interaction with local terminal 130 and remote terminal 130 through communication connections.
[0027] Through the interaction between the local terminal 130 and the remote terminal 130 and the server 110, each node in the local terminal 130 and the remote terminal 130 sends data to the server 110. The transmitted data includes, but is not limited to: point cloud data collected by the depth camera from the real device, interactive data obtained by the VR controller, etc.
[0028] For server 110, after receiving the data uploaded by each node, it calls the virtual remote live streaming service to generate virtual scenes and virtual models, and then sends the virtual models to local terminal 130 and remote terminal 130 through communication connection, so that local terminal 130 and remote terminal 130 can perform synchronous remote linkage and collaboration based on the display of virtual models on MR devices and VR devices.
[0029] Please see Figure 2 This application provides a virtual remote live streaming method, which is applicable to electronic devices, and the electronic devices can be... Figure 1 The server 110 in the implementation environment is shown.
[0030] In the following method embodiments, for ease of description, the execution subject of each step of the method is an electronic device, but this does not constitute a specific limitation.
[0031] like Figure 2 As shown, the method may include the following steps: Step 210: Construct a virtual scene based on the real model in the real scene.
[0032] The virtual scene includes virtual models that correspond to the real model.
[0033] Specifically, the real-world scene refers to the actual environment in which the local device is located. The local device models the devices or objects in the real-world scene that are related to the virtual remote live broadcast using 3D reconstruction or modeling techniques, generating a virtual scene that includes virtual models. For example, point cloud data is collected from the real-world environment of the local device using a depth camera, and then the point cloud data is refined to reconstruct the virtual scene and various virtual models appearing in the virtual scene.
[0034] In one possible implementation, the real-world scenario is a human anatomical surgical environment. In this case, the virtual model includes a virtual model of human anatomy / pathology reconstructed from CT / MRI images, virtual models of the actual surgical instruments to be used during the live stream, and virtual models of the medical staff involved in the live stream.
[0035] Step 230: Based on the pose of the real model in the real scene, the virtual model in the virtual scene is displayed synchronously on the local and remote ends.
[0036] This invention combines the display requirements of both local and remote ends, displaying the poses of real models in the real world on a virtual model, so that both local and remote ends can accurately obtain the poses of each real model in the real scene based on the same virtual model.
[0037] The number of local and remote endpoints can be multiple, and there is no limit here.
[0038] Step 250: Continuously acquire pose data generated by local end for pose monitoring of real model, and / or interaction data generated by remote end for remote interaction of virtual model to adjust the pose of virtual model.
[0039] Specifically, during virtual remote live streaming, interactions between local users and the real model cause changes in the real model's posture, such as pose shifts, rotations, and deformations. The virtual model is then adjusted to match the real model's posture using pose data collected from the real model.
[0040] Meanwhile, remote users can interact with virtual models in the virtual scene, which will also cause changes in the pose of the virtual models. For example, remote users can input control commands to the virtual models in the virtual scene, causing the virtual models in the virtual scene to execute control commands and change their pose, and the pose changes of the virtual models will be synchronously displayed in the local virtual models.
[0041] Through the above process, the linkage between real and virtual models in the real scene is realized, enabling remote linkage and collaboration between the local and remote ends on the virtual model, and allowing both parties to obtain accurate feedback. This enables immersive operation and guidance of virtual reality devices, allowing both the local and remote ends to obtain rich pose information of the real model in the real scene, ensuring information synchronization between the local and remote ends, thereby improving the interactivity and effect of virtual remote live streaming.
[0042] In one exemplary embodiment, such as Figure 3 As shown, step 210 may include the following steps: Step 211: Generate feature labels for the virtual model.
[0043] Among them, the feature marker is a marker that uniquely corresponds to the virtual model. For example, the feature marker is the feature text information bound to the corresponding virtual model, or a QR code image generated based on the feature text information corresponding to the virtual model.
[0044] Step 213: Register the virtual model to the real model in the real scene based on feature tags.
[0045] By binding real models with feature tags, virtual models can be registered and bound to real models in both real and virtual scenes. In subsequent virtual remote live broadcasts, the registration information can be called at any time to quickly obtain the correspondence between virtual and real models.
[0046] In one exemplary embodiment, such as Figure 4 As shown, step 213 may include the following steps: Step 2131: Set feature labels for the corresponding virtual model for the real model.
[0047] In one possible implementation, based on the generation relationship between the real model and the virtual model, the feature identifier of the virtual model is set onto the real model that generated the virtual model.
[0048] Step 2133: Detect feature markers in the real scene to determine the spatial coordinates of the real model carrying the feature markers in the real scene.
[0049] Specifically, before a virtual live stream, feature markers corresponding to the virtual model can be set on the real model. During the virtual live stream, the real model can be identified and tracked by detecting these feature markers. For example, when the real scene is a human anatomical surgical environment, QR code image markers can be set on the human body model and surgical instrument model in the real scene. During the virtual live stream, a depth camera is used to capture the QR code image markers to determine their spatial coordinates. By recognizing the QR code, the specific real model at those spatial coordinates can be identified.
[0050] Step 2135: Register the virtual model corresponding to the feature marker to the real model at the spatial coordinates.
[0051] After determining the real model and its corresponding spatial coordinates in the real scene, the virtual model is registered to those spatial coordinates and bound to the real model. Specifically, after determining the virtual model and the registered real model in the real scene, the virtual model can be displayed at the registered spatial coordinates during the display process.
[0052] Through the above process, by marking the features of the virtual model, the virtual model is registered with the real model in the real scene. Through the virtual-real fusion interaction between the virtual model and the real model, the accuracy of the virtual model display in the virtual remote live broadcast is ensured, thus guaranteeing the effect of the virtual remote live broadcast.
[0053] Step 215: Generate a virtual scene based on the completed virtual model.
[0054] Specifically, the registered virtual model is displayed in a preset background to generate a virtual scene.
[0055] In the above process, by binding virtual models and real models to realize the generation of virtual scenes and virtual models within them, the interaction between virtual and reality is achieved, which can meet the different application needs of local and remote interaction in virtual remote live streaming.
[0056] In one exemplary embodiment, such as Figure 5 As shown, step 230 may include the following steps: Step 231: Bind the camera model in the real model to the character model in the virtual model.
[0057] It's important to note that during virtual remote live streaming, the remote and local terminals need to display the virtual model within the same coordinate system. Therefore, the coordinate systems of the remote and local terminals need to be unified relative to their respective viewpoints. Specifically, the camera model in the real-world model is a virtual model generated by the local terminal's camera view of the real scene, while the character model in the virtual model is a character model generated based on the remote terminal's perspective of observing the virtual environment. By binding the camera model in the real-world model and the character model in the virtual model, the observation perspectives of the remote and local terminals on the scene are relatively synchronized.
[0058] Step 233: Construct a spatial coordinate system with the camera model as the base point to determine the relative spatial relationship between each real model.
[0059] The relative spatial relationships are used to indicate the spatial coordinates of each real model relative to the base point. Specifically, by constructing a spatial coordinate system with the camera model as the base point, the spatial coordinates of the other real models relative to the camera model in the real scene are detected, and the relative distances and angular deviations between the real models are determined.
[0060] Step 235: Using the character model as a base point, display the corresponding virtual models on the local and remote ends according to the relative spatial relationship between the real models.
[0061] By constructing a spatial coordinate system identical to that in the real scene using the character model as the base point, and displaying the virtual model in the corresponding position, a virtual scene is constructed and displayed on the remote end, allowing the remote end to observe each virtual model in the virtual scene.
[0062] Meanwhile, by displaying the virtual model at the corresponding location in the real scene and displaying it on the local device, the local device can observe the virtual model in the real scene.
[0063] In one possible implementation, the virtual model is displayed locally via an MR device, and the virtual model is displayed remotely via a VR device.
[0064] In one exemplary embodiment, such as Figure 6 As shown, step 235 may include the following steps: Step 2351: Construct a spatial coordinate system with the character model in the virtual space as the base point, and / or obtain the camera spatial coordinates mapped by the camera model in the virtual space, and construct a spatial coordinate system with the camera spatial coordinates as the base point.
[0065] Since the character model is already bound to the camera model, a spatial coordinate system that is consistent with the real scene can be constructed by using the coordinates of either one in the virtual space as a base point.
[0066] Step 2353: Based on the relative spatial relationship between each real model, display the corresponding virtual model at the spatial coordinates corresponding to each real model in the virtual scene.
[0067] Since the spatial coordinate systems in the real scene and the virtual scene are unified, the corresponding virtual model is generated in the virtual scene based on the relative spatial relationship between each real model, which can accurately display the relative relationship between each real model in the real scene.
[0068] In one possible implementation, during the display process between the remote and local ends, real-time streaming is performed through desktop programs on both the local and remote ends, and the virtual model is rendered by a rendering workstation to achieve large-capacity display of the virtual model.
[0069] Through the above process, the spatial coordinate system in the virtual scene and the real scene is unified, which ensures the accurate generation of the virtual model in the virtual scene and improves the accuracy of virtual remote live broadcast.
[0070] In one exemplary embodiment, such as Figure 7 As shown, step 250 may include the following steps: Step 251: Perform pose monitoring on the real model on the local end to obtain the pose data of the real model.
[0071] Specifically, pose data includes model size, coordinates, and rotation angles.
[0072] During virtual remote live streaming, the local end continuously monitors the real model to determine the pose changes of the real model.
[0073] In one possible implementation, by fixing the Vuforia image onto the corresponding real model, the local user can scan the Vuforia image in real time through the camera model while operating the real model, thereby confirming the pose change of the real model relative to the camera model and realizing the monitoring of the real model.
[0074] Step 253: Update the pose of the corresponding virtual model in the virtual scene based on the pose data of the real model.
[0075] After confirming the pose changes of the real model, the latest changes of the virtual model are synchronized in real time on both the remote and local ends.
[0076] Through the above process, the virtual model is updated based on the local terminal's operations, information sharing is achieved between the local and remote terminals, and the accuracy of virtual remote live streaming is improved.
[0077] In one exemplary embodiment, such as Figure 8 As shown, step 250 may also include the following steps: Step 255: Obtain interactive data based on remote interaction between the remote terminal and the virtual model.
[0078] Specifically, after receiving the virtual model, the remote end can manipulate the virtual model in the virtual scene to change its pose.
[0079] Step 257: Predict the pose of the virtual model based on the interactive data, and update the pose of the corresponding virtual model.
[0080] By updating the pose of the virtual model, the information obtained by the remote end and the local end can be synchronized.
[0081] Through the above process, the virtual model is updated based on the operations of the remote end, realizing information sharing between the local and remote ends and improving the accuracy of virtual remote live streaming.
[0082] Figure 9a This illustration shows a specific implementation diagram of this application in an application scenario. Figure 9a In this setup, the local device is a mixed reality device, and the remote device is a virtual reality device. The mixed reality device displays virtual models within a real-world scene, while the virtual reality device displays virtual models within a virtual scene.
[0083] During virtual remote live streaming, a virtual scene is constructed using the background space on the local end during the live stream, and a corresponding virtual model is constructed using real objects such as preoperative images to reconstruct human anatomy / pathology models and real surgical instruments to be used during the operation.
[0084] By creating two Unity projects, one for mixed reality devices and one for virtual reality devices, the virtual scenes are imported into the virtual reality Unity application. All virtual models are then imported into both the mixed reality and virtual reality Unity applications. A hardware connectivity application is established on a photon cloud server to visualize the number of connected devices and provides software-independent identifiers (SIOs) to support the application's connection. Connection requests are initiated on both the local and remote ends by calling connection functions. By setting the same SIO in the application, the connection between the mixed reality and virtual reality devices is achieved.
[0085] By setting the virtual scene as the background of the remote end, and using the position of a depth camera in the virtual scene as the initial position of the main camera of the virtual reality Unity program, the mixed reality device starts the program at the spatial coordinates of the depth camera to complete the unification of the spatial coordinate system. This enables the local doctor and the remote expert to calibrate their positions in the virtual world, and allows both devices to observe the virtual model and the virtual character model and actions corresponding to the other party's position in real time.
[0086] The local mixed reality device achieves virtual-real interaction, such as registration and tracking of the virtual model and the real model, by scanning image markers on the real model. The pose of the virtual model is synchronized to the remote expert.
[0087] The remote expert's virtual reality (VR) device allows manipulation of the virtual model via controllers, changing its pose. The VR device streams in real-time through a Steam VR-compatible desktop application, with model rendering taking place on a rendering workstation and display and interaction occurring on the VR device itself. The local mixed reality (MVR) device manipulates real surgical instruments, tracking the virtual surgical instrument model through image markers. This MVR device also streams in real-time with the desktop application, with model rendering on a rendering workstation and display and interaction occurring on the MVR device.
[0088] Changes in the virtual model are synchronized in real time to the virtual reality device on the remote end and the mixed reality device on the local end.
[0089] In this application scenario, the realistic effect of the mixed reality device is as follows: Figure 9b As shown in the upper half, the realistic effect of virtual reality devices is as follows: Figure 9b As shown in the lower part, this application enables immersive operation and guidance of virtual reality devices through the interaction of the virtual model between the local and remote ends. This allows both the local and remote ends to obtain rich pose information of the real model in the real scene, ensuring information synchronization between the local and remote ends, thereby improving the interactivity and effect of virtual remote live streaming.
[0090] The following are embodiments of the apparatus described in this application, which can be used to execute the virtual remote live streaming method involved in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments of the virtual remote live streaming method involved in this application.
[0091] Please see Figure 10 This application provides a virtual remote live streaming device 1000, including but not limited to: a scene construction module 1010, a synchronous display module 1030, and a pose adjustment module 1050.
[0092] The scene construction module 1010 is used to construct a virtual scene based on the real model in the real scene. The virtual scene includes a virtual model corresponding to the real model.
[0093] The synchronous display module 1030 is used to synchronously display the virtual model in the virtual scene on both the local and remote ends based on the pose of the real model in the real scene.
[0094] The pose adjustment module 1050 is used to continuously acquire pose data generated by local end for pose monitoring of the real model, and / or interaction data generated by remote end for remote interaction of the virtual model in order to adjust the pose of the virtual model.
[0095] It should be noted that the virtual remote live streaming device provided in the above embodiments is only illustrated by the division of the above functional modules when performing virtual remote live streaming. In actual applications, the above functions can be assigned to different functional modules as needed. That is, the internal structure of the virtual remote live streaming device will be divided into different functional modules to complete all or part of the functions described above.
[0096] Furthermore, the virtual remote live streaming device and the virtual remote live streaming method provided in the above embodiments belong to the same concept, and the specific way in which each module performs operations has been described in detail in the method embodiments, and will not be repeated here.
[0097] Figure 11 A schematic diagram of the structure of a server is shown according to an exemplary embodiment. This server is suitable for... Figure 1 The server 110 in the implementation environment is shown.
[0098] It should be noted that this server is merely an example adapted to this application and should not be construed as providing any limitation on the scope of use of this application. Nor should this server be interpreted as requiring or depending on any specific feature. Figure 11 One or more components of the exemplary server 2000 shown.
[0099] The hardware architecture of Server 2000 can vary significantly due to differences in configuration or performance, such as... Figure 11 As shown, the server 2000 includes: a power supply 210, an interface 230, at least one memory 250, and at least one central processing unit (CPU) 270.
[0100] Specifically, power supply 210 is used to provide operating voltage for the various hardware devices on server 2000.
[0101] Interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices. For example, to perform... Figure 1 The diagram illustrates the interaction between terminal 100 and server 200 in the implementation environment.
[0102] Of course, in other examples adapted in this application, interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input / output interface 235, and at least one USB interface 237, etc. Figure 11 As shown, this does not constitute a specific limitation.
[0103] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored on it include the operating system 251, application programs 253, and data 255, etc., and the storage method can be temporary storage or permanent storage.
[0104] The operating system 251 is used to manage and control the various hardware devices and application programs 253 on the server 2000, so as to enable the central processing unit 270 to perform calculations and processing on the massive data 255 in the memory 250. It can be Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0105] Application 253 is a program instruction or code based on operating system 251 that performs at least one specific task, and may include at least one module ( Figure 11 (Not shown), each module can contain program instructions or code for server 2000. For example, a virtual remote live streaming device can be considered as application 253 deployed on server 2000.
[0106] Data 255 can be photos, pictures, etc. stored on a disk, or it can be the target input image, etc., stored in memory 250.
[0107] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read program instructions or code stored in the memory 250, thereby performing operations and processing on massive amounts of data 255 stored in the memory 250. For example, a virtual remote live streaming method can be implemented by the central processing unit 270 reading a series of program instructions or code stored in the memory 250.
[0108] Furthermore, this application can also be implemented through hardware circuits or a combination of hardware circuits and software. Therefore, the implementation of this application is not limited to any specific hardware circuit, software, or combination thereof.
[0109] Please see Figure 12 This application provides an electronic device 4000, which may include: a desktop computer, a laptop computer, a server, etc.
[0110] exist Figure 12 In this context, the electronic device 4000 includes at least one processor 4001 and at least one memory 4003.
[0111] The data interaction between the processor 4001 and the memory 4003 can be achieved through at least one communication bus 4002. This communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0112] Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one unit, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.
[0113] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0114] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program instructions or code in the form of instructions or data structures and accessible by the electronic device 400, but not limited thereto.
[0115] The memory 4003 stores program instructions or code, and the processor 4001 can read the program instructions or code stored in the memory 4003 through the communication bus 4002.
[0116] When the program instructions or code are executed by the processor 4001, the virtual remote live streaming method in the above embodiments is implemented.
[0117] Furthermore, this application provides a storage medium storing program instructions or code, which is loaded and executed by a processor to implement the virtual remote live streaming method described above.
[0118] This application provides an application product, which includes program instructions or code stored in a storage medium. The processor of an electronic device reads the program instructions or code from the storage medium, loads and executes the program instructions or code, enabling the electronic device to implement the virtual remote live streaming method described above.
[0119] Compared with related technologies, this application achieves the linkage between real and virtual models in a real scene, enabling remote collaborative operation of the virtual model by both local and remote ends, and providing both parties with accurate feedback. This allows for immersive operation and guidance of virtual reality devices, enabling both local and remote ends to obtain rich pose information of the real model in the real scene, ensuring information synchronization between the local and remote ends, thereby improving the interactivity and effect of virtual remote live streaming. By marking the features of the virtual model, the virtual model is registered with the real model in the real scene. Through the virtual-real fusion interaction between the virtual and real models, the accuracy of the virtual model display in virtual remote live streaming is ensured, guaranteeing the effect of virtual remote live streaming. The application unifies the spatial coordinate systems in the virtual and real scenes, ensuring the accurate generation of the virtual model in the virtual scene and improving the accuracy of virtual remote live streaming. It updates the virtual model based on the operation of the remote end, realizing information sharing between the local and remote ends, and improving the accuracy of virtual remote live streaming. Finally, it updates the virtual model based on the operation of the local end, realizing information sharing between the local and remote ends, and improving the accuracy of virtual remote live streaming.
[0120] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0121] The above are only some embodiments of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A virtual remote live streaming method, characterized in that, The method includes: A virtual scene is constructed based on a real model in a real scene, wherein the virtual scene includes a virtual model corresponding to the real model; the construction of the virtual scene based on the real model in the real scene specifically includes: generating feature tags for the virtual model; registering the virtual model to the real model in the real scene based on the feature tags; and generating the virtual scene based on the registered virtual model. The real model includes a camera model, and the virtual model includes a character model. The camera model in the real model is a virtual model generated by a camera on the real scene from the local end, and the character model in the virtual model is a character model generated based on the perspective of the remote end observing the virtual environment. Binding the camera model in the real model and the character model in the virtual model synchronizes the observation perspectives of the scene between the remote end and the local end. The mixed reality device on the local end registers and tracks the virtual model by scanning image markers on the real model. The mixed reality device on the local end can operate real surgical instruments and tracks virtual surgical instrument models through image markers. Based on the pose of the real model in the real scene, the virtual model in the virtual scene is synchronously displayed on both the local and remote ends. Specifically, this includes: binding the camera model in the real model to the character model in the virtual model; constructing a spatial coordinate system using the camera model as a base point to determine the relative spatial relationship between each real model, where the relative spatial relationship indicates the spatial coordinates of each real model relative to the base point; and displaying the corresponding virtual models on both the local and remote ends based on the relative spatial relationship between each real model and the character model as a base point. The pose data generated by local terminal for pose monitoring of the real model and remote terminal for remote interaction of the virtual model are continuously acquired to adjust the pose of the virtual model. The step of continuously acquiring pose data generated by local end for pose monitoring of the real model to adjust the pose of the virtual model includes: responding to the interactive operation of the real model on the local end, performing pose monitoring of the real model on the local end, and continuously acquiring the pose data of the real model; updating the pose of the virtual model displayed on the local end and the remote end according to the pose data of the real model. The step of continuously acquiring interaction data generated by the remote terminal in remote interaction with the virtual model to adjust the pose of the virtual model includes: acquiring interaction data based on the remote terminal's remote interaction with the virtual model; predicting the pose of the virtual model based on the interaction data; and updating the pose of the virtual model displayed on the remote terminal and the local terminal.
2. The method as described in claim 1, characterized in that, The step of registering the virtual model to a real model in a real scene based on the feature tags includes: For the real model, set corresponding feature tags for the virtual model; Detect the feature markers in a real-world scene to determine the spatial coordinates of the real model carrying the feature markers in the real-world scene; Register the virtual model corresponding to the feature marker to the real model at the spatial coordinates.
3. The method as described in claim 1 or 2, characterized in that, The step of displaying the corresponding virtual models on the local and remote ends based on the relative spatial relationships between the real models, using the character model as a base point, includes: A spatial coordinate system is constructed using the character model in the virtual space as the base point, and / or, the camera spatial coordinates mapped by the camera model in the virtual space are obtained, and a spatial coordinate system is constructed using the camera spatial coordinates as the base point; Based on the relative spatial relationship between the real models, the corresponding virtual model is displayed in the virtual scene at the spatial coordinates corresponding to each real model.
4. A virtual remote live streaming method apparatus, used to implement the method as described in any one of claims 1-3, characterized in that, The device includes: The scene construction module is used to construct virtual scenes based on real models in real scenes. The virtual scenes include virtual models that correspond to the real models. A synchronous display module is used to synchronously display the virtual model in the virtual scene on both the local and remote ends based on the pose of the real model in the real scene. The pose adjustment module is used to continuously acquire pose data generated by local terminal for pose monitoring of the real model, and / or interaction data generated by remote terminal for remote interaction of the virtual model, in order to adjust the pose of the virtual model.
5. An electronic device, characterized in that, include: At least one processor and at least one memory, wherein, The memory stores program instructions or code; The program instructions or code are loaded and executed by the processor, causing the electronic device to implement the virtual remote live streaming method as described in any one of claims 1 to 3.
6. A storage medium storing program instructions or code thereon, characterized in that, The program instructions or code are loaded and executed by the processor to implement the virtual remote live streaming method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Method, device and equipment for virtual-real fusion of object under mixed reality and medium
CN117765209A