AR-based live broadcast real-time interaction system and method

Through the collaborative acquisition of motion capture devices and environment perception devices and dual-host parallel processing, combined with physical engines and timestamp synchronization, the problem of insufficient real-time interaction capabilities between virtual elements and anchor actions and limited synchronization processing capabilities of multi-dimensional environmental information is solved, and a high consistency immersion experience between virtual elements and real scenes is achieved.

CN120434412APending Publication Date: 2025-08-05王建国
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510611758.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing virtual elements have insufficient real-time interaction capabilities with the anchor, and the synchronization processing capabilities of multi-dimensional environmental information are limited, which makes it difficult to match the lighting and shadows of real characters naturally, weakening the immersive experience of live broadcasts and film and television content.

Method used

Motion capture equipment and environment perception equipment are used to collect anchor action and scene data in collaboration, combined with dual host parallel processing and virtual and real fusion module, eliminate delays through time stamp synchronization, simulate the motion trajectory of virtual elements using the physics engine, and realize optical superposition of virtual elements through the AR display terminal, and combine image morphology processing to eliminate aliasing edges to achieve high-precision virtual and real fusion.

Benefits of technology

It realizes low latency and natural interaction between virtual elements and anchor actions, improves the light and shadow consistency between virtual elements and real scenes, and enhances the immersive experience of live broadcast scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120434412A_ABST
    Figure CN120434412A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an AR-based live broadcast real-time interaction system and method, and the system comprises a motion capture device, an environment perception device, a dual-host processing module, a virtual-real fusion module, an interaction response module, a data distribution module and an AR display terminal. And the double hosts parse the action data and the environment data in parallel and simulate a physical feedback track of a virtual element in combination with a physical engine, so that low-delay response of an interaction instruction is realized. And the virtual-real fusion module adopts a timestamp synchronization mechanism to align the rendered picture and the anchor picture after green screen matting, and eliminates edge sawteeth through image morphological processing to generate high-precision fusion streaming media data. And the AR display terminal projects the virtual elements to real scene space coordinates by adopting an optical superposition technology, so that an anchor previews a virtual-real fusion effect in real time. According to the invention, the dynamic interaction capability and shadow consistency of the virtual elements and the real scene are improved, and the immersive experience of the live scene is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an AR-based live broadcast real-time interaction system and method. Background Art

[0002] With the rapid development of augmented reality (AR) technology, the integration of virtual elements and real-world scenes is widely adopted in live broadcasts and film and television production to enhance the expressiveness of content. Existing technologies, through methods such as green screen keying and real-time rendering, have enabled the seamless integration of live streamers and dynamic virtual backgrounds. For example, these can place live streamers in scenes such as virtual forests and futuristic cities, significantly enhancing the sense of visual immersion.

[0003] However, this technology still has significant limitations: The real-time interaction between virtual elements and the host is insufficient. For example, when the host needs to interact with objects in the virtual scene (such as picking up virtual props or triggering special effects), existing solutions often rely on pre-set animations or fixed program responses. They are unable to dynamically adjust the state of the virtual elements (such as position changes and physical feedback) based on the host's actual actions (such as gestures and movement). This results in a stilted and unrealistic interaction process, which is particularly important in scenarios such as game live streaming and virtual performances, severely restricting the depth and flexibility of the user experience.

[0004] Furthermore, existing virtual-reality fusion technologies have limited ability to simultaneously process multi-dimensional environmental information (such as spatial depth and light and shadow variations). This makes it difficult for virtual elements to naturally match the lighting and shadows of real people, further weakening the realism of the scene. Therefore, a technical solution that can perceive the host's movements in real time and dynamically interact with virtual elements is urgently needed to overcome the static limitations of virtual-reality fusion and enhance the immersive experience of live broadcasts and film and television content. Summary of the Invention

[0005] In response to the shortcomings of existing technologies, the present invention provides an AR-based live broadcast real-time interactive system and method. The present invention solves the problem of lack of immersion caused by the insufficient real-time interaction ability between virtual elements and host actions and the limited synchronous processing ability of multi-dimensional environmental information in existing virtual-reality fusion technologies.

[0006] In order to solve the above technical problems, the specific technical solutions of the present invention are as follows: In a first aspect, the present invention provides an AR-based live broadcast real-time interactive system, comprising: A motion capture device collects the host's body movement data in real time and transmits the movement data to the first host; An environmental sensing device synchronously acquires depth information and ambient light data of a scene, and transmits the depth information and light data to the first host; The first host parses the received motion data to generate interaction instructions for the virtual element, parses the depth information and illumination data to generate light and shadow parameters for the virtual element, and outputs the interaction instructions and light and shadow parameters to the second host; The second host deploys a virtual scene rendering engine, receives the interactive instructions and light and shadow parameters, and then renders the virtual elements. The rendered virtual elements and the anchor image after green screen matting are input into the virtual-reality fusion module. A virtual-reality fusion module, which superimposes the rendered virtual elements onto the anchor screen after the green screen is matted through a transparent playback window to generate virtual-reality fusion streaming media data; an interactive response module, which calls a physics engine to simulate the motion trajectory of the virtual element according to the interactive instruction, and feeds back the trajectory data to the second host to update the position and state of the virtual element; A data distribution module receives the virtual-reality fusion streaming media data, compresses it and pushes it to the live broadcast platform; The AR display terminal receives the virtual-reality fusion streaming media data and displays it in real time, allowing the host to preview the superposition effect of virtual elements and real scenes.

[0007] Furthermore, in the AR-based live broadcast real-time interactive system described in the present invention, the motion capture device includes an infrared camera and an inertial sensor, which generates grabbing or releasing instructions of virtual elements through joint coordinate mapping.

[0008] Furthermore, in the AR-based live broadcast and real-time interactive system described in the present invention, the environmental perception device is an RGB-D camera, which extracts the scene light intensity through depth map analysis and dynamically matches the shadow direction of the virtual element.

[0009] Furthermore, in the AR-based live broadcast real-time interactive system described in the present invention, the virtual-reality fusion module is based on the Unreal Engine or the Unity rendering engine, and eliminates the delay between the action data and the rendered image through a timestamp synchronization mechanism.

[0010] Furthermore, in the AR-based live broadcast real-time interactive system described in the present invention, the interactive response module calls the physical engine to simulate the motion trajectory of the virtual element, and the physical feedback includes the adsorption or dropping effect of the virtual props.

[0011] Furthermore, in the AR-based live broadcast real-time interactive system described in the present invention, the first host and the second host transmit interactive instructions and light and shadow parameters via Gigabit Ethernet, and the motion data is processed in parallel via a hardware acceleration card.

[0012] Furthermore, in the AR-based live broadcast real-time interactive system described in the present invention, the transparent playback window uses image morphological processing to smooth the edges of the anchor's outline to eliminate the jagged edges or color differences left over from the green screen matte.

[0013] Furthermore, in the AR-based live broadcast real-time interactive system described in the present invention, the AR display terminal is a semi-transparent head-mounted display device, which projects virtual elements to the real scene position in the host's field of view through optical overlay technology.

[0014] Furthermore, in the AR-based live broadcast real-time interactive system described in the present invention, the data distribution module performs H.265 encoding compression on the fused streaming media data and distributes it to multiple terminal users through the live broadcast streaming server.

[0015] In a second aspect, the AR-based live broadcast real-time interaction method provided by the present invention is applied to the AR-based live broadcast real-time interaction system, comprising: Step 1: Collect the host's body movement data in real time and transmit the movement data to the first host; Step 2: synchronously obtain depth information and ambient light data of the scene, and transmit the depth information and light data to the first host; Step 3: parse the received motion data to generate interaction instructions for the virtual element, parse the depth information and illumination data to generate light and shadow parameters for the virtual element, and output the interaction instructions and light and shadow parameters to the second host; Step 4: deploy a virtual scene rendering engine, render virtual elements after receiving the interactive instructions and light and shadow parameters, and input the rendered virtual elements and the anchor image after green screen matting into the virtual-reality fusion module; Step 5: Overlaying the rendered virtual elements onto the live broadcast screen after green screen matting through a transparent playback window to generate virtual-reality fusion streaming media data; Step 6: Calling a physics engine to simulate the motion trajectory of the virtual element according to the interaction instruction, and feeding back the trajectory data to the second host to update the position and state of the virtual element; Step 7: Receive the virtual-reality fusion streaming media data, compress it and push it to the live broadcast platform; Step 8: Receive the virtual-reality fusion streaming media data and display it in real time, so that the host can preview the superposition effect of virtual elements and real scenes.

[0016] Beneficial effects of the present invention: The present invention realizes real-time dynamic interaction between virtual elements and anchor actions through the collaborative data collection of motion capture equipment and environmental perception equipment, combined with a dual-host parallel processing mechanism; based on the skeleton binding algorithm, the joint coordinates are mapped into the displacement and rotation instructions of the virtual elements, and the motion trajectory of the virtual elements is physically simulated in combination with the physics engine, which significantly improves the real-time response of the virtual elements to action instructions and the realism of the motion trajectory; the environmental analysis module constructs a three-dimensional scene grid through depth map analysis and dynamically matches the light and shadow parameters of the virtual elements, solving the problem of inconsistency between the virtual shadow and the direction of the real environment light source; the virtual-reality fusion module adopts a timestamp synchronization mechanism and image morphological processing to effectively eliminate the timing deviation between the action data and the rendered image and the jagged edges remaining in the green screen, thereby enhancing the fusion accuracy of the virtual and real images; the data distribution module reduces the transmission delay of streaming media data through H.265 encoding compression and multi-point transmission of the content distribution network. At the same time, the optical overlay technology of the AR display terminal realizes the spatial position matching of the virtual elements in the real scene, and finally achieves a low-latency, high-consistency virtual-reality interaction effect in the live broadcast scene, enhancing the user's immersive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on the drawings without paying any creative labor.

[0018] Figure 1 This is a system architecture diagram of an AR-based live broadcast and real-time interactive system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. The technical solutions provided by each embodiment of the present invention are described in detail below in conjunction with the drawings. In order to better understand the purpose of the present invention, the present invention is further described in detail below.

[0020] First, see Figure 1 The present invention provides an AR-based live broadcast real-time interactive system, comprising: The present invention provides an AR-based live broadcast real-time interactive system. It uses motion capture equipment and environmental perception equipment to collect live broadcaster movements and scene data in real time. Through dual-host processing and virtual-real integration, it enables dynamic interaction between virtual elements and real scenes. The following describes the technical solutions and logical relationships of each module of the system: The motion capture device uses a combination of an infrared camera and an inertial sensor. The infrared camera captures the coordinates of the host's key joints, while the inertial sensor collects data on limb acceleration and angular velocity. These two data are then fused to generate a high-precision motion trajectory. After the motion data is transmitted to the primary host, a pre-set skeletal rigging algorithm maps the joint coordinates into interactive commands for virtual elements. For example, approaching a virtual prop triggers a "grab" command, while moving away triggers a "release" command.

[0021] The environmental perception device is an RGB-D camera, which uses infrared structured light or time-of-flight (ToF) technology to obtain scene depth information and analyzes ambient light intensity based on the RGB image. The depth data is used to construct the three-dimensional scene spatial coordinates, and the lighting data is processed through grayscale histogram equalization to generate the lighting parameters of virtual elements. This ensures that the direction of virtual shadows aligns with the direction of the real-world light source. For example, when there is strong light on the left side of the real scene, the projection of the virtual element automatically shifts to the right.

[0022] The first host, equipped with a GPU accelerator, performs parallel parsing of motion and environmental data. The motion parsing module converts the raw motion data into displacement and rotation parameters for virtual elements using a skeletal animation model. The environmental parsing module generates a 3D scene mesh based on the depth map and adjusts the reflectivity and transparency of the virtual elements' materials based on lighting intensity. The parsed interaction commands and lighting parameters are transmitted to the second host via Gigabit Ethernet.

[0023] The second host, running the Unreal Engine or Unity rendering engine, receives interactive commands and calls the virtual element library to load the corresponding 3D model. For example, a grab command triggers the display of a virtual prop, while a release command hides it. The rendering engine dynamically adjusts the virtual element's shader parameters based on lighting parameters, such as adjusting highlight intensity to match ambient light. The rendered virtual elements and the live stream's green screen keyed image are simultaneously input into the virtual-reality fusion module. A timestamp synchronization mechanism aligns the two, preventing delays in movement and image quality.

[0024] The virtual-reality fusion module overlays virtual elements onto the live stream using a transparent playback window. This transparent playback window utilizes alpha channel blending technology, creating a semi-transparent virtual element overlay onto the live stream after green screen keying. To avoid jagged edges left behind by green screen keying, the fusion module applies image morphological processing to the live stream's outline, using, for example, dilation to fill edge gaps and erosion to remove stray green edges, resulting in a smooth edge for the overlaid image.

[0025] The interactive response module integrates a physics engine to simulate the physical behavior of virtual elements based on interactive commands. For example, when a grab command is triggered, the physics engine calculates the virtual prop's adsorption trajectory, allowing it to smoothly move to the host's hand position. When a release command is triggered, the prop's drop trajectory is simulated based on gravity parameters. The trajectory data output by the physics engine is fed back to the secondary host in real time, updating the virtual element's position and collision status, forming a closed-loop interaction.

[0026] The data distribution module compresses the integrated virtual and real streaming media data using H.265 encoding and distributes it to end users via the live streaming server. Inter-frame prediction technology is used during the encoding process to reduce data redundancy, and the video stream is transmitted to multiple locations via a content delivery network (CDN) to reduce live streaming latency. The AR display terminal is a semi-transparent head-mounted display (HMD) that uses waveguide optical technology to project virtual elements onto a specific location in the host's field of view, such as overlaying virtual props on a real desktop. The host observes the virtual and real fusion effect in real time through the HMD and adjusts interactive actions.

[0027] In the second aspect, the AR-based live broadcast real-time interaction method provided by the present invention, when applied to the system, the technical solutions and logical relationships of each step are supplemented as follows: In step 1, the motion capture device uses an infrared camera to capture the spatial coordinate data of the host's joint key points. Combined with the limb acceleration and angular velocity data collected by inertial sensors, it generates motion trajectory data containing skeletal node displacements and rotation angles. This motion trajectory data is then transmitted to the first host via USB 3.0 or Gigabit Ethernet for subsequent analysis and processing.

[0028] In step 2, the RGB-D camera generates a scene depth map using structured light or time-of-flight (ToF) technology, while also analyzing the ambient light intensity based on the RGB image. The depth map and light intensity data are transmitted to the first host simultaneously with the motion data via the same transmission channel. The two are aligned using hardware-level timestamps to prevent timing deviations that could cause spatial misalignment between virtual elements and the real scene.

[0029] In step 3, the first host applies a skeletal binding algorithm to the motion trajectory data, mapping joint coordinates into displacement and rotation instructions for virtual elements. For example, changes in hand joint coordinates trigger a grabbing motion for a virtual prop. The environmental data analysis module constructs a 3D scene mesh model based on the depth map and extracts ambient lighting parameters through grayscale histogram equalization. It dynamically adjusts the diffuse reflectance and shadow direction of virtual elements to match the real-world lighting.

[0030] In step 4, the second host's virtual scene rendering engine loads a pre-configured library of virtual element resources (such as 3D models and particle effects) and instantiates virtual elements based on the interaction instructions. The rendering engine adjusts the virtual element's shader parameters based on lighting parameters, such as adjusting the brightness of the specular map based on ambient light intensity. It also uses a timestamp synchronization mechanism to align the rendered image with the live stream's green screen keyed image, eliminating delays between the action and the image.

[0031] In step 5, the virtual-reality fusion module uses alpha channel blending within a transparent playback window to superimpose the rendered virtual elements onto the live stream, which has been keyed off the green screen. To eliminate jagged edges left by the green screen keying, the fusion module dilates and erodes the live stream's outline, filling in gaps and removing stray green edges, resulting in seamlessly blended streaming data.

[0032] In step 6, the interactive response module invokes a physics engine (such as PhysX) to simulate the motion trajectory of the virtual element. For example, when a grab command is triggered, the physics engine calculates the virtual prop's adsorption trajectory, allowing it to smoothly move to the host's hand position. When a release command is triggered, the prop's free-fall trajectory is simulated based on gravity parameters. This trajectory data is fed back to the secondary host in real time, updating the virtual element's spatial coordinates and collision status in the rendering engine.

[0033] In step 7, the data distribution module uses the H.265 encoding standard to perform inter-frame prediction compression on the virtual-reality fusion streaming media data, reducing transmission bandwidth usage. The compressed data stream is distributed to the content delivery network (CDN) via the live streaming server. CDN nodes transmit the video stream to end users at multiple points, reducing live broadcast latency.

[0034] In step 8, the AR display terminal uses waveguide optical technology to project a fused virtual and real image onto the display area of the semi-transparent headset. Virtual elements automatically match the spatial position of the real scene based on the scene depth information. For example, a virtual prop is projected to a specific coordinate on the real desktop. The host observes the occlusion relationship between the prop and the desktop through the headset and adjusts interactive actions in real time to enhance the sense of immersion.

[0035] The present invention addresses the challenges of insufficient interaction between virtual elements and real scenes and limited synchronous processing of multi-dimensional environmental information in the prior art. The invention utilizes the following technical solutions: A motion capture device utilizes a combination of an infrared camera and an inertial sensor. The infrared camera captures the three-dimensional coordinates of the anchor's joints at a rate of 60 frames per second, while the inertial sensor collects acceleration and angular velocity data from limb movements. The two are then fused using a Kalman filter algorithm to generate motion trajectory data containing displacement and rotation angles, which is then transmitted to a first host. The first host then maps the joint coordinates into interaction commands for the virtual elements using a skeletal rigging algorithm. For example, changes in hand joint coordinates trigger a grabbing action for a virtual prop, while a release command is generated when the displacement exceeds a preset threshold. The environmental perception device utilizes an RGB-D camera, using structured light technology to acquire a scene depth map with a resolution of 1280×720. The ambient light intensity is analyzed based on the RGB image, and illumination parameters are extracted using a grayscale histogram equalization algorithm. The virtual element's diffuse reflectance and shadow direction are dynamically adjusted to match the light source direction in the real scene. For example, when a strong light source is located to the left of the real environment, the virtual element's projection direction automatically shifts to the right. Interaction commands and lighting parameters are transmitted between the first and second hosts via Gigabit Ethernet. The second host uses Unreal Engine to render virtual elements. The rendering engine dynamically adjusts the brightness and transparency of the virtual element's specular map based on the lighting parameters, with a rendering frame rate of 60 fps to ensure real-time performance. The virtual-reality fusion module uses alpha channel blending in a transparent playback window to overlay the rendered virtual elements onto the live stream after green screen shading. Simultaneously, image morphology processing is used to dilate and erode the live stream's outline, using a dilation kernel size of 3×3 pixels and an erosion kernel size of 5×5 pixels to eliminate jagged edges and color aberration left by the green screen shading. The interactive response module uses the PhysX physics engine to simulate the motion trajectory of the virtual element. When a grab command is triggered, the physics engine calculates the virtual prop's attachment trajectory based on a spring-damper model. When a release command is triggered, the free-fall trajectory is simulated based on gravitational acceleration parameters. This trajectory data is fed back to the second host at a rate of 120 updates per second, enabling real-time adjustments to the virtual element's spatial coordinates. The data distribution module applies inter-frame prediction compression to the fused streaming media data using the H.265 encoding standard, setting the encoding bitrate to 8Mbps. This data is then distributed to the content distribution network via the live streaming server, with CDN nodes enabling multi-point transmission. The AR display terminal utilizes waveguide optical technology to project a fused virtual and real image onto the display area of a semi-transparent headset. Virtual elements are aligned with the spatial coordinates of the real scene based on scene depth information. For example, a virtual prop can be projected onto a specific 3D coordinate point on a real desktop with a projection accuracy of less than 0.1 mm, allowing the host to observe the occlusion relationship between virtual elements and real objects in real time through the headset.The above technical solution solves the problem of real-time interaction delay between virtual elements and host actions by processing motion data and environmental data in parallel on two hosts, combined with dynamic simulation of the physics engine and timestamp synchronization mechanism. At the same time, through depth map analysis and dynamic matching of light and shadow parameters, consistency between virtual shadows and real environment light sources is achieved.

[0036] The present invention solves the problem of lack of immersion caused by the insufficient real-time interaction between virtual elements and anchor actions and the limited ability to synchronously process multi-dimensional environmental information in existing virtual-reality fusion technologies through the following technical solutions: Motion capture equipment and environmental perception equipment collaborate to collect the host's motion data and multi-dimensional scene information. The motion capture equipment generates high-precision skeletal joint coordinate data through the fusion of infrared cameras and inertial sensors, mapping it to displacement and rotation instructions for virtual elements. The environmental perception equipment uses RGB-D cameras to obtain scene depth information and analyze ambient light intensity, dynamically adjusting the lighting parameters of virtual elements to match the real-world scene light source. A dual-host parallel processing mechanism analyzes motion data and scene data in separate channels, and uses a timestamp synchronization mechanism to align the timing of the rendered image with the host's movements, eliminating interaction delays and achieving low-latency response of virtual elements to body movements.

[0037] The virtual-reality fusion module combines a physics engine with image morphological processing technology to enhance the realism of interactions between virtual elements and real scenes, as well as the accuracy of image fusion. The physics engine simulates the motion trajectory of virtual elements based on interactive commands, such as calculating adsorption trajectories using a spring-damper model or simulating drop trajectories using gravity parameters, ensuring that the physical feedback of virtual elements conforms to the laws of real mechanics. The virtual-reality fusion module uses alpha channel blending technology within a transparent playback window to overlay virtual elements with the live stream after green screen cutout. Dilation and erosion operations correct for jagged edges and chromatic aberration in the live stream's outline, generating seamlessly integrated streaming media data. This design ensures consistent lighting, shadow, and spatial occlusion between virtual elements and real scenes, enhancing the naturalness of virtual-reality interactions.

[0038] The data distribution module and AR display terminal optimize the transmission and presentation of virtual-reality fusion data. The data distribution module uses H.265 encoding to compress streaming media data, and combined with a content distribution network (CDN), it achieves multi-point low-latency transmission, reducing the risk of screen freezes on the user side. The AR display terminal uses waveguide optical technology to project virtual elements into the spatial coordinates of the real scene, and dynamically matches virtual and real positions based on depth information, allowing the host to preview the occlusion relationship between virtual elements and real objects in real time through a translucent headset. Through efficient encoding, spatial coordinate mapping, and optical superposition, these technologies address the bandwidth and accuracy issues associated with the synchronous processing of multi-dimensional environmental information, ultimately achieving a highly consistent immersive experience of virtual-reality fusion images.

Claims

1. A live broadcast real-time interactive system based on AR, characterized by: include: A motion capture device collects the host's body movement data in real time and transmits the movement data to the first host; An environmental sensing device synchronously acquires depth information and ambient light data of a scene, and transmits the depth information and light data to the first host; The first host parses the received motion data to generate interaction instructions for the virtual element, parses the depth information and illumination data to generate light and shadow parameters for the virtual element, and outputs the interaction instructions and light and shadow parameters to the second host; The second host deploys a virtual scene rendering engine, receives the interactive instructions and light and shadow parameters, and then renders the virtual elements. The rendered virtual elements and the anchor image after green screen matting are input into the virtual-reality fusion module. A virtual-reality fusion module, which superimposes the rendered virtual elements onto the anchor screen after the green screen is matted through a transparent playback window to generate virtual-reality fusion streaming media data; an interactive response module, which calls a physics engine to simulate the motion trajectory of the virtual element according to the interactive instruction, and feeds back the trajectory data to the second host to update the position and state of the virtual element; A data distribution module receives the virtual-reality fusion streaming media data, compresses it and pushes it to the live broadcast platform; The AR display terminal receives the virtual-reality fusion streaming media data and displays it in real time, allowing the host to preview the superposition effect of virtual elements and real scenes.

2. The AR-based live broadcast real-time interactive system according to claim 1, characterized in that: The motion capture device includes an infrared camera and an inertial sensor, and generates grabbing or releasing instructions of virtual elements through joint coordinate mapping.

3. The AR-based live broadcast real-time interactive system according to claim 1, characterized in that: The environment perception device is an RGB-D camera, which extracts the scene lighting intensity through depth map analysis and dynamically matches the shadow direction of the virtual element.

4. The AR-based live broadcast and real-time interactive system according to claim 1, characterized in that: The virtual-reality fusion module is based on the Unreal Engine or Unity rendering engine, and eliminates the delay between action data and rendered images through a timestamp synchronization mechanism.

5. The AR-based live broadcast and real-time interactive system according to claim 1, characterized in that: The interactive response module calls the physical engine to simulate the motion trajectory of the virtual element, and the physical feedback includes the adsorption or dropping effect of the virtual prop.

6. The AR-based live broadcast and real-time interactive system according to claim 1, characterized in that: The first host and the second host transmit interactive instructions and light and shadow parameters via Gigabit Ethernet, and perform parallel processing on motion data via a hardware acceleration card.

7. The AR-based live broadcast and real-time interactive system according to claim 1, characterized in that: The transparent playback window uses image morphological processing to smooth the edges of the anchor's outline and eliminate the jagged edges or color differences left by the green screen keying.

8. The AR-based live broadcast and real-time interactive system according to claim 1, characterized in that: The AR display terminal is a translucent head-mounted display device that projects virtual elements onto the real scene position in the host's field of view through optical overlay technology.

9. The AR-based live broadcast and real-time interactive system according to claim 1, characterized in that: The data distribution module performs H.265 encoding compression on the merged streaming media data and distributes it to multiple terminal users through the live streaming server.

10. A live broadcast and real-time interaction method based on AR, applied to the live broadcast and real-time interaction system based on AR according to any one of claims 1 to 9, characterized in that: include: Step 1: Collect the host's body movement data in real time and transmit the movement data to the first host; Step 2: synchronously obtain depth information and ambient light data of the scene, and transmit the depth information and light data to the first host; Step 3: parse the received motion data to generate interaction instructions for the virtual element, parse the depth information and illumination data to generate light and shadow parameters for the virtual element, and output the interaction instructions and light and shadow parameters to the second host; Step 4: deploy a virtual scene rendering engine, render virtual elements after receiving the interactive instructions and light and shadow parameters, and input the rendered virtual elements and the anchor image after green screen matting into the virtual-reality fusion module; Step 5: Overlaying the rendered virtual elements onto the live broadcast screen after green screen matting through a transparent playback window to generate virtual-reality fusion streaming media data; Step 6: Calling a physics engine to simulate the motion trajectory of the virtual element according to the interaction instruction, and feeding back the trajectory data to the second host to update the position and state of the virtual element; Step 7: Receive the virtual-reality fusion streaming media data, compress it and push it to the live broadcast platform; Step 8: Receive the virtual-reality fusion streaming media data and display it in real time, so that the host can preview the superposition effect of virtual elements and real scenes.

Citation Information

Cited By

  • Dynamic target synchronous tracking and interaction system and method

    CN121033947A

  • A dynamic target synchronous tracking and interaction system and method

    CN121033947B