Video Texture Rendering in Shared AR Video Calls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video calling systems are limited to non-interactive video calls, only allowing user devices to present and view captured videos between each other without the capability to render participants as augmented reality (AR) effects within a shared AR scene.
Innovation Solution
The system enables shared augmented reality scenes during video calls by transmitting video data and video processing data through streaming channels, allowing client devices to render participants as video textures within AR effects in a shared AR scene.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional video calling systems are used, then video calls can be established between user devices, but the system only allows presentation and viewing of captured videos without AR effects capability
Solution Approach 1:
The system segments video processing by having the first client device perform video capture, face detection, and AR effect rendering, while the second client device performs viewing. This division allows AR functionality to be added without requiring both devices to have equal processing capabilities, resolving the contradiction between enhanced adaptability and device complexity.
Solution Approach 2:
The patent introduces an intermediary AR effect processing mechanism where the first client device acts as a mediator that receives video input, processes it through face detection and AR rendering, then transmits the processed video to the second device. This intermediary processing layer enables AR effects capability while managing system complexity through specialized processing.
2Adaptability or versatility
If video data and video processing data are transmitted through streaming channels, then shared AR scenes can be rendered, but data transmission requirements increase
Solution Approach 1:
The system performs preliminary video processing at the source device by detecting faces and rendering AR effects before transmission. This preliminary action ensures that the video data is pre-processed into a format suitable for AR display, reducing the need for complex real-time processing at the receiving end and optimizing data transmission requirements.
Solution Approach 2:
The first client device serves itself by performing video capture, face detection, and AR effect rendering locally before transmitting the processed video. This self-service approach allows the device to prepare its own video data with embedded AR effects, reducing the processing burden on the network and receiving devices while enabling shared AR scene rendering.
3Speed
If face detection and video processing are performed at the client device, then AR effects can be rendered in real-time, but processing requirements increase
Solution Approach 1:
The system applies local quality processing by performing face detection and AR effect rendering only on relevant portions of the video feed (specifically on detected face regions) rather than processing the entire video stream. This localized processing approach enables real-time AR rendering while significantly reducing the overall processing energy consumption compared to full-frame processing.
Data Source
AI summary
Systems, methods, client devices, and non-transitory computer-readable media are disclosed for utilizing video data and video processing data to enable shared augmented reality scenes having video textures depicting participants of video calls as augmented reality (AR) effects during the video calls. For instance, the disclosed systems can establish a video call between client devices that include streaming channels (e.g., a video and audio data channel). In one or more implementations, the disclosed systems enable the client devices to transmit video processing data and video data of a participant through the streaming channel during a video call. Indeed, in one or more embodiments, the disclosed systems cause the client devices to utilize video data streams and video processing data to render videos as video textures within AR effects in a shared AR scene (or AR space) of the video call (e.g., to depict participants within the AR scene).


