Video Texture Rendering in Shared AR Video Calls

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video calling systems are limited to non-interactive video calls, only allowing user devices to present and view captured videos between each other without the capability to render participants as augmented reality (AR) effects within a shared AR scene.

Innovation Solution

The system enables shared augmented reality scenes during video calls by transmitting video data and video processing data through streaming channels, allowing client devices to render participants as video textures within AR effects in a shared AR scene.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional video calling systems are used, then video calls can be established between user devices, but the system only allows presentation and viewing of captured videos without AR effects capability

Engineering Contradiction:
ImproveAR effects capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments video processing by having the first client device perform video capture, face detection, and AR effect rendering, while the second client device performs viewing. This division allows AR functionality to be added without requiring both devices to have equal processing capabilities, resolving the contradiction between enhanced adaptability and device complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary AR effect processing mechanism where the first client device acts as a mediator that receives video input, processes it through face detection and AR rendering, then transmits the processed video to the second device. This intermediary processing layer enables AR effects capability while managing system complexity through specialized processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If video data and video processing data are transmitted through streaming channels, then shared AR scenes can be rendered, but data transmission requirements increase

Engineering Contradiction:
Improveshared AR scene renderingVSAvoiddata transmission volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system performs preliminary video processing at the source device by detecting faces and rendering AR effects before transmission. This preliminary action ensures that the video data is pre-processed into a format suitable for AR display, reducing the need for complex real-time processing at the receiving end and optimizing data transmission requirements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The first client device serves itself by performing video capture, face detection, and AR effect rendering locally before transmitting the processed video. This self-service approach allows the device to prepare its own video data with embedded AR effects, reducing the processing burden on the network and receiving devices while enabling shared AR scene rendering.

Inventive Principle:
Principle #25Self-service

3Speed

If face detection and video processing are performed at the client device, then AR effects can be rendered in real-time, but processing requirements increase

Engineering Contradiction:
Improvereal-time AR renderingVSAvoidprocessing energy consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system applies local quality processing by performing face detection and AR effect rendering only on relevant portions of the video feed (specifically on detected face regions) rather than processing the entire video stream. This localized processing approach enables real-time AR rendering while significantly reducing the overall processing energy consumption compared to full-frame processing.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12211121B2Generating shared augmented reality scenes utilizing video textures from video streams of video call participants
Publication Date: 2025.01.28 META PLATFORMS INC
  • US12211121B2 patent drawing
  • US12211121B2 patent drawing
  • US12211121B2 patent drawing

AI summary

Systems, methods, client devices, and non-transitory computer-readable media are disclosed for utilizing video data and video processing data to enable shared augmented reality scenes having video textures depicting participants of video calls as augmented reality (AR) effects during the video calls. For instance, the disclosed systems can establish a video call between client devices that include streaming channels (e.g., a video and audio data channel). In one or more implementations, the disclosed systems enable the client devices to transmit video processing data and video data of a participant through the streaming channel during a video call. Indeed, in one or more embodiments, the disclosed systems cause the client devices to utilize video data streams and video processing data to render videos as video textures within AR effects in a shared AR scene (or AR space) of the video call (e.g., to depict participants within the AR scene).