Recipient-Side Video Content Segmentation for Distraction Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing videoconferencing technologies provide limited control to recipients over the presentation of multimedia content, leading to distractions and discomfort due to background artifacts, network conditions, and individual sensitivities such as color blindness or attention disorders.

Innovation Solution

A content recognition model classifies video frames into primary and auxiliary content, allowing recipients to modify or replace background pixels and adjust text appearance based on personal preferences, with processing occurring on the recipient's device or server.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If videoconferencing content is displayed with full detail including all backgrounds and artifacts, then information completeness is improved, but user distraction and discomfort increase

Engineering Contradiction:
Improveinformation completenessVSAvoiduser distraction
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent segments video content into distinct regions: foreground regions (speakers, presenters) and background regions (artifacts, distractions). This segmentation allows selective processing where background regions can be modified or replaced while preserving foreground content, thus maintaining information completeness while reducing user distraction through targeted background modification.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If background modification is performed on the recipient's device, then user control and customization are improved, but processing resource consumption increases

Engineering Contradiction:
Improveuser controlVSAvoidprocessing resources
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent implements partial action by performing background modification only on specific background regions rather than processing the entire video frame. The system identifies and modifies only the background portions while leaving foreground content unchanged, thereby providing user control and customization while significantly reducing processing resource consumption compared to full-frame processing.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If content recognition processing is performed to classify video frames, then content perception is improved, but network bandwidth requirements increase

Engineering Contradiction:
Improvecontent perceptionVSAvoidnetwork bandwidth
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential classification information (foreground vs. background regions) from video frames and transmits this extracted data to the recipient's device. Instead of transmitting fully processed content recognition results, the system sends compact region classification masks that can be applied locally, thereby improving content perception while minimizing network bandwidth consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20260038168A1Recipient-side modification of videoconferencing content
Publication Date: 2026.02.05 NVIDIA CORP
  • US20260038168A1 patent drawing
  • US20260038168A1 patent drawing
  • US20260038168A1 patent drawing

AI summary

Disclosed are apparatuses, systems, and techniques for implementing recipient's control over displayed content received in videoconferencing applications. In one embodiment, the techniques include receiving, by a first processing device, media frames depicting participant(s) of a videoconference and generated by a sending processing device communicatively coupled to the first processing device over a network. The techniques further include processing, using a content recognition model, the media frames to identify auxiliary content in the media frames and replacing at least a portion of the identified auxiliary content with a replacement auxiliary content to generate a plurality of modified media frames. The techniques further include causing the modified plurality of media frames to be displayed using the first processing device or a second processing device.