Real-Time Video Stream Object Modeling via Image Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current telecommunications systems face challenges in capturing, measuring, and modeling aspects within video streams during communication sessions, as existing methods do not effectively modify or model video content in real-time.

Innovation Solution

A system and method for automated image segmentation of video streams, which identifies and tracks objects of interest, such as facial features, and generates a model of these objects to modify the video stream in real-time, allowing for the transmission of stylized video content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If automated image segmentation is implemented to generate object models in real-time, then video communication quality and object representation are improved, but system complexity and computational requirements increase

Engineering Contradiction:
Improvevideo communication qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The video stream is divided into individual frames, and each frame is segmented into multiple segments based on object boundaries. This segmentation approach allows the system to process and model specific objects of interest independently, improving video communication quality while managing computational complexity through localized processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-identifying objects of interest and generating their models before final video transmission. This advance processing enables real-time modification of video streams with improved object representation, while the computational burden is distributed over time through frame-by-frame processing.

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If real-time video stream modification is performed to generate stylized content, then user interaction and content delivery are enhanced, but processing time and computational resources increase

Engineering Contradiction:
Improveuser interactionVSAvoidprocessing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system applies periodic action by processing video streams in discrete frame intervals rather than continuously. Each frame is segmented, objects are identified and modeled, and modifications are applied in periodic cycles. This approach enhances user interaction through stylized content delivery while managing processing time through efficient batch processing of video frames.

Inventive Principle:
Principle #19Periodic action

3Measurement precision

If object tracking and modeling is implemented to capture facial features, then measurement precision and object representation are improved, but computational complexity and processing requirements increase

Engineering Contradiction:
Improveobject detection precisionVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system applies local quality by focusing computational resources on specific regions of interest within each video frame, such as facial features. Instead of processing the entire frame uniformly, the segmentation and modeling operations are concentrated on localized areas containing objects of interest, thereby improving measurement precision while reducing overall processing complexity.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10896346B1Image segmentation for object modeling
Publication Date: 2021.01.19 SNAP INC
  • US10896346B1 patent drawing
  • US10896346B1 patent drawing
  • US10896346B1 patent drawing

AI summary

Systems, devices, and methods are presented for segmenting an image of a video stream with a client device by accessing a set of images within a video stream, identifying an object of interest within one or more images of the set of images, and detecting a region of interest within the one or more images. The systems, devices, and method identify a first set of median pixels in a first portion of the object of interest and a second set of median pixels in a second portion of the object of interest. The systems, devices, and methods determine a polyline approximating the first and second sets of median pixels and generate a model for the polyline.