Dynamic Video Conference Layout via Down-sampled Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video conferencing software is limited in configuring the layout of output images, as it can only assign a single region of interest to a single image, even when capturing panoramic images with multiple people, restricting flexibility in displaying relevant participants.

Innovation Solution

An image processing system that includes multiple image capture devices and a computing device, which generates and transmits down-sampled images for object detection, allowing dynamic cropping and mapping of regions of interest based on user instructions, audio synchronization, and object detection results to flexibly configure the output image layout.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional video conferencing software assigns a single region of interest to a single image, then the system maintains simple processing logic, but the layout flexibility of output images is limited

Engineering Contradiction:
Improvelayout flexibilityVSAvoidprocessing logic complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the image processing by separating region of interest detection (performed on down-sampled images) from the final image composition. Multiple regions of interest can be detected and assigned to different display areas, enabling flexible layouts while keeping each processing stage relatively simple.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of processing by generating down-sampled images from original images. This additional processing layer enables efficient object detection and region of interest identification without directly complicating the main image composition workflow.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If down-sampled images are generated for object detection, then the computing device can efficiently process images to identify regions of interest, but additional processing steps and data transmission are required

Engineering Contradiction:
Improveobject detection efficiencyVSAvoidprocessing steps
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by generating down-sampled images before object detection. This preprocessing step reduces image complexity and enables faster, more efficient object detection and region of interest identification, as the computing device processes smaller, simplified image data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a simplified copy (down-sampled image) of the original image for the specific purpose of object detection. This copy contains the essential structural information needed for region identification while requiring significantly less processing power and memory than the full-resolution original image.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240323042A1Image processing system and image processing method for video conferencing software
Publication Date: 2024.09.26 CUPOLA360 INC
  • US20240323042A1 patent drawing
  • US20240323042A1 patent drawing
  • US20240323042A1 patent drawing

AI summary

An image processing system and an image processing method for a video conferencing software are provided. The image processing method includes: capturing a first original image by a first image capture device and capturing a second original image by a second image capture device; generating first information corresponding to the first original image and transmitting the first information to the first image capture device; cropping a first cropped image from the first original image according to a first mapping relationship in the first information by the first image capture device; and outputting an output image including the first cropped image and a second cropped image corresponding to the second original image to the video conferencing software according to a second mapping relationship in the first information by the first image capture device.