Real-Time Screen Translation for Virtual Meetings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional software solutions are unable to translate presentation slides in real-time during virtual meetings and fail to detect changes in screen-sharing content, thereby not providing immediate translation of new content.

Innovation Solution

The implementation of a system that uses a machine-learning model to identify and translate visual presentation content in real-time by analyzing visual attributes, segmenting images, and recognizing text, while also monitoring for changes in the shared content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional software solutions are used for translation, then translation capability is provided, but real-time translation of presentation slides is not achieved and manual file identification is required

Engineering Contradiction:
Improvetranslation speedVSAvoidtime for manual file identification and preparation
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system automatically detects and translates presentation slides without requiring users to manually identify or prepare files. The machine learning model autonomously monitors screen content, identifies presentation regions, detects slide changes, and performs translation in real-time, eliminating the need for manual intervention in file identification and translation initiation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary actions by continuously monitoring and analyzing screen content in real-time before translation is needed. The machine learning model pre-identifies presentation regions and tracks slide changes, so that translation can immediately commence when new content appears, eliminating delays associated with manual file identification.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If manual translation preparation is required, then translation accuracy can be maintained, but the burden shifts to participants and screen-shared content is not accounted for

Engineering Contradiction:
Improvetranslation accuracyVSAvoiduser burden for file identification
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system autonomously performs content identification, region segmentation, and translation without requiring user action. Participants simply share their screens, and the system automatically detects the presentation content, identifies text regions using image segmentation, and provides translations, completely eliminating the burden of manual file identification and translation setup.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The machine learning model acts as an intermediary between the screen-shared content and the translation output. It automatically analyzes the visual content, segments the presentation region from other screen elements, identifies text, and generates translations, serving as an intelligent mediator that bridges the gap between raw screen content and translated output without user intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Extent of automation

If real-time monitoring of screen content is implemented, then automatic detection of new content is achieved, but system complexity increases

Engineering Contradiction:
Improveautomatic content detectionVSAvoidsystem complexity for real-time monitoring
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The machine learning model serves as an intermediary that simplifies the complex task of real-time screen monitoring. It automatically segments the screen to identify presentation regions, distinguishes them from other visual elements, and tracks content changes by comparing sequential frames. This intermediary layer handles the computational complexity internally while presenting a simple automated translation service to users.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system segments the screen content into distinct regions, separating the presentation area from other visual elements such as video feeds, chat windows, and toolbars. The machine learning model divides the complex screen capture into manageable segments, focusing translation resources only on the identified presentation region, thereby reducing overall system complexity while maintaining automated detection capability.

Inventive Principle:
Principle #1Segmentation

4Reliability

If translation is performed for all screen content, then completeness is improved, but irrelevant content is also translated wasting resources

Engineering Contradiction:
Improvetranslation completenessVSAvoidcomputational resources for translating irrelevant content
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The machine learning model segments the screen content to identify and isolate the presentation region from other visual elements. By dividing the screen into distinct functional areas and selectively processing only the presentation segment, the system ensures translation completeness for relevant content while avoiding unnecessary translation of video feeds, chat interfaces, and other non-presentation elements, thereby conserving computational resources.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies different processing qualities to different screen regions. The identified presentation region receives full translation processing with high reliability, while other regions are either excluded from processing or receive minimal processing. This local quality approach ensures that translation resources are concentrated on the content that requires translation, maintaining completeness for relevant content while reducing energy consumption for irrelevant content.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12309211B1Automatic image translation for virtual meetings
Publication Date: 2025.05.20 KUDO INC
  • US12309211B1 patent drawing
  • US12309211B1 patent drawing
  • US12309211B1 patent drawing

AI summary

Disclosed herein are methods and systems to translate text on images shared between participants online to the preferred language of the viewer. The system determines the material shared between participants. The system then uses optical character recognition (OCR) to determine which portion of the shared image is text. The system translates the text into the preferred language of the viewer and overlays the translated text onto the image. The system monitors the shared screen to determine a revision in the image. When a revision is detected, the system finds and translates the text to the preferred language of the viewer.