Video Conference Screen Comparison for Off-Task Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to adequately track whether remote learning students are on task during live video conferences, as there are no effective solutions to determine if participants are engaged in the same activity.

Innovation Solution

A system utilizing computer vision and content recognition technologies to analyze images from student screens, comparing them for similarities in text, graphical objects, color schemes, and shapes to determine if participants are engaged in the same task, and providing notifications or alerts when discrepancies are detected.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If computer vision technology is used to analyze screen images and determine participant engagement, then the ability to track whether students are on task is improved, but the device complexity and processing requirements increase

Engineering Contradiction:
Improveability to track participant engagementVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses an intermediary system (the engagement detection system with computer vision algorithms) that mediates between the screen images and the determination of participant engagement. This intermediary processes visual information from multiple screens, compares content similarities, and generates engagement status without requiring direct complex analysis at each participant's device, thus improving measurement precision while managing system complexity through centralized processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If real-time image analysis is performed to detect off-task behavior, then the timeliness of intervention is improved, but the processing time and computational resources increase

Engineering Contradiction:
Improveresponse time for interventionVSAvoidcomputational processing power
Core Design Contradiction:
Loss of timeVSPower

Solution Approach 1:

The system performs partial analysis by focusing on key visual features and content similarities rather than complete frame-by-frame analysis of all screens. It compares essential elements (text, graphical objects, color schemes, shapes) across screens to detect engagement status, using sufficient but not excessive processing to achieve timely detection without requiring maximum computational power continuously.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If multiple screen images are analyzed and compared for content similarity, then the accuracy of engagement detection is improved, but the quantity of data to be processed increases

Engineering Contradiction:
Improveaccuracy of engagement detectionVSAvoidvolume of image data
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system extracts and compares only the essential visual features from multiple screen images (text content, graphical objects, color schemes, and shapes) rather than processing complete high-resolution images. This extraction approach maintains high accuracy in detecting whether participants are viewing the same content while significantly reducing the volume of data that needs to be transmitted and processed.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12555379B2Computer vision to determine when video conference participant is off task
Publication Date: 2026.02.17 LENOVO (SINGAPORE) PTE LTD
  • US12555379B2 patent drawing
  • US12555379B2 patent drawing
  • US12555379B2 patent drawing

AI summary

In one aspect, a device includes a processor assembly and storage accessible to the processor assembly. The storage includes instructions executable by the processor assembly to use one or more content recognition/computer vision algorithms to determine whether first and second images from respective client devices of first and second video conference participants indicate the first and second video conference participants being engaged in a same task. Based on a determination that the first and second images do not indicate the first and second video conference participants being engaged in the same task, the instructions are executable to present an electronic notification indicating that the first and second video conference participants are not engaged in the same task. For example, the electronic notification may be presented at one of the participants' devices and/or the device of a separate conference organizer.