Video Interaction Using Spatial Actions for Intent Understanding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI systems face limitations in understanding user intentions during human-computer interaction in scenarios like video calls and screen sharing.

Innovation Solution

A method and apparatus for video interaction using a large model that determines a target object through spatially directional actions, processes data based on input information, and performs data processing using a large model to obtain a result.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing AI systems are used for video interaction, then the system structure is relatively simple, but the understanding accuracy of user intentions is limited

Engineering Contradiction:
Improveunderstanding accuracy of user intentionsVSAvoidsystem structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the video interaction process into distinct components: video frame analysis, spatially directional action recognition, input information processing, and large model-based data processing. Each component handles a specific aspect of intention understanding, improving overall accuracy while maintaining manageable system complexity through modular design

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces spatially directional actions as an additional dimension of input beyond traditional video frames and text. This multi-dimensional approach (combining visual, spatial, and textual information) enables more comprehensive understanding of user intentions without requiring complete system redesign

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If spatially directional actions and input information are combined for processing, then the accuracy of understanding user intentions is improved, but the communication cost increases

Engineering Contradiction:
Improveaccuracy of understanding user intentionsVSAvoidcommunication cost
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system merges spatially directional actions with input information processing in a unified large model framework. By combining these data types rather than processing them separately, the system reduces redundant communication overhead while improving intention understanding accuracy through integrated multi-modal analysis

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If traditional AI systems process video interaction data, then the processing speed is moderate, but the interaction efficiency is limited

Engineering Contradiction:
Improveinteraction efficiencyVSAvoiddata processing speed
Core Design Contradiction:
ProductivityVSSpeed

Solution Approach 1:

The patent replaces traditional mechanical AI processing methods with a large model-based system that can handle multiple data types (video frames, spatially directional actions, input information) simultaneously. This substitution enables parallel processing of diverse inputs, significantly improving interaction efficiency while maintaining high processing speed through optimized model architecture

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20260017948A1Method and apparatus for video interaction based on large model, and product
Publication Date: 2026.01.15 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20260017948A1 patent drawing
  • US20260017948A1 patent drawing
  • US20260017948A1 patent drawing

AI summary

A method for video interaction, an electronic device, and a storage medium are provided. The method may include: during a video interaction with a large model, determining a target object targeted by a spatially directional action associated with a video frame in an interaction process; determining a data processing instruction for the target object based on input information linked to the spatially directional action; and using the large model to perform data processing on the target object according to the data processing instruction, thereby obtaining a data processing result.