Video Calling Action Recognition for Emotion Animation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video calling technologies lack the ability to automatically and seamlessly integrate emotional expressions and scenario-related images into video calls, limiting the richness of user interaction and emotional communication.
Innovation Solution
A video calling method and apparatus that utilizes action recognition to match user actions with preset animations, allowing for the automatic display of corresponding emotion images or scenarios on the receiving terminal, enhancing the emotional expression and interaction during video calls.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If action recognition technology is integrated into video calling to automatically send emotion images, then user interaction richness and emotional communication are improved, but system complexity and processing time increase
Solution Approach 1:
The system segments the video calling functionality into multiple independent modules: video capture module, action recognition module, emotion image selection module, and transmission module. Each module operates independently with defined interfaces, allowing the action recognition capability to be added without redesigning the entire video calling system.
Solution Approach 2:
The system performs preliminary action recognition on video frames in real-time during the video call. The action recognition module continuously analyzes user gestures and expressions, matches them against predefined action templates, and triggers corresponding emotion images before the user needs to manually select them, enabling seamless automated emotional expression.
2Extent of automation
If action recognition is performed on video images in real-time, then automated emotion expression is achieved, but processing time and computational resources increase
Solution Approach 1:
The action recognition is performed periodically at specific intervals rather than continuously on every video frame. The system samples video frames at optimized time intervals, performs action recognition only on these sampled frames, and interpolates results between samples, reducing computational load while maintaining automated emotion expression functionality.
Solution Approach 2:
The system dynamically adjusts action recognition parameters such as detection sensitivity, frame sampling rate, and matching threshold based on call context and user behavior patterns. This optimization reduces unnecessary processing while ensuring accurate emotion detection, balancing automation level with processing time requirements.
3Adaptability or versatility
If multiple preset actions and animations are supported, then emotional expression capability is enhanced, but device complexity and memory requirements increase
Solution Approach 1:
The system implements a universal emotion image library where a single set of preset emotion images serves multiple purposes across different action types. The same emotion image can be triggered by multiple different actions, and the system dynamically selects and displays appropriate images based on the recognized action context, reducing the total number of stored emotion images while maintaining diverse emotional expression capability.
Data Source
AI summary
A video calling method and video calling apparatus are provided. The video calling method includes obtaining a first video image acquired by a first terminal; performing action recognition on the first video image; and sending, in response to determining an action recognition result matches a first preset action, a first preset animation corresponding to the first preset action and the first video image to a second terminal performing video calling with the first terminal for displaying by the second terminal. With the video calling apparatus, an animation related to a scenario may be generated according to the scenario, for example, a body action of a user, provided by a video, and the animation is sent to a peer device for displaying.


