AI-Enhanced Video Calling With Real-Time Animation Overlay
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio/video calling technologies lack intelligent functions beyond basic functionalities, and require users to install additional applications to access enhanced features.
Innovation Solution
An AI component recognizes specific content in audio and video streams during a call, superimposing animation effects corresponding to the recognized content, without the need for additional client applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If basic audio/video calling functionality is provided, then call capability is achieved, but intelligence and趣味性 are insufficient
Solution Approach 1:
An AI component is introduced as an intermediary between the media server and the client. The media server copies audio and video streams to this AI component, which performs content recognition and triggers animation effects. This intermediary handles the complex intelligence functions centrally, keeping the client apparatus simple while adding versatile intelligent capabilities to the calling system.
2Adaptability or versatility
If additional client applications are installed to provide intelligent functions, then intelligence is improved, but ease of operation deteriorates
Solution Approach 1:
The patent integrates multiple functions into a single universal calling system. The media server performs both basic audio/video transmission and intelligent content recognition through the AI component. Animation effects are superimposed on the existing call interface, allowing users to access intelligent functions without installing separate applications, thus maintaining ease of operation while enhancing adaptability.
3Adaptability or versatility
If content recognition and animation superimposition are added, then user experience is improved, but processing complexity increases
Solution Approach 1:
The AI component serves as a dedicated intermediary that handles all content recognition and animation trigger decisions. The media server simply copies streams to this component and superimposes animations based on its outputs. This centralized processing architecture manages the complexity of content analysis and animation generation in one location, improving user experience without distributing complex processing across multiple devices.
Data Source
AI summary
Provided is a method and device for audio/video calling. According to the present disclosure, after an audio/video call between a calling user and a called user is anchored to a media server, an AI component is used to receive an audio stream and a video stream of the audio/video call between the calling user and the called user, which are copied by the media server; and the AI component recognizes specific content in the audio stream and/or the video stream, and the media server superimposes on the audio/video call between the calling user and the called user an animation effect corresponding to the specific content. The problem of single audio/video calling functionality in the related art is solved, and the interestingness and intellectualization level of audio/video calls are increased.


