Digital Human Media Playback With Entity-Based Avatar Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital human technologies are limited to single application scenarios and primarily replace traditional voice assistant avatars, lacking versatility and interactivity.
Innovation Solution
A server and display apparatus system that receives speech input, recognizes it to obtain entity and media data, and plays digital human data, including image and speech, to enhance interaction and customization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If digital human technology is limited to single application scenarios and traditional voice assistant avatars, then device complexity is reduced, but adaptability and versatility are limited
Solution Approach 1:
The patent implements a universal digital human system that can operate across multiple application scenarios including but not limited to virtual anchors, educational lecturers, and customer service assistants. The system uses a common digital human generation platform that can be adapted to different scenarios through configurable parameters and templates, allowing one system to serve multiple functions without requiring separate specialized systems for each application.
Solution Approach 2:
The system segments the digital human functionality into independent modular components including speech recognition modules, entity extraction modules, digital human generation modules, and playback modules. This segmentation allows the system to maintain complexity management while achieving versatility, as each module can be independently configured and combined for different application scenarios.
2Adaptability or versatility
If digital human avatar display only replaces traditional voice assistant avatar, then ease of operation is maintained, but adaptability and interactivity are limited
Solution Approach 1:
The system implements dynamic avatar selection and configuration capabilities where users can choose from multiple digital human avatars with different characteristics, styles, and personalities. The system dynamically adapts the avatar display based on the application scenario and user preferences, transforming the static avatar replacement into a dynamic, interactive experience that maintains ease of operation through automated selection options.
Solution Approach 2:
The system incorporates feedback mechanisms where user interactions with digital humans are analyzed and used to improve future interactions. The speech recognition and entity extraction systems provide feedback loops that enable the digital humans to learn from and adapt to user behavior patterns, enhancing interactivity while maintaining operational simplicity through automated learning.
3Loss of information
If speech recognition and entity extraction are implemented, then information processing capability is improved, but processing time and computational resources increase
Solution Approach 1:
The system performs preliminary speech recognition and entity extraction processing as soon as speech input is received, preparing recognition results and entity data in advance before they are needed for digital human generation. This preliminary action reduces the perceived processing time by having critical information extraction completed proactively rather than reactively.
Solution Approach 2:
The system implements partial processing strategies where speech recognition and entity extraction are performed to the necessary degree of accuracy for the specific application scenario. Not all speech inputs require full entity extraction and analysis - the system adjusts the depth of processing based on the context, performing only the necessary level of information extraction to balance accuracy with processing efficiency.
Data Source
AI summary
A server is provided. The server is configured to: receive speech data input from a user and sent from a display apparatus; recognize the speech data to obtain a recognition result; based on that the recognition result includes entity data, obtain media resource data corresponding to the recognition result, and digital human data corresponding to the entity data; wherein the entity data includes a human name and/or a media resource name, the digital human data includes image data and a broadcast speech of a digital human, and the media resource data includes audio and video data or interface data; and send the digital human data and the media resource data to the display apparatus for the display apparatus to play the audio and video data or display the interface data, and play an image and a speech of the digital human according to the digital human data.


