Speech Animation Engine Automatic Data Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech animation systems require user intervention for inputting text, limiting scalability and flexibility in generating speech synchronized with facial expressions.

Innovation Solution

A system that automatically retrieves and generates speech animation data in real-time without user intervention, using a speech animation engine and client application communication, which includes retrieving data for speech animation, generating speech animation, and sending responses to the client application, allowing for dynamic and personalized content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual text input is required for speech animation, then the system can generate speech synchronized with facial expressions, but the scalability and ease of operation deteriorate due to user intervention requirements

Engineering Contradiction:
Improveease of operationVSAvoidextent of automation
Core Design Contradiction:
Ease of operationVSExtent of automation

Solution Approach 1:

The system automatically retrieves text content from data sources without requiring manual user input. The speech animation engine autonomously fetches data, processes it through text-to-speech conversion, and generates synchronized facial animations, making the system self-sufficient and eliminating the need for user intervention in the text input process.

Inventive Principle:
Principle #25Self-service

2Productivity

If manual text input is required for speech animation, then the system can generate speech synchronized with facial expressions, but the productivity and scalability worsen due to limited text input capabilities

Engineering Contradiction:
ImproveproductivityVSAvoidextent of automation
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The system automatically retrieves text content from data sources without requiring manual user input. The speech animation engine autonomously fetches data, processes it through text-to-speech conversion, and generates synchronized facial animations, making the system self-sufficient and eliminating the need for user intervention in the text input process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system can accept text from multiple sources including direct input, file uploads, and automated data retrieval from various formats. This multi-functional capability allows the same speech animation engine to handle diverse data sources, increasing productivity and scalability without requiring separate processing pipelines for each input method.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If static text data is used for speech animation, then the system can generate speech content, but the adaptability and personalization deteriorate due to inability to dynamically adapt content

Engineering Contradiction:
ImproveadaptabilityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system dynamically adapts speech content based on retrieved data characteristics, user preferences, and context information. Instead of using fixed static text, the engine processes variable data sources and adjusts the generated speech and facial animations in real-time, enabling personalization and dynamic content adaptation without requiring complex manual configuration.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS7529674B2Speech animation
Publication Date: 2009.05.05 SAP SE
  • US7529674B2 patent drawing
  • US7529674B2 patent drawing
  • US7529674B2 patent drawing

AI summary

Methods and systems, including computer program products, for speech animation. The system includes a speech animation engine and a client application in communication with the speech animation engine. The client application sends a request for speech animation to the speech animation engine. The request identifies data to be used to generate the speech animation, where speech animation is speech synchronized with facial expressions. The client application receives a response from the speech animation engine. The response identifies the generated speech animation. The client application uses the generated speech animation to animate a talking agent displayed on a user interface of the client application. The speech animation engine receives the request for speech animation from the client application, retrieves the data identified in the request without user intervention, generates the speech animation using the retrieved data and sends the response identifying the generated speech animation to the client application.