AI 3D Agency Motion Control for Spoken Content Indication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing artificial intelligence technologies cannot effectively control the motion of three-dimensional (3D) agencies to indicate specific positions of 3D content in response to spoken sentences, limiting the provision of engaging and fun 3D content services.
Innovation Solution
An artificial intelligence device and method that extracts keywords from spoken sentences, detects positions of objects and text in related content, maps these positions to three-dimensional coordinates, and controls the motion of a 3D agency to perform utterance and indication operations accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing AI technology is used to implement 3D images from 2D images, then 3D content can be generated, but the motion of the 3D target object cannot be controlled in accordance with spoken sentences
Solution Approach 1:
The system segments the motion control task into distinct components: keyword extraction from spoken sentences, object detection in related content, position mapping to 3D coordinates, and motion generation. This segmentation allows each component to be handled by specialized modules, enabling motion control without overwhelming system complexity.
Solution Approach 2:
The patent introduces intermediary elements including a keyword extraction module that bridges spoken sentences and content understanding, and a position mapping module that translates 2D positions to 3D coordinates. These intermediaries enable controlled interaction between the spoken command and the 3D agency motion without requiring direct complex coupling.
2Productivity
If the system maps positions to 3D coordinates and controls agency motion, then user engagement is enhanced, but processing complexity increases
Solution Approach 1:
The system performs preliminary actions by pre-detecting objects and text positions in the related content before generating the spoken sentence response. The position mapping to 3D coordinates is prepared in advance, allowing the 3D agency to immediately execute motion commands without real-time computation delays, thus enhancing user engagement while managing processing complexity.
Solution Approach 2:
The system employs self-service mechanisms through automated keyword extraction from spoken sentences and automatic position detection in content. The AI device independently performs these processing tasks without requiring manual intervention or complex external processing, improving responsiveness and user engagement while keeping the overall system manageable.
3Measurement precision
If keyword extraction and position detection are performed, then accurate indication is achieved, but processing time increases
Solution Approach 1:
The system extracts only the essential keyword from the spoken sentence that is relevant to position indication, rather than processing the entire sentence. This selective extraction focuses computational resources on the critical information needed for accurate position detection, achieving measurement precision while minimizing unnecessary processing time.
Solution Approach 2:
The system performs partial action by detecting only the positions of objects and text that are directly related to the extracted keyword, rather than analyzing all elements in the content. This selective detection approach maintains accurate position indication for relevant elements while reducing overall processing time by excluding irrelevant computations.
Data Source
AI summary
An artificial intelligence device is configured to: when a spoken sentence of a three-dimensional agency is generated, extract a keyword of the spoken sentence; acquire related content associated with the keyword of the spoken sentence to detect positions of an object and text in the related content; when an object and text corresponding to the keyword of the spoken sentence exist in the related content, map the positions of the object and the text corresponding to the keyword of the spoken sentence to three-dimensional coordinates; output the related content to a surrounding space of the three-dimensional agency; and control the operation of the three-dimensional agency so that the three-dimensional agency performs an utterance operation corresponding to the spoken sentence and an indication operation of indicating three-dimensional coordinates at which the object and the text of the related content are located.


