Voice and Image Recognition Integration in Electronic Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current electronic devices are unable to efficiently integrate voice recognition and image input to provide a comprehensive artificial intelligence service, limiting their ability to perform tasks that require both voice and image processing simultaneously.
Innovation Solution
An integrated intelligence system comprising a user terminal, intelligence server, and service server that processes voice and image inputs through a client module, natural language platform, and capsule database to generate plans for task execution, allowing for the association of voice and image inputs to perform unified operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If voice recognition service and image analysis service are provided separately, then each service can be optimized independently, but the electronic device cannot process voice and image inputs in association to perform unified tasks
Solution Approach 1:
The patent merges voice recognition service and image analysis service into a unified processing architecture. The voice input module and image input module both feed into the same natural language processing module, enabling combined processing of voice and image inputs for unified task execution, thereby resolving the contradiction between adaptability and complexity.
Solution Approach 2:
The natural language processing module serves multiple functions: processing voice inputs, analyzing image inputs, and generating unified responses. This multi-functional design enables the system to handle both separate and combined inputs efficiently, improving adaptability while managing complexity through shared processing resources.
2Productivity
If separate voice and image services are provided, then service optimization is easier, but the electronic device cannot perform tasks requiring both voice and image processing simultaneously
Solution Approach 1:
The patent combines voice and image processing pipelines into a unified architecture where both input modalities are processed simultaneously through the natural language processing module, enabling efficient task execution that requires both voice and image analysis capabilities.
Solution Approach 2:
The system performs preliminary processing of both voice and image inputs through their respective modules before feeding them to the natural language processing module, enabling efficient simultaneous processing and task execution without requiring sequential handling of different modalities.
3Adaptability or versatility
If integrated voice and image processing is implemented, then comprehensive AI service capability is achieved, but system complexity increases
Solution Approach 1:
The patent merges voice recognition and image analysis services into a unified processing architecture where both input types converge in the natural language processing module, achieving integrated capability while managing complexity through shared processing components.
Solution Approach 2:
The natural language processing module acts as an intermediary that receives processed information from both voice and image modules, translating them into unified responses. This intermediary design simplifies the overall system architecture by providing a single convergence point for multi-modal inputs.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An electronic device is provided. The electronic device includes a microphone, a display, a camera, a processor, and a memory. The processor is configured to receive a first utterance input through the microphone. The processor is also configured to obtain first recognized data from a first image displayed on the display or stored in the memory. The processor is further configured to store the first recognized data in association with the first utterance input when the obtained first recognized data matches the first utterance input. Additionally, the processor is configured to activate the camera when the first recognized data does not match the first utterance input. The processor is also configured to obtain second recognized data from a second image collected through the camera and store the second recognized data in association with the first utterance input when the obtained second recognized data matches the first utterance input.