Offline Speech Chip Architecture for Local Voice Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech chips rely solely on remote servers for speech recognition, limiting their functionality to online environments and failing to support offline applications.
Innovation Solution
A speech chip with integrated modules for speech-to-text and text-to-speech conversion, including a first processing module for system control, a second processing module for model-based conversions, and a third module for digital signal processing, along with a power management system for efficient power usage and an image processing module for extended functionality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition is performed through remote server, then speech recognition function is achieved, but offline application support is lost
Solution Approach 1:
The speech chip is divided into multiple functional modules: audio engine module for audio processing, first processing module for control functions, second processing module for speech-to-text conversion, and third processing module for digital signal processing. This segmentation allows offline speech recognition capabilities to be integrated without overwhelming device complexity.
Solution Approach 2:
The speech chip integrates multiple functions including speech-to-text conversion, text-to-speech conversion, wake-up word detection, and digital signal processing within a single device. The second processing module handles both speech recognition and text synthesis, enabling the device to support various offline applications with a unified architecture.
2Use of energy by moving object
If all modules are powered on continuously, then processing speed is maintained, but power consumption increases
Solution Approach 1:
The power supply module dynamically adjusts the power state of different modules based on operational needs. The third processing module can be selectively powered on for wake-up word detection, and the second processing module is activated only when speech recognition or text-to-speech conversion is required. This dynamic power management reduces overall power consumption while maintaining processing speed when needed.
Solution Approach 2:
The speech detection unit periodically monitors for speech input and wake-up words, activating the main processing modules only when speech is detected. This periodic activation pattern allows the system to consume minimal power during idle periods while ensuring rapid response when speech input occurs.
Data Source
Figure 1~2
Figure 3~4
Figure 5~7
AI summary
The present disclosure discloses a speech chip and an electronic device, relates to a field of data processing technologies, and more particularly, relates to speech technologies. In detail, the speech chip includes a first processing module, a second processing module and a third processing module, wherein the first processing module is configured to run an operating system, and to perform data scheduling on modules other than the first processing module in the chip; the second processing module is configured to perform a mutual conversion between speech and text based on a speech model; and the third processing module is configured to perform digital signal processing on inputted speech. Embodiments of the present disclosure provide the speech chip and the electronic device, such that a smart speech product supports applications in offline scenes.