Offline Speech Chip Architecture for Local Voice Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech chips rely solely on remote servers for speech recognition, limiting their functionality to online environments and failing to support offline applications.

Innovation Solution

A speech chip with integrated modules for speech-to-text and text-to-speech conversion, including a first processing module for system control, a second processing module for model-based conversions, and a third module for digital signal processing, along with a power management system for efficient power usage and an image processing module for extended functionality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition is performed through remote server, then speech recognition function is achieved, but offline application support is lost

Engineering Contradiction:
Improveoffline application supportVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The speech chip is divided into multiple functional modules: audio engine module for audio processing, first processing module for control functions, second processing module for speech-to-text conversion, and third processing module for digital signal processing. This segmentation allows offline speech recognition capabilities to be integrated without overwhelming device complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The speech chip integrates multiple functions including speech-to-text conversion, text-to-speech conversion, wake-up word detection, and digital signal processing within a single device. The second processing module handles both speech recognition and text synthesis, enabling the device to support various offline applications with a unified architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Use of energy by moving object

If all modules are powered on continuously, then processing speed is maintained, but power consumption increases

Engineering Contradiction:
Improvepower consumptionVSAvoidprocessing speed
Core Design Contradiction:
Use of energy by moving objectVSSpeed

Solution Approach 1:

The power supply module dynamically adjusts the power state of different modules based on operational needs. The third processing module can be selectively powered on for wake-up word detection, and the second processing module is activated only when speech recognition or text-to-speech conversion is required. This dynamic power management reduces overall power consumption while maintaining processing speed when needed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The speech detection unit periodically monitors for speech input and wake-up words, activating the main processing modules only when speech is detected. This periodic activation pattern allows the system to consume minimal power during idle periods while ensuring rapid response when speech input occurs.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentEP3866162B1Speech chip and electronic device
Publication Date: 2025.09.24 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • EP3866162B1 patent drawingFigure 1~2
  • EP3866162B1 patent drawingFigure 3~4
  • EP3866162B1 patent drawingFigure 5~7

AI summary

The present disclosure discloses a speech chip and an electronic device, relates to a field of data processing technologies, and more particularly, relates to speech technologies. In detail, the speech chip includes a first processing module, a second processing module and a third processing module, wherein the first processing module is configured to run an operating system, and to perform data scheduling on modules other than the first processing module in the chip; the second processing module is configured to perform a mutual conversion between speech and text based on a speech model; and the third processing module is configured to perform digital signal processing on inputted speech. Embodiments of the present disclosure provide the speech chip and the electronic device, such that a smart speech product supports applications in offline scenes.