Shared Speech Processing Network for Resource-Constrained Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-performance neural network-based speech processing is challenging in memory- or computation-constrained devices due to large memory footprints and increased computation requirements, limiting the concurrent execution of multiple speech processing applications.
Innovation Solution
A shared speech processing network generates a common output representation that can be used by multiple speech application modules, reducing duplication of processing and resource usage, and is trained to optimize performance for various speech processing tasks such as speech recognition and speaker recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple separate speech processing models are used for different speech applications, then speech processing performance for each application is improved, but computation and memory usage increase significantly
Solution Approach 1:
The patent combines multiple speech processing models into a single shared speech processing network that serves multiple speech applications simultaneously. The shared network includes input processing layers and task-specific output layers that can handle different speech tasks (e.g., speech recognition, speaker identification, keyword detection) using a single unified model, thereby reducing the overall computation and memory footprint compared to running separate models for each application.
Solution Approach 2:
The shared speech processing network is designed with universal functionality to support multiple speech applications through a single model. The network architecture includes configurable task-specific output layers that can be adapted to different speech processing tasks, allowing one model to perform multiple functions that previously required separate specialized models, thus reducing resource consumption while maintaining performance.
2Quantity of substance
If the speech processing network is designed to support multiple speech applications, then device resource efficiency is improved, but the network architecture complexity increases
Solution Approach 1:
The shared speech processing network is segmented into distinct functional components: input processing layers that handle common speech features, task-specific output layers that handle different speech applications, and configurable connection mechanisms. This segmentation allows the network to support multiple applications while maintaining manageable architecture complexity through modular design, where each component has a specific responsibility and can be independently optimized.
Data Source
AI summary
A device to process speech includes a speech processing network that includes an input configured to receive audio data. The speech processing network also includes one or more network layers configured to process the audio data to generate a network output. The speech processing network includes an output configured to be coupled to multiple speech application modules to enable the network output to be provided as a common input to each of the multiple speech application modules.


