Distributed Speech Processing Nodes for Cross-Room Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition technologies face limitations in recognizing speech across rooms and environments due to distance constraints, network instability, and privacy concerns, with current solutions relying heavily on central devices or cloud connections that can lead to recognition failures and security issues.

Innovation Solution

A distributed speech processing system and method utilizing a network of node devices with sound acquisition and processing modules, where each device preprocesses audio signals and shares preprocessed results to perform speech recognition, enabling decentralized and concurrent recognition across multiple devices without relying on a single central node or internet connection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If speech recognition is performed locally by a single device, then recognition speed is improved, but recognition distance is limited and cross-room recognition fails

Engineering Contradiction:
Improverecognition speedVSAvoidrecognition distance
Core Design Contradiction:
SpeedVSLength of stationary object

Solution Approach 1:

The patent divides the speech recognition system into multiple independent node devices distributed across different locations. Each node performs local speech recognition independently, allowing the system to cover a larger spatial area while maintaining fast local recognition speed. The segmentation of the recognition function across multiple devices resolves the contradiction between speed and distance.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent combines the speech recognition capabilities of multiple node devices into a unified distributed system. By merging the recognition functions of multiple devices that are spatially separated, the system achieves both fast local recognition and extended coverage distance, resolving the contradiction between recognition speed and recognition distance.

Inventive Principle:
Principle #5Merging (Combining)

2Length of stationary object

If a central device collects raw audio from multiple locations, then cross-room recognition is enabled, but network bandwidth requirements increase and transmission delay increases

Engineering Contradiction:
Improverecognition distanceVSAvoidtransmission delay
Core Design Contradiction:
Length of stationary objectVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing speech recognition locally at each node device before transmitting results to the central coordinator. This preprocessing step extracts only the essential recognition information rather than transmitting raw audio data, significantly reducing network transmission time and delay while maintaining cross-room recognition capability.

Inventive Principle:
Principle #10Preliminary action

3Length of stationary object

If a central device collects raw audio from multiple locations, then cross-room recognition is enabled, but network bandwidth requirements increase

Engineering Contradiction:
Improverecognition distanceVSAvoidnetwork bandwidth
Core Design Contradiction:
Length of stationary objectVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential speech recognition results from each node device rather than transmitting the complete raw audio data. This extraction approach transmits only the necessary information (recognition outcomes) over the network, dramatically reducing bandwidth consumption while enabling cross-room recognition through the distributed node system.

Inventive Principle:
Principle #2Taking out (Extraction)

4Measurement precision

If speech is uploaded to cloud server, then recognition accuracy is improved, but user privacy and security issues arise

Engineering Contradiction:
Improverecognition accuracyVSAvoidprivacy and security risks
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent implements self-service by enabling each node device to perform speech recognition independently using local computational resources. This eliminates the need to upload speech data to external cloud servers, thereby maintaining user privacy and security while achieving recognition accuracy through distributed local processing at each node.

Inventive Principle:
Principle #25Self-service

5Device complexity

If a single device serves as control center, then system complexity is reduced, but system reliability decreases due to single point of failure

Engineering Contradiction:
Improvesystem complexityVSAvoidsystem reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the control function from a single centralized device to multiple distributed node devices. Each node maintains independent speech recognition capability, eliminating the single point of failure while keeping system complexity manageable through standardized node implementations with clear communication protocols.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by enabling each node device to possess independent speech recognition capabilities tailored to its local environment. This distributed architecture improves system reliability through redundancy while maintaining reasonable complexity through consistent local processing at each node.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240062764A1Distributed speech processing system and method
Publication Date: 2024.02.22 ESPRESSIF SYST SHANGHAI
  • US20240062764A1 patent drawing
  • US20240062764A1 patent drawing
  • US20240062764A1 patent drawing

AI summary

A distributed speech processing system and a method therefor is provided. The system includes: a plurality of node devices in a network, wherein each node device includes a processor, a memory, a communication module and a sound processing module, and at least one node device comprises a sound acquisition module configured to acquire an audio signal; the sound processing module is configured to preprocess the audio signal to obtain a first sound preprocessed result; the communication module is configured to send the first sound preprocessed result to one or more node devices in the network; the communication module is further configured to receive one or more second sound preprocessed results from at least one other node device over the network; and the sound processing module is further configured to perform speech recognition based on the first sound preprocessed result and/or the one or more second sound preprocessed results.