Offline Voice Recognition via Compressed Acoustic Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice recognition technologies rely on network connectivity, limiting their functionality and usability, especially in mobile devices where resource constraints hinder the execution of complex acoustic models.

Innovation Solution

Implementing a method and device for offline voice recognition by collecting and decoding voice information using pre-compressed acoustic and language models, which enables voice recognition without network dependency and optimizes data processing through data compression and parallel computation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice recognition is performed online using cloud servers, then recognition accuracy is maintained, but network dependency increases and offline functionality is lost

Engineering Contradiction:
Improverecognition accuracyVSAvoidoffline functionality
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent divides the voice recognition system into separate functional modules: acoustic model, language model, and decoding module. These segmented components can be selectively deployed and executed on mobile devices, enabling offline functionality while maintaining recognition accuracy through distributed computation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a data compression technology as an intermediary between the complex acoustic models and mobile device resources. This intermediary enables the transmission and execution of compressed models on resource-constrained devices, bridging the gap between cloud-based accuracy and offline capability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If complex acoustic models are deployed on mobile devices, then offline voice recognition is achieved, but device resources are exceeded and processing becomes infeasible

Engineering Contradiction:
Improveoffline voice recognitionVSAvoidmodel complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies data compression to transform the acoustic model parameters into a compressed format. This parameter transformation reduces the model size by approximately 95% while preserving recognition accuracy, making complex models executable on mobile devices with limited resources

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent performs data compression on acoustic models in advance before deployment to mobile devices. This preliminary processing reduces model complexity beforehand, enabling offline voice recognition without requiring significant device resources during actual recognition operations

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If uncompressed acoustic models are used, then model accuracy is maintained, but decoding time increases significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoiddecoding time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent transforms model parameters through compression, creating a compact representation that maintains accuracy while enabling faster processing. The compressed parameters are optimized for rapid computation on mobile devices, reducing decoding time without sacrificing recognition precision

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9805712B2Method and device for recognizing voice
Publication Date: 2017.10.31 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US9805712B2 patent drawing
  • US9805712B2 patent drawing
  • US9805712B2 patent drawing

AI summary

A method for recognizing a voice and a device for recognizing a voice are provided. The method includes: collecting voice information input by a user; extracting characteristics from the voice information to obtain characteristic information; decoding the characteristic information according to an acoustic model and a language model obtained in advance to obtain recognized voice information, wherein the acoustic model is obtained by data compression in advance.