Speech Recognition System for Voice-Driven Application Launching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing interaction devices require users to manually tap icons to start third-party applications, lacking intelligent speech recognition capabilities for initiating these applications.

Innovation Solution

A method and system for speech recognition that parses speech signals to determine target semantics, identifies corresponding third-party applications, and starts them without manual intervention by accessing a third-party application registry, utilizing semantic analysis and API invocation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech recognition is used to start third-party applications, then ease of operation is improved, but device complexity increases due to the need for semantic parsing and application registry integration

Engineering Contradiction:
Improveease of operationVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system segments the speech recognition functionality into separate modules: speech signal acquisition, semantic parsing, application object determination, and application starting. This modular approach allows each component to be optimized independently while maintaining overall system functionality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an application object as an intermediary layer between the speech recognition system and the actual third-party applications. This intermediary enables the system to manage application launching without direct integration with each application, reducing overall system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If manual icon tapping is used to start applications, then device complexity is reduced, but loss of time increases due to the need for user navigation and selection

Engineering Contradiction:
Improveloss of timeVSAvoiddevice complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by maintaining an application registry that pre-maps application names to their corresponding application objects. When a user speaks an application name, the system can quickly retrieve and start the application without requiring user navigation through menus or icon searching.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The speech recognition system automatically parses the speech signal, determines the target application object, and starts the application without requiring user intervention for each step. The system serves itself by autonomously completing the entire application launching process based on the user's spoken command.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11676605B2Method, interaction device, server, and system for speech recognition
Publication Date: 2023.06.13 HUAWEI TECH CO LTD
  • US11676605B2 patent drawing
  • US11676605B2 patent drawing
  • US11676605B2 patent drawing

AI summary

A method, an apparatus, and a system for speech recognition are provided. a third-party application corresponding to a speech signal of a user can be determined according to the speech signal and by means of semantic analysis; and third-party application registry information is searched for and a third-party program is started, so that the user does not need to tap the third-party application to start the corresponding program, thereby providing more intelligent service for the user and facilitating use for the user.