Voice Keyword Identification Using Dual Neural Classifiers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech keyword recognition methods are sensitive to manually set decision logic, making them less universal and requiring frequent tuning for different application scenarios or keywords, which limits their flexibility and generalization ability.

Innovation Solution

The method involves determining first and second speech segments from a recognized speech signal using preset classification models, generating prediction characteristics, and performing classification to determine the presence of a pre-determined keyword, thereby reducing reliance on manual decision logic and improving adaptability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If manual decision logic is used for speech keyword recognition, then the system can be simple to implement, but the system becomes sensitive to decision logic and requires frequent tuning for different scenarios

Engineering Contradiction:
Improveimplementation simplicityVSAvoidscenario adaptability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent replaces manual decision logic (mechanical/systematic approach) with a deep neural network model (intelligent/adaptive system). The DNN automatically learns optimal decision boundaries from training data, eliminating the need for manual tuning of decision logic while maintaining implementation feasibility through standardized model training procedures.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameter from manually configured decision thresholds to learned probability distributions from DNN. By transforming the decision-making parameter from fixed manual values to dynamic learned parameters, the system adapts to different scenarios without requiring manual reconfiguration.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If manual decision logic is tuned for each application scenario, then recognition accuracy can be optimized, but the complexity and time required for system configuration increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidconfiguration time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by training the deep neural network model in advance on comprehensive training data that covers multiple application scenarios. This pre-training enables the model to automatically adapt to different scenarios without requiring time-consuming manual tuning during deployment, thus maintaining high recognition accuracy while reducing configuration time.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If conventional speech keyword recognition is used, then the system structure is simple, but the system lacks universality across different keywords and scenarios

Engineering Contradiction:
Improvesystem structureVSAvoidkeyword universality
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements universality by designing a deep neural network model that can recognize multiple different keywords and adapt to various application scenarios through a single unified architecture. The DNN's ability to learn general patterns from training data enables it to handle diverse keywords and scenarios without requiring separate systems, thus achieving multi-functionality while managing complexity through modular design.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3748629B1Identification method for voice keywords, computer-readable storage medium, and computer device
Publication Date: 2023.09.06 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • EP3748629B1 patent drawingFigure 1
  • EP3748629B1 patent drawingFigure 2
  • EP3748629B1 patent drawingFigure 3~4

AI summary

An identification method for voice keywords, comprising: acquiring each first voice clip on the basis of voice signals to be identified; acquiring each first probability corresponding to the each first voice clip by means of a pre-configured first classifying model, the first probabilities comprising each probability of the first voice clips respectively corresponding to each pre-determined participle unit of pre-determined keywords; acquiring each second voice clip on the basis of the voice signals to be identified, generating a first prediction characteristic of each second voice clip on the basis of the first probabilities corresponding to first voice clips that correspond to the each second voice clip, carrying out classification by means of a pre-configured second classifying model on the basis of the each first prediction characteristic, and obtaining each second probability corresponding to each second voice clip respectively, wherein the second probabilities comprise at least one of the probability of the second voice clips corresponding to the pre-determined keywords and/or the probability of not corresponding to the pre-determined keywords; and determining whether the pre-determined keywords are present in the voice signals to be identified on the basis of the second probability. The described identification method may improve the universality.