Speech Recognition via Merged Word Trees for Position and Language

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition methods face inaccuracies due to the lack of integration of user-related position and language information during the decoding process, leading to incomplete resource involvement and incorrect answer activation.

Innovation Solution

The method integrates user position and language information by constructing and merging word trees of different dimensions, allowing resource information to participate in the decoding process and improving recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If resource information is introduced after obtaining candidate words through decoding, then the decoding process can be completed, but the introduced resource information cannot participate in the decoding and activation cannot be obtained when no correct answer exists in the candidate words

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidresource information involvement
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent applies preliminary action by constructing the word tree from resource information (position, language, authorization) before the decoding process begins. This pre-construction allows the word tree to be integrated into the decoding process, enabling resource information to participate in guiding candidate word generation and activation determination, thereby resolving the issue of resource information being introduced too late to affect decoding outcomes

Inventive Principle:
Principle #10Preliminary action

2Reliability

If user position and language information are integrated through word tree merging, then speech recognition accuracy is improved, but the system complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidword tree construction and merging complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the resource information into distinct dimensions (position, language, authorization) and constructing separate word trees for each dimension. These segmented word trees are then merged to form a comprehensive structure. This segmentation approach manages complexity by organizing information dimensionally while still achieving integrated speech recognition improvement through the merging process

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4216209B1Speech recognition method, terminal and storage medium
Publication Date: 2024.09.11 GUANGZHOU XIAOPENG MOTORS TECH CO LTD
  • EP4216209B1 patent drawingFigure 1
  • EP4216209B1 patent drawingFigure 2
  • EP4216209B1 patent drawingFigure 3

AI summary

Provided are a speech recognition method and apparatus, a terminal, and a storage medium. The method includes: obtaining, in response to a speech request of a user, position information of the user and language information of the speech request, the position information of the user and the language information of the speech request having each a word tree of a corresponding dimension of information, word trees of different dimensions of information being used to merge into a merged word tree; and recognizing the speech request of the user based on the merged word tree. With the word tree-based approach, the position information of the user and the language information that are relevant to user information are integrated and introduced to recognize the speech request of the user, allowing resource information introduced in the form of word tree to participate in a decoding process of the speech request, and avoiding a problem of inaccurate language recognition.