Semantic Prediction Network for Direct Speech-to-Semantics Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition technologies, particularly in smart home systems, rely on a three-level cascade solution that is inefficient and resource-intensive, requiring significant computational resources for converting speech to text and then to semantics, and lacks effective decoding methods.

Innovation Solution

A method and apparatus for training a semantic prediction network with an encoder and decoder network structure, incorporating convolutional and long short-term memory layers, where the syllable classification network is jointly trained with the semantic prediction network using semantic and syllable labels as constraints, reducing the need for conventional acoustic decoding and improving prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional three-level cascade solution is used for speech recognition, then speech to text conversion can be achieved, but computational resources and time consumption increase significantly

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent merges the acoustic model, language model, and semantic model into a unified neural network architecture that processes speech directly to semantic understanding, eliminating the sequential processing stages of conventional cascade solutions and reducing overall processing time

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent extracts and removes the text intermediate representation from the speech recognition pipeline, enabling direct mapping from speech features to semantic meaning without requiring speech-to-text conversion as a separate step

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If conventional three-level cascade solution with decoding methods is used, then speech to semantics conversion can be achieved, but device complexity and computational burden increase

Engineering Contradiction:
Improvesemantic recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple separate models (acoustic, language, semantic) into a single integrated neural network with shared layers, reducing system complexity while maintaining the functional capabilities of each component through joint training

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements preliminary action by pre-training the shared encoder layers on large-scale speech data before fine-tuning the entire system for specific semantic recognition tasks, reducing the computational burden during deployment

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If speech is converted to text before semantic recognition, then conventional processing pipeline can be followed, but resource overhead increases

Engineering Contradiction:
Improveimplementation feasibilityVSAvoidcomputational energy consumption
Core Design Contradiction:
Ease of manufactureVSUse of energy by moving object

Solution Approach 1:

The patent extracts and eliminates the unnecessary speech-to-text conversion step from the processing pipeline, enabling direct semantic recognition from speech features and reducing computational energy consumption while maintaining implementation feasibility through a streamlined architecture

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11823660B2Method, apparatus and device for training network and storage medium
Publication Date: 2023.11.21 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11823660B2 patent drawing
  • US11823660B2 patent drawing
  • US11823660B2 patent drawing

AI summary

Embodiments of the present disclosure disclose a method, apparatus and device for training a network, and a storage medium, relate to the field of artificial intelligence technology such as deep learning and speech analysis. A semantic prediction network comprises: an encoder network and at least one decoder network; and a particular solution is: acquiring a first speech feature of a target speech sample; the target speech sample being a synthesized speech sample or a real speech sample, the synthesized speech sample being attached with a sample syllable label and a semantic label comprising a value of the domain, and the real speech sample being attached with a sample syllable label; and jointly training an initial semantic prediction network and a syllable classification network using the first speech feature of the target speech sample, to obtain a trained semantic prediction network.