Semantic Prediction Network for Direct Speech-to-Semantics Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies, particularly in smart home systems, rely on a three-level cascade solution that is inefficient and resource-intensive, requiring significant computational resources for converting speech to text and then to semantics, and lacks effective decoding methods.
Innovation Solution
A method and apparatus for training a semantic prediction network with an encoder and decoder network structure, incorporating convolutional and long short-term memory layers, where the syllable classification network is jointly trained with the semantic prediction network using semantic and syllable labels as constraints, reducing the need for conventional acoustic decoding and improving prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional three-level cascade solution is used for speech recognition, then speech to text conversion can be achieved, but computational resources and time consumption increase significantly
Solution Approach 1:
The patent merges the acoustic model, language model, and semantic model into a unified neural network architecture that processes speech directly to semantic understanding, eliminating the sequential processing stages of conventional cascade solutions and reducing overall processing time
Solution Approach 2:
The patent extracts and removes the text intermediate representation from the speech recognition pipeline, enabling direct mapping from speech features to semantic meaning without requiring speech-to-text conversion as a separate step
2Measurement precision
If conventional three-level cascade solution with decoding methods is used, then speech to semantics conversion can be achieved, but device complexity and computational burden increase
Solution Approach 1:
The patent combines multiple separate models (acoustic, language, semantic) into a single integrated neural network with shared layers, reducing system complexity while maintaining the functional capabilities of each component through joint training
Solution Approach 2:
The patent implements preliminary action by pre-training the shared encoder layers on large-scale speech data before fine-tuning the entire system for specific semantic recognition tasks, reducing the computational burden during deployment
3Ease of manufacture
If speech is converted to text before semantic recognition, then conventional processing pipeline can be followed, but resource overhead increases
Solution Approach 1:
The patent extracts and eliminates the unnecessary speech-to-text conversion step from the processing pipeline, enabling direct semantic recognition from speech features and reducing computational energy consumption while maintaining implementation feasibility through a streamlined architecture
Data Source
AI summary
Embodiments of the present disclosure disclose a method, apparatus and device for training a network, and a storage medium, relate to the field of artificial intelligence technology such as deep learning and speech analysis. A semantic prediction network comprises: an encoder network and at least one decoder network; and a particular solution is: acquiring a first speech feature of a target speech sample; the target speech sample being a synthesized speech sample or a real speech sample, the synthesized speech sample being attached with a sample syllable label and a semantic label comprising a value of the domain, and the real speech sample being attached with a sample syllable label; and jointly training an initial semantic prediction network and a syllable classification network using the first speech feature of the target speech sample, to obtain a trained semantic prediction network.


