Speech Analysis Apparatus for Context-Aware Structured Data Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition technologies fail to utilize context information embedded in sounds generated in various home spaces, limiting their effectiveness in personalized speech recognition services.

Innovation Solution

A speech analysis method and apparatus that analyze sounds in consideration of space, time, and speaker-specific characteristics, dividing speech data into segments, aligning them based on meta information, extracting keyword lists, and modeling topic information to generate structured speech data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition technology is applied without considering space, time, and speaker characteristics, then the system complexity is reduced, but the accuracy and personalization of speech recognition deteriorates

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The speech data processing is divided into multiple segments including keyword extraction, topic information modeling, and structured data generation. Each segment handles specific aspects of speech analysis (space, time, speaker characteristics) independently, allowing the system to maintain high accuracy while managing complexity through modular processing steps.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces multiple dimensions (space, time, speaker characteristics) to traditional speech recognition. By adding these dimensional layers to the analysis framework, the system achieves more accurate and personalized recognition without proportionally increasing overall system complexity, as each dimension is processed through dedicated but integrated modules.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If speech data is analyzed in raw format without segmentation and structuring, then the processing speed is faster, but the ability to extract meaningful context information deteriorates

Engineering Contradiction:
Improvecontext information retentionVSAvoiddata processing efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system performs preliminary actions by pre-segmenting speech data into meaningful units and pre-extracting keywords before full analysis. This preliminary structuring of speech data into segments with associated metadata enables faster subsequent processing while preserving context information, as the data is already organized in a analysis-ready format.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms speech data from raw audio format into structured parameters including keywords, topic information, and metadata about space, time, and speaker characteristics. This parameter transformation maintains all context information while making the data more efficient for various types of analysis and queries.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11710497B2Method and apparatus for speech analysis
Publication Date: 2023.07.25 LG ELECTRONICS INC
  • US11710497B2 patent drawing
  • US11710497B2 patent drawing
  • US11710497B2 patent drawing

AI summary

Disclosed are method and apparatus for speech analysis. The speech analysis apparatus and a server are capable of communicating with each other in a 5G communication environment by executing mounted artificial intelligence (AI) algorithms and/or machine learning algorithms. The speech analysis method and apparatus may collect and analyze speech data to build a database of structured speech data.