Voice Recognition Using Structured User Data for Intent Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

User utterances can be interpreted differently based on context and user intent, leading to inaccuracies in voice recognition systems.

Innovation Solution

An electronic device processes natural language inputs through an encoder and data sensing unit to convert them into structured data, identifying specific data types and generating additional data based on this identification to determine intended tasks accurately.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice recognition systems process natural language inputs directly, then the system structure remains simple, but the accuracy of recognizing user intent deteriorates due to various interpretations of the same utterance

Engineering Contradiction:
Improveaccuracy of recognizing user intentVSAvoidsystem structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the voice recognition system into multiple processing units: an encoder that converts natural language to first input data, a data sensing unit that identifies data corresponding to the input, and a task determination unit that processes second input data. This segmentation allows each unit to specialize in specific processing tasks, improving overall intent recognition accuracy while managing system complexity through modular design

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary data structures (first input data from encoder, second input data from data sensing unit) that mediate between the raw natural language input and the final task determination. These intermediaries transform and enrich the information progressively, enabling more accurate intent recognition without requiring the entire system to process all information simultaneously

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the system uses basic natural language processing, then the processing speed remains fast, but the ability to handle context-dependent interpretations deteriorates

Engineering Contradiction:
Improveability to handle context-dependent interpretationsVSAvoidprocessing speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The encoder performs preliminary action by converting natural language inputs into structured first input data before the main processing occurs. This pre-processing step organizes the information in advance, enabling the data sensing unit and task determination unit to work more efficiently with pre-structured data, thus improving adaptability without proportionally increasing processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the one-dimensional natural language input into multi-dimensional data representations through the encoder and data sensing unit. By converting text into structured formats with multiple attributes and relationships, the system gains the ability to process context-dependent interpretations across multiple dimensions while maintaining processing efficiency

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12400082B2Electronic device for providing voice recognition service using user data and operating method thereof
Publication Date: 2025.08.26 SAMSUNG ELECTRONICS CO LTD
  • US12400082B2 patent drawing
  • US12400082B2 patent drawing
  • US12400082B2 patent drawing

AI summary

An electronic device including an input device, a processor, and a memory that stores instructions are provided. The instructions, when executed by the processor, cause the electronic device to obtain a natural language input by using the input device, to convert the natural language input into first input data, to identify data corresponding to at least part of the natural language input in a specified type of data included in the memory, to generate second input data based on the identification result, and to determine at least one task according to the natural language input based on the first input data and the second input data.