Virtual Assistant Data Loading for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic devices face challenges in improving response speed for user speech while minimizing memory load during speech recognition and natural language understanding processes, as current methods either overload memory or delay response due to data loading and initialization.

Innovation Solution

An electronic device with a non-volatile memory storing virtual assistant model data classified by domains and commonly used data, where a processor initiates loading of relevant data into volatile memory upon receiving a trigger input, performs speech recognition, and maintains the loaded data for a predetermined period to efficiently process user speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If all data for speech recognition and natural language understanding is pre-loaded in background, then response speed for user speech is improved, but memory load increases significantly

Engineering Contradiction:
Improveresponse speedVSAvoidmemory load
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The patent segments the virtual assistant model data into multiple categories: trigger input recognition data, domain-specific data, and commonly used data. This segmentation allows the system to load only the necessary segments into volatile memory at any given time, reducing memory load while maintaining fast response capability for user speech processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action by pre-classifying and organizing data in non-volatile memory before actual speech processing is needed. The data is structured with classification information that enables rapid identification and selective loading of required data segments when a trigger input is detected, avoiding the need to load all data into memory.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If data is loaded and initialized before receiving user speech, then processing capability is improved, but response time is delayed due to loading and initialization time

Engineering Contradiction:
Improveprocessing capabilityVSAvoidresponse time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary organization of data in non-volatile memory with classification metadata before actual speech processing. When a trigger input is received, the system can immediately identify which pre-organized data segments are needed and load them rapidly, eliminating the delay associated with loading and initializing all data from scratch.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamic data loading strategies where the system adjusts which data segments to load into volatile memory based on the actual speech processing requirements. The system can load domain-specific data and commonly used data selectively rather than statically loading all data, optimizing both processing capability and response time.

Inventive Principle:
Principle #15Dynamics

3Quantity of substance

If only trigger input recognition data is loaded into volatile memory, then memory load is reduced, but speech recognition and natural language understanding processing is insufficient

Engineering Contradiction:
Improvememory loadVSAvoidspeech processing capability
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent segments data into trigger input recognition data, domain-specific data, and commonly used data. The system loads trigger input recognition data and commonly used data into volatile memory to handle basic speech processing, while domain-specific data can be loaded on-demand based on the actual speech content, balancing memory load with processing capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by loading only the essential data segments (trigger input recognition data and commonly used data) into volatile memory for basic speech processing operations. Domain-specific data is loaded only when needed based on the speech content, avoiding the excessive action of loading all possible data while maintaining sufficient processing capability for the current task.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11967325B2Electronic device and method for controlling the electronic device
Publication Date: 2024.04.23 SAMSUNG ELECTRONICS CO LTD
  • US11967325B2 patent drawing
  • US11967325B2 patent drawing
  • US11967325B2 patent drawing

AI summary

Disclosed are an electronic device capable of efficiently performing speech recognition and natural language understanding and a method for controlling thereof. The electronic device includes: a microphone; a non-volatile memory configured to store virtual assistant model data comprising data that is classified according to a plurality of domains and data that is commonly used for the plurality of domains; a volatile memory; and a processor configured to: based on receiving, through the microphone, a trigger input to perform speech recognition for a user speech, initiate loading the virtual assistant model data from the non-volatile memory into the volatile memory, load, into the volatile memory, first data from among the data classified according to the plurality of domains and, while loading the first data into the volatile memory, load at least a part of the data commonly used for the plurality of domains into the volatile memory.