Local Speech Recognition Response Control Unit

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice interaction systems face delays in response time due to network access, general-purpose speech recognition supporting various utterances, and increased processing loads on servers, leading to reduced user convenience.

Innovation Solution

An information processing device and method that utilize a response control unit to manage user responses based on both first and second utterance interpretation results, where the second result is acquired through learning data associated with the first result, allowing for faster processing and reduced response time by performing local-side voice interaction tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition processing is performed by both client and server to improve accuracy, then recognition accuracy is improved, but response time increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary speech recognition processing on the client side before server processing. The client generates a first recognition result locally, which is then used by the server to generate a second recognition result. This preliminary local processing reduces the overall response time while maintaining accuracy through subsequent server-side refinement.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The speech recognition system is divided into two segments: client-side processing and server-side processing. The client handles initial recognition and generates first recognition results, while the server performs additional processing to generate second recognition results. This segmentation allows parallel processing and reduces the total time required for accurate recognition.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If general-purpose speech recognition is used to support various utterances, then adaptability is improved, but processing load on server increases

Engineering Contradiction:
Improveutterance support capabilityVSAvoidserver processing load
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system performs partial speech recognition processing on the client side rather than relying entirely on server processing. The client generates first recognition results locally, reducing the amount of processing required on the server while still supporting various utterances through the combined client-server approach.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If network access is required for speech processing, then recognition accuracy is improved, but response time increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidnetwork access time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary speech recognition processing on the client side before server processing. The client generates a first recognition result locally, which is then used by the server to generate a second recognition result. This preliminary local processing reduces the overall response time while maintaining accuracy through subsequent server-side refinement.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11948564B2Information processing device and information processing method
Publication Date: 2024.04.02 SONY GROUP CORP
  • US11948564B2 patent drawing
  • US11948564B2 patent drawing
  • US11948564B2 patent drawing

AI summary

Provided is an information processing device including a response control unit that controls a response to a user's utterance based on a first utterance interpretation result and a second utterance interpretation result. The first utterance interpretation result is a result of natural language understanding processing for an utterance text generated by automatic speech recognition processing based on the user's utterance and the second utterance interpretation result is an interpretation result acquired based on learning data in which the first utterance interpretation result and the utterance text used to acquire the first utterance interpretation result are associated with each other. The response control unit further controls the response to the user's utterance based on the second utterance interpretation result in a case where the second utterance interpretation result is acquired based on the user's utterance before acquisition of the first utterance interpretation result.