Near-End Voice Interaction Module Using Sound Energy Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice interaction methods using AI voice assistants face delays and judgment errors due to the time required for signals to travel between near-end electronic devices and remote servers, causing prolonged processing times.

Innovation Solution

A near-end electronic device with a microphone, transmission module, near-end processing module, and sound energy calculation module that determines when a voice input has ended by analyzing sound energy values, allowing for early termination of recording and transmission, thereby reducing processing time and errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the remote artificial intelligence server determines whether the voice has ended, then the judgment accuracy is improved, but the processing time is prolonged due to signal travel time

Engineering Contradiction:
Improvevoice end detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The voice end detection function is segmented between the near-end electronic device (which performs initial detection using sound energy calculation) and the remote artificial intelligence server (which performs final confirmation). This segmentation allows the near-end device to make preliminary judgments locally, reducing the time spent waiting for remote server responses while maintaining overall detection accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The near-end electronic device performs preliminary voice end detection using sound energy calculation before the remote server completes its analysis. This preliminary action enables the device to stop recording and transmission earlier, reducing processing time while the remote server subsequently validates the decision to maintain accuracy.

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If the near-end device stops transmitting voice data earlier, then the processing time is reduced, but the judgment accuracy may deteriorate

Engineering Contradiction:
Improveprocessing timeVSAvoidvoice end detection accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The remote artificial intelligence server provides feedback to the near-end electronic device regarding the accuracy of voice end detection. The server analyzes the complete voice data and confirms or corrects the near-end device's preliminary judgment, ensuring high accuracy while allowing the near-end device to operate with reduced processing time.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach speeds up the voice interaction process by allowing the near-end device to stop transmitting voice data earlier, reducing processing time and minimizing judgment errors, while the remote server still confirms the completeness of sentences, thus enhancing the efficiency of AI voice interactions.

Implementation Method 1

near-end processing module and a sound energy calculation module that determines when a voice input has ended by analyzing sound energy values

Methodology Applied
Scientific EffectSound energy analysis: Acoustics

Data Source

PatentUS11244697B2Artificial intelligence voice interaction method, computer program product, and near-end electronic device thereof
Publication Date: 2022.02.08 AIROHA TECHNOLOGY CORPORATION
  • US11244697B2 patent drawing
  • US11244697B2 patent drawing
  • US11244697B2 patent drawing

AI summary

An artificial intelligence voice interaction method and a near-end electronic device thereof are disclosed. The method includes the following steps: receiving a voice input by a user; transmitting the voice to a remote artificial intelligence server; determining whether the voice has ended; when determining that the voice has ended and has not received a stop recording signal transmitted by the remote artificial intelligence server, it stops transmitting the voice to the remote artificial intelligence server; before determining that the voice has ended, and has received the stop recording signal from the remote artificial intelligence server, it stops transmitting the voice to the remote artificial intelligence server; and receiving a response signal send back from the remote artificial intelligence server.