Audio Processing Method for Continuous Voice Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice interaction systems, such as voice assistants, face complications in smooth conversation flow due to the need to repeatedly wake up applications for processing subsequent audio data, leading to increased errors in responding to isolated audio data segments.

Innovation Solution

An audio processing method that combines first and second audio data to obtain target audio data without re-waking the target application, allowing for continuous conversation processing and improved accuracy in understanding user needs by analyzing semantic content and determining the completeness of input audio data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the application is woken up for each audio data processing, then the audio data can be processed, but the conversation flow becomes complicated and not smooth

Engineering Contradiction:
Improveaudio data processing efficiencyVSAvoidconversation flow complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The target application is woken up in advance and remains active to continuously acquire audio data. This preliminary action eliminates the need to repeatedly wake up the application for each audio segment, thereby simplifying the conversation flow while maintaining processing efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The application maintains continuous audio data acquisition and processing without interruption. By keeping the application running and continuously processing audio streams, the system achieves smooth conversation flow while efficiently handling multiple audio segments.

Inventive Principle:
Principle #20Continuity of useful action

2Ease of operation

If audio data are processed separately, then each audio segment can be handled individually, but the responding accuracy decreases due to isolated processing errors

Engineering Contradiction:
Improveaudio data processing simplicityVSAvoidaudio response accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

Multiple audio data segments are merged and processed together as a unified input. The system combines first audio data, second audio data, and any additional audio segments into a single target audio data stream, which is then processed collectively to improve recognition accuracy while maintaining operational simplicity.

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If the application re-wakes for continuous conversation, then new audio data can be processed, but the interaction smoothness deteriorates

Engineering Contradiction:
Improveaudio data processing capabilityVSAvoidvoice interaction smoothness
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The application is preliminarily activated and maintains its processing state continuously. This preliminary action enables the system to handle subsequent audio segments without re-waking, thereby maintaining voice interaction smoothness while preserving full processing capability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The audio processing function operates continuously without interruption or re-initialization. The application maintains its useful action of processing audio data throughout the conversation, ensuring both productivity and interaction smoothness.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentEP4184506A1Audio processing
Publication Date: 2023.05.24 BEIJING XIAOMI MOBILE SOFTWARE CO LTD
  • EP4184506A1 patent drawingFigure 1
  • EP4184506A1 patent drawingFigure 2
  • EP4184506A1 patent drawingFigure 3

AI summary

Provided are an audio processing method and apparatus, and a storage medium. The method includes: acquiring first audio data associated with a first audio signal after waking up a target application; when second audio data associated with a second audio signal is detected in the process of acquiring the first audio data, acquiring the second audio data; and obtaining target audio data according to the first audio data and the second audio data. According to the technical solution of the present disclosure, a conversation flow can be simplified without waking up a target application again, the first audio data and the second audio data are combined to obtain target audio data, and audio response is made to the target audio data, which can more accurately get real needs of a user, reduce the rate of isolated responding errors, and improve the accuracy of the audio response.