Voice Service System Adapting Response to Speaking Speed

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice service technologies fail to accurately match voice responses with user demands due to ignoring the speed variations of speakers, leading to a poor match between the provided service and user needs.

Innovation Solution

A method and apparatus that analyze the time-domain waveform of voice input signals to determine current speaking speed, compare it with a user's standard speed information set, and generate a voice response signal based on the matched demand information, improving the match between voice service and user demands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice service technology converts voice signals into characters and analyzes them to determine response strategies, then the basic voice recognition function is achieved, but the potential demand information contained in different speaking speeds is ignored, leading to poor match between provided service and user demand

Engineering Contradiction:
Improvedemand information recognition accuracyVSAvoidservice adaptation to user characteristics
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the voice signal analysis into multiple dimensions: converting voice to text for semantic analysis, simultaneously extracting speaking speed from time-domain waveforms, and analyzing frequency-domain characteristics. This segmentation allows each aspect to be processed independently and then integrated to form a comprehensive understanding of user demand, resolving the contradiction between basic recognition accuracy and adaptability to user characteristics.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the analysis parameters by not only converting voice to text but also extracting temporal parameters (speaking speed, pauses) and frequency parameters (pitch, tone) from the voice signal. By incorporating these additional parameters into the demand analysis, the system achieves both precise demand recognition and adaptability to individual user characteristics such as emotional state and speaking habits.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If the system analyzes only the text content of voice signals, then the processing complexity is low, but the speaking speed information and emotional state information are lost

Engineering Contradiction:
Improvesignal processing complexityVSAvoidspeed and emotional information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent merges multiple analysis approaches by combining text-based semantic analysis with waveform-based temporal and frequency analysis. The system processes voice signals through parallel pathways: one for text conversion and semantic understanding, another for time-domain waveform analysis to extract speaking speed, and a third for frequency-domain analysis to detect emotional tone. These merged results are then integrated to comprehensively determine user demand while maintaining manageable processing complexity through modular architecture.

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If the voice service provides standardized responses, then the service delivery is efficient, but the responses do not adapt to different user emotional states and speaking speeds

Engineering Contradiction:
Improveservice response efficiencyVSAvoidresponse adaptation to user state
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic response adaptation by making the response generation process adjustable based on real-time analysis of user speaking speed and emotional state. The system dynamically selects from multiple response strategies: for fast speakers it may provide more concise responses, for slow speakers more detailed explanations; for excited emotional states it may use more energetic tone and shorter responses; for calm states it may provide more comprehensive information. This dynamic adaptation maintains service efficiency while significantly improving adaptability to user characteristics.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms where the analyzed user characteristics (speaking speed, emotional state) feed into the response generation process. The response strategy is continuously adjusted based on this feedback loop, allowing the system to learn and adapt to individual user preferences over time while maintaining efficient service delivery through standardized response templates that can be dynamically modified.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10403282B2Method and apparatus for providing voice service
Publication Date: 2019.09.03 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US10403282B2 patent drawing
  • US10403282B2 patent drawing
  • US10403282B2 patent drawing

AI summary

The present disclosure discloses a method and apparatus for providing a voice service. A specific implementation of the method for providing a voice service comprises: acquiring a voice input signal; analyzing the time-domain waveform of the voice input signal to determine current speed information of the voice input signal; comparing the current speed information with an acquired standard speed information set of a user outputting the voice input signal, and determining first demand information from a preset demand information set according to the comparison result, wherein the standard speed information set comprises at least one piece of standard speed information, and the preset demand information set comprises demand information corresponding to each piece of standard speed information in the standard speed information set; and generating a voice response signal based on the first demand information and the second demand information acquired by analyzing the voice input signal.