Dynamic Speech Section Detection Level Adjustment for Voice Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice interaction systems face inefficiencies in noise reduction and power consumption when performing voice recognition in noisy environments, as existing speech section detection methods fail to accurately distinguish speech from noise, leading to increased communication costs and power usage.

Innovation Solution

A control apparatus that dynamically adjusts the identification level of a speech section detector based on noise levels and distance, lowering the detection level when it's likely that a target person is speaking to improve accuracy while minimizing unnecessary data transmission and power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If speech section detection is performed with a high identification level to reduce noise transmission, then communication cost and power consumption are reduced, but voice recognition accuracy deteriorates due to false identification of speech as noise

Engineering Contradiction:
Improvepower consumptionVSAvoidvoice recognition accuracy
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The identification level of the speech section detector is dynamically adjusted based on the operational state of the voice interaction apparatus. When voice recognition is actively performed, the identification level is lowered to improve speech detection accuracy; when voice recognition is not active, the identification level is raised to reduce noise transmission and save power. This dynamic adjustment resolves the contradiction between power efficiency and recognition accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the identification level parameter of the speech section detector according to different operational conditions. By adjusting this parameter based on whether voice recognition is being performed, the system optimizes both power consumption and voice recognition accuracy for different states of operation.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If speech section detection is performed with a low identification level to improve speech detection accuracy, then voice recognition accuracy is improved, but communication cost and power consumption increase due to transmission of noise data

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The identification level is dynamically adjusted based on operational needs. During active voice recognition, the level is lowered to capture all speech; during inactive periods, the level is raised to filter noise and conserve power, thus resolving the contradiction between detection sensitivity and energy efficiency.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system modifies the identification level parameter according to the operational state, lowering it during voice recognition activities to improve accuracy and raising it during idle periods to reduce power consumption and communication overhead.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If constant communication with voice recognition server is performed to ensure accurate voice recognition, then voice recognition accuracy is improved, but communication cost and power consumption are wastefully increased during non-speech periods

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidcommunication data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system extracts and transmits only the relevant speech sections to the voice recognition server by using the speech section detector to identify and separate actual speech from noise. This selective transmission reduces the volume of communication data while maintaining voice recognition accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of constant communication, the system performs periodic speech section detection and transmits data only during periods when speech is detected. This periodic approach reduces communication frequency and data volume during non-speech periods while ensuring accurate voice recognition when needed.

Inventive Principle:
Principle #19Periodic action

4Quantity of substance

If speech section detection is performed to reduce data transmission, then communication cost is reduced, but voice recognition accuracy deteriorates due to insufficient or false speech detection

Engineering Contradiction:
Improvecommunication data volumeVSAvoidvoice recognition accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The speech section detector extracts actual speech from the mixed audio signal by comparing against noise profiles. This extraction process identifies genuine speech sections for transmission while filtering out noise, thus reducing communication data volume without compromising voice recognition accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses feedback from the voice recognition server to adjust the speech section detection parameters. When the server indicates successful recognition, the system continues current detection settings; when recognition fails, the system adjusts detection sensitivity to improve future speech identification accuracy.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11081114B2Control method, voice interaction apparatus, voice recognition server, non-transitory storage medium, and control system
Publication Date: 2021.08.03 TOYOTA JIDOSHA KK
  • US11081114B2 patent drawing
  • US11081114B2 patent drawing
  • US11081114B2 patent drawing

AI summary

The control apparatus includes: a calculation unit configured to control a voice interaction apparatus including a speech section detector, the speech section detector being configured to identify whether an acquired voice includes a speech made by a target person by a set identification level and perform speech section detection, in which the calculation unit instructs, when an estimation result indicating that it is highly likely that the speech made by the target person is included in the acquired voice has been acquired from a voice recognition server, the voice interaction apparatus to change a setting in such a way as to lower the identification level of the speech section detector, and to perform communication with the voice recognition server by speech section detection in accordance with the identification level after the change.