Dynamic Audio Type Recognition Thresholds Using Continuous Frames

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio recognition methods using constant thresholds suffer from poor recognition accuracy.

Innovation Solution

An audio processing method that dynamically adjusts recognition thresholds based on characteristic information of continuous audio frames, using a threshold adjustment condition to improve the accuracy of audio type recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a constant threshold is used for audio recognition, then the recognition process is simple and fast, but the recognition accuracy is poor

Engineering Contradiction:
Improverecognition accuracyVSAvoidthreshold adjustment mechanism
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies dynamics by transforming the static constant threshold into a dynamic adaptive threshold that automatically adjusts based on audio frame characteristics. The threshold is no longer fixed but evolves with the input signal properties, allowing the system to maintain high recognition accuracy across varying audio conditions without manual intervention.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements feedback by using the characteristics of continuously recognized audio frames to adjust the recognition threshold. The system monitors recognition results and audio frame properties, then feeds this information back to modify the threshold, creating a closed-loop control system that continuously optimizes recognition accuracy.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If the recognition threshold is adjusted dynamically based on audio frame characteristics, then the recognition accuracy is improved, but the computational complexity increases

Engineering Contradiction:
Improveaudio type recognition accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies parameter changes by modifying the recognition threshold parameter based on audio frame characteristics such as energy, zero-crossing rate, or spectral features. Instead of changing the entire recognition system, only the threshold parameter is adjusted, which reduces computational overhead while maintaining accuracy improvements.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If continuous audio frames are analyzed for threshold adjustment, then the recognition reliability is improved, but the processing time increases

Engineering Contradiction:
Improverecognition reliabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by analyzing only the necessary characteristics of continuous audio frames rather than processing the entire audio signal. The system selectively extracts relevant features (such as energy levels, frequency components, or temporal patterns) to adjust the threshold, achieving improved reliability without proportionally increasing processing time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250252969A1Audio processing method and apparatus, storage medium, and electronic device
Publication Date: 2025.08.07 DOUYIN VISION CO LTD
  • US20250252969A1 patent drawing
  • US20250252969A1 patent drawing
  • US20250252969A1 patent drawing

AI summary

An audio processing method and apparatus, a storage medium, and an electronic device. The audio processing method includes: acquiring an audio frame to be processed and determining an audio type of the audio frame based on a current recognition threshold; in response to a determination that a current audio frame satisfies a threshold adjustment condition, determining a determination state of a recognized audio type based on characteristic information of recognized continuous audio frames; and adjusting the current recognition threshold according to the determination state, wherein the adjusted recognition threshold is configured to recognize the audio type of a next audio frame.