语音端点检测方法及相关装置、设备和介质

By employing a four-class state machine and refining noise types in speech endpoint detection, the problem of semantic fragmentation in speech endpoint detection is solved, thereby improving the semantic coherence of speech detection and the accuracy of downstream tasks.

CN121306102BActive Publication Date: 2026-07-17ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD
Filing Date
2025-10-13
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing speech endpoint detection technologies lack the ability to model the overall semantic level, leading to misjudgments of sentence termination during short pauses, causing semantic fragmentation and affecting the execution of downstream tasks.

Method used

A four-class state machine based on streaming audio is used for continuous monitoring. By refining the noise type into first noise, middle noise and last noise, and combining the energy value and probability value of the audio frame, the four-class state machine is used for continuous monitoring to determine the semantic position of the speech frame.

Benefits of technology

It improves the semantic coherence after speech endpoint detection, reduces incomplete sentence segmentation, and improves the accuracy of downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306102B_ABST
    Figure CN121306102B_ABST
Patent Text Reader

Abstract

本申请公开了一种语音端点检测方法及相关装置、设备和介质,其中语音端点检测方法包括:基于流式音频进行持续预测,得到流式音频中音频帧分别属于若干帧类型的概率值;其中,若干帧类型包括人声、首噪声、中间噪声、尾噪声;基于音频帧分别属于若干帧类型的概率值和音频帧的能量值,在若干帧类型中确定音频帧所属的目标类型;基于四分类状态机对流式音频中各个音频帧的目标类型进行持续监测,得到当前帧的判断结果;其中,判断结果包括:当前帧处是否整句结束、当前帧处是否子句结束中至少一者,且四分类状态机中相邻状态对应于任意两个帧类型。上述方案,能够提升执行语音端点检测之后的语义连贯性。
Need to check novelty before this filing date? Find Prior Art