语音处理方法、装置、电子设备和可读存储介质

By determining the starting endpoint using a speech endpoint detection model without back view and combining it with a model with back view to determine the ending endpoint, the latency problem caused by deep neural networks is solved, achieving efficient endpoint detection and a user experience with low latency.

CN116030796BActive Publication Date: 2026-07-17IFLYTEK CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2022-12-29
Publication Date
2026-07-17

Smart Images

  • Figure CN116030796B_ABST
    Figure CN116030796B_ABST
Patent Text Reader

Abstract

本发明提供一种语音处理方法、装置、电子设备和可读存储介质,涉及语音处理技术领域,所述方法包括:将待检测语音信号对应的原始语音特征输入至不带后视野的语音端点检测模型中,得到原始语音特征中的人声语音特征对应的目标起始端点;将待检测语音信号对应的原始语音特征输入至带后视野的语音端点检测模型中,得到原始语音特征中的人声语音特征对应的目标终止端点;基于目标起始端点和目标终止端点,从原始语音特征中截取待检测语音信号对应的目标人声语音特征,以解决现有技术中无法兼顾提高端点检测效果以及降低用户使用的延迟感的技术问题。
Need to check novelty before this filing date? Find Prior Art