语音关键词检索方法、装置、设备及存储介质

By extracting features and performing cluster analysis on audio samples of keywords, combined with a support vector regression model, the number of states and the verification of keyword segments are automatically determined. This solves the problem of phoneme annotation dependency in existing technologies and achieves high-accuracy keyword retrieval across languages ​​and dialects.

CN122132592BActive Publication Date: 2026-07-17GUANGZHOU DAYIN ZHIYUAN DIGITAL TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU DAYIN ZHIYUAN DIGITAL TECH CO LTD
Filing Date
2026-05-06
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing speech keyword retrieval methods rely on phoneme annotation, which limits their applicability and makes it difficult to meet the actual needs of cross-language and cross-dialect scenarios.

Method used

By performing feature extraction and cluster analysis on user-uploaded keyword sample audio, the number of keyword states is automatically determined, a network model containing keyword state sequences and non-keyword absorption states is established, and a support vector regression model is used for secondary verification to determine whether the suspected keyword fragments are training keywords.

Benefits of technology

A network model can be built without manual annotation of phoneme information, and it is applicable to keyword retrieval of any language, dialect, and even non-verbal sounds, significantly expanding the scope of application and improving the accuracy of retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132592B_ABST
    Figure CN122132592B_ABST
Patent Text Reader

Abstract

本发明提供了一种语音关键词检索方法、装置、设备及存储介质,该方法包括:对N个关键词样本音频的标注区间进行特征提取和聚类分析,根据聚类结果确定关键词对应的状态数量,建立包含关键词状态序列和非关键词吸收状态的网络模型;对待检索音频进行特征提取,将提取的特征输入网络模型中进行解码,根据解码路径中关键词状态序列的概率值确定疑似关键词片段;对疑似关键词片段进行特征提取并输入预先训练的支持向量回归模型,根据支持向量回归模型输出的分类得分判断疑似关键词片段是否为训练的关键词。本发明通过对关键词样本的特征进行时序聚类分析自动确定状态数量,能够适用于任何语言甚至非语言声音的关键词检索,扩展适用范围检索的准确率。
Need to check novelty before this filing date? Find Prior Art