语音关键词检索方法、装置、设备及存储介质
By extracting features and performing cluster analysis on audio samples of keywords, combined with a support vector regression model, the number of states and the verification of keyword segments are automatically determined. This solves the problem of phoneme annotation dependency in existing technologies and achieves high-accuracy keyword retrieval across languages and dialects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU DAYIN ZHIYUAN DIGITAL TECH CO LTD
- Filing Date
- 2026-05-06
- Publication Date
- 2026-07-17
AI Technical Summary
Existing speech keyword retrieval methods rely on phoneme annotation, which limits their applicability and makes it difficult to meet the actual needs of cross-language and cross-dialect scenarios.
By performing feature extraction and cluster analysis on user-uploaded keyword sample audio, the number of keyword states is automatically determined, a network model containing keyword state sequences and non-keyword absorption states is established, and a support vector regression model is used for secondary verification to determine whether the suspected keyword fragments are training keywords.
A network model can be built without manual annotation of phoneme information, and it is applicable to keyword retrieval of any language, dialect, and even non-verbal sounds, significantly expanding the scope of application and improving the accuracy of retrieval.
Smart Images

Figure CN122132592B_ABST