基于多任务方言识别的智能设备语音交互方法及系统

By combining deep learning and LSTM models with geolocation information, the problem of dialect recognition in complex acoustic environments for smart wearable devices was solved, achieving accurate dialect recognition and device control, and improving the flexibility and personalization of user interaction.

CN120766677BActive Publication Date: 2026-07-17CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
Filing Date
2025-07-31
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing smart wearable devices lack dialect recognition accuracy in complex acoustic environments, failing to effectively separate dialect features from the spectral overlap of environmental noise, leading to incorrect recognition of key commands. Furthermore, they cannot construct a closed-loop decision-making logic between dialect regions and user behavior, resulting in poor comprehension capabilities.

Method used

By combining geolocation information with a deep learning framework, a probabilistic correlation model between geolocation and dialect distribution is constructed, generating a fusion feature tensor of acoustic and spatial features. The LSTM model is used to capture long-term temporal dependencies, and adaptive interactive decisions are output through a multi-level decision tree to generate voice response data.

Benefits of technology

Achieve accurate dialect recognition and device control in complex speech environments, improve recognition accuracy, increase flexibility and personalization, reduce false triggering rate, and enhance user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766677B_ABST
    Figure CN120766677B_ABST
Patent Text Reader

Abstract

本发明涉及基于多任务方言识别的智能设备语音交互方法及系统,该方法包括:获取来自用户的语音数据,对语音数据进行预处理,以提取语音数据中的声学特征。通过深度学习框架对用户的地理位置信息和声学特征进行耦合,以构建地理位置与方言分布的概率关联模型,调用概率关联模型生成声学特征与空间特征的融合特征张量。调用LSTM模型基于融合特征张量捕获语音数据中的长时序依赖关系,建立参数共享机制,对概率关联模型进行微调,以输出方言识别结果。将方言识别结果转换为设备控制指令,并响应于设备控制指令调用多级决策树输出自适应交互决策。基于自适应交互决策,结合声学特征生成对应方言的语音回复数据,将语音回复数据反馈至用户终端。
Need to check novelty before this filing date? Find Prior Art