说话对象的识别方法及装置、电子设备和存储介质
By fusing feature information from multimodal data, the problems of high requirements and low recognition efficiency in speech object recognition are solved, achieving efficient recognition without registration and wide application, and improving recognition accuracy and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MOORE THREADS TECH CO LTD
- Filing Date
- 2023-03-07
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies for speech object recognition suffer from problems such as high usage requirements, limited application scenarios, poor recognition results, and low recognition efficiency.
By identifying the feature information corresponding to each modality in the multimodal data containing audio, the feature information is fused to perform speaker recognition. This includes feature extraction and similarity matrix calculation of multiple modal information such as audio, video, and audio acquisition device information, and recognition is performed using the fused feature information.
It enables the expansion of application scenarios, improves recognition accuracy and efficiency, reduces the possibility of recognition errors, and enhances user experience without requiring prior registration of the speaking subject.
Smart Images

Figure CN116386645B_ABST