说话对象的识别方法及装置、电子设备和存储介质

By fusing feature information from multimodal data, the problems of high requirements and low recognition efficiency in speech object recognition are solved, achieving efficient recognition without registration and wide application, and improving recognition accuracy and user experience.

CN116386645BActive Publication Date: 2026-07-17MOORE THREADS TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MOORE THREADS TECH CO LTD
Filing Date
2023-03-07
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies for speech object recognition suffer from problems such as high usage requirements, limited application scenarios, poor recognition results, and low recognition efficiency.

Method used

By identifying the feature information corresponding to each modality in the multimodal data containing audio, the feature information is fused to perform speaker recognition. This includes feature extraction and similarity matrix calculation of multiple modal information such as audio, video, and audio acquisition device information, and recognition is performed using the fused feature information.

Benefits of technology

It enables the expansion of application scenarios, improves recognition accuracy and efficiency, reduces the possibility of recognition errors, and enhances user experience without requiring prior registration of the speaking subject.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386645B_ABST
    Figure CN116386645B_ABST
Patent Text Reader

Abstract

本公开实施例公开了一种说话对象的识别方法及装置、电子设备和存储介质,该方法包括:确定包含音频的第一多模态数据中每一模态信息对应的特征信息;其中,所述第一多模态数据中具有至少两种模态信息;基于每一所述模态信息对应的特征信息,确定融合特征信息;利用所述融合特征信息,对所述第一多模态数据中的说话对象进行识别,得到第一识别结果。
Need to check novelty before this filing date? Find Prior Art