Guangzhou-hybrid speech recognition method and device, computer equipment and storage medium

By constructing a mixed speech training set and utilizing a deep neural network model, the challenges of Mandarin and Cantonese speech recognition were solved, achieving accurate speech recognition in the medical field and improving the effectiveness of patient follow-up.

CN116564314BActive Publication Date: 2026-05-29PING AN TECH (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2023-05-31
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing speech recognition technology has difficulty distinguishing between Mandarin and Cantonese, which leads to the inability to accurately recognize Cantonese speech during patient follow-ups in the medical field, affecting the effectiveness of disease monitoring and follow-up.

Method used

By acquiring training sets of Mandarin and Cantonese speech, adding identifiers to Cantonese text data, constructing a mixed speech training set, and training it using a deep neural network model, a mixed speech recognition model is generated that can distinguish between Mandarin and Cantonese speech.

Benefits of technology

It improves the accuracy of speech recognition, enabling accurate recognition of Cantonese speech during telephone follow-ups, thereby enhancing the effectiveness of disease monitoring and improving the patient experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116564314B_ABST
    Figure CN116564314B_ABST
Patent Text Reader

Abstract

The embodiment of the application belongs to the technical field of artificial intelligence, is applied to the field of digital medical treatment, relates to a Cantonese-Mandarin mixed speech recognition method, and comprises the following steps: adding an identifier to Cantonese text in obtained Cantonese text data to obtain extended Cantonese text data; combining obtained Mandarin audio data, Mandarin text data, Cantonese audio data and the extended Cantonese text data into a mixed speech training set, inputting the mixed speech training set into a pre-constructed deep neural network model for training to obtain a mixed speech recognition model; obtaining to-be-recognized speech data, inputting the to-be-recognized speech data into the mixed speech recognition model, and outputting a speech recognition result. The application also provides a Cantonese-Mandarin mixed speech recognition device, a computer device and a storage medium. In addition, the application also relates to the technology of blockchains, and the mixed speech training set can be stored in a blockchain. The application can distinguish whether the currently recognized speech type is Cantonese or Mandarin.
Need to check novelty before this filing date? Find Prior Art