A dialect medical question and answer method based on multi-modal base model adaptation

By training with a hybrid data of lightweight speech models and large language models, combined with rule-based dialogue management, the low recognition accuracy and deployment challenges of low-resource dialects in intelligent medical question-and-answer systems have been solved, enabling safe and professional multi-dialect medical consultation services and expanding the service coverage population.

CN122436268APending Publication Date: 2026-07-21NEUSOFT INST GUANGDONG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NEUSOFT INST GUANGDONG
Filing Date
2026-04-08
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing intelligent medical question-answering systems have low recognition accuracy when faced with dialects with limited resources, high deployment thresholds, and are difficult to implement in edge computing scenarios. Furthermore, they lack a collaborative mechanism between rule-based dialogue management and a large model knowledge base, making it impossible to provide secure, professional, and concise multi-dialect medical consultation services.

Method used

We employ a lightweight speech model for low-resource dialect speech signal processing, combined with hybrid data training of a large language model and rule-based dialogue management, to construct a dialect medical question-answering method based on a multimodal base model. Through intent parsing, confidence scoring, and security review, we achieve a complete closed-loop interaction from dialect speech input to output.

Benefits of technology

It improves the accuracy of speech recognition for low-resource dialects, provides natural and personalized intelligent medical consultation services, ensures the security and professionalism of medical content, and lowers the deployment threshold, expanding the coverage of intelligent medical services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122436268A_ABST
    Figure CN122436268A_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence and medical information processing technology, in particular to a dialect medical question and answer method based on multi-modal base model adaptation, which comprises converting original dialect signals into recognized text through a voice model enhanced by low-resource data and optimized parameters; then analyzing the intent and entity, and generating candidate replies and their confidence scores based on the dialogue state; through comparing the scores with the preset threshold, realizing double-track intelligent switching. When the bottom link is triggered, the system retrieves knowledge fragments from the medical knowledge base combined with the context, and injects them into a large language model fine-tuned by medical reasoning chain and consultation mixed data to generate text replies. Finally, the replies after security review are converted into matching dialect voice output. The present application effectively solves the shortcomings of traditional medical question and answer systems in dialect recognition, complex medical reasoning and long-tail consultation processing, and significantly improves the professionalism and controllability of medical advice through RAG retrieval enhancement and double-track collaborative strategy.
Need to check novelty before this filing date? Find Prior Art