This invention discloses a heterogeneous
hybrid expert
large model architecture and alignment method for high-
concurrency online medical dialogue. First, the method proposes a dual-
pool heterogeneous routing strategy of "4 general experts + 2 departmental experts." Through physically isolated
parallel routing channels, it forces the model to simultaneously activate general language capabilities and specialized medical reasoning when
processing input, fundamentally ensuring the focus and professionalism of responses. Second, it adopts a
structural evolution strategy based on parameter expansion and selective freezing. By completely freezing the pre-trained basic parameters and training only newly added departmental experts, it completely eliminates catastrophic forgetting. Finally, in the model alignment stage, it employs the Group Relative Policy Optimization (GRPO)
algorithm for valueless models to reduce memory overhead and designs a dual-track
hybrid reward function combining low-level
semantic similarity rewards and high-level structured
large model referee scoring, effectively solving the "reward hacking" problem and guiding the model to generate refined responses that combine clinical accuracy and humanized interaction. This invention significantly improves the performance of medical dialogue systems. Its sparse activation and parameter freezing characteristics further ensure low latency and high
throughput in online high-
concurrency scenarios, providing a complete solution for the deployment of reliable medical dialogue systems.