基于多库联动的领域垂类大语言模型微调数据集构建方法
By using a multi-database linkage approach, the problems of knowledge silos, multi-source data conflicts, and low sample quality in the construction of fine-tuning datasets for large language models in specific domains were solved, achieving efficient and professional dataset construction and improving the knowledge consistency and sample quality of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN SHAOFENG INST OF APPLIED MATHEMATICS
- Filing Date
- 2026-05-07
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies for constructing large language model fine-tuning datasets for specific domains suffer from several problems, including knowledge silos, non-standardized prompting engineering, potential conflicts between multiple data sources, lack of domain-specific priority configuration, and low quality of fine-tuning samples. These issues result in incomplete model knowledge, insufficient accuracy, and limited practicality.
By employing a multi-database linkage approach, a dynamic priority mechanism and a four-tuple instruction system are established to generate structured instruction question-and-answer samples through the association mechanism of structured business data, unstructured text data, and knowledge relationship graph data. Combined with a large language model, diverse rewriting and quality filtering are performed to construct a high-quality fine-tuning dataset.
It has achieved the organic integration of multi-source data, solved the problem of inconsistency between historical data and current business rules, significantly improved sample quality and generation efficiency, and enhanced the knowledge consistency and professionalism of the model.
Smart Images

Figure CN122133822B_ABST