基于多库联动的领域垂类大语言模型微调数据集构建方法

By using a multi-database linkage approach, the problems of knowledge silos, multi-source data conflicts, and low sample quality in the construction of fine-tuning datasets for large language models in specific domains were solved, achieving efficient and professional dataset construction and improving the knowledge consistency and sample quality of the model.

CN122133822BActive Publication Date: 2026-07-17HUNAN SHAOFENG INST OF APPLIED MATHEMATICS

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN SHAOFENG INST OF APPLIED MATHEMATICS
Filing Date
2026-05-07
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies for constructing large language model fine-tuning datasets for specific domains suffer from several problems, including knowledge silos, non-standardized prompting engineering, potential conflicts between multiple data sources, lack of domain-specific priority configuration, and low quality of fine-tuning samples. These issues result in incomplete model knowledge, insufficient accuracy, and limited practicality.

Method used

By employing a multi-database linkage approach, a dynamic priority mechanism and a four-tuple instruction system are established to generate structured instruction question-and-answer samples through the association mechanism of structured business data, unstructured text data, and knowledge relationship graph data. Combined with a large language model, diverse rewriting and quality filtering are performed to construct a high-quality fine-tuning dataset.

Benefits of technology

It has achieved the organic integration of multi-source data, solved the problem of inconsistency between historical data and current business rules, significantly improved sample quality and generation efficiency, and enhanced the knowledge consistency and professionalism of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133822B_ABST
    Figure CN122133822B_ABST
Patent Text Reader

Abstract

本发明公开了基于多库联动的领域垂类大语言模型微调数据集构建方法,包括以下步骤:采集结构化业务数据、非结构化文本数据、知识关系图数据,分别构成结构化业务数据库、非结构化文本库和知识关系图数据库;建立多库联动关联机制;建立动态优先级机制,当不同数据源冲突时,按照优先级顺序进行冲突消解;基于多库联动关联结果,采用四元组指令生成结构化指令问答样本;使用大语言模型进行多样化改写和质量过滤,从而构建领域垂类大语言模型微调数据集。本发明采用数据优先级进行冲突消解有效解决了历史数据与现行业务规则不一致的技术问题;采用四元组指令生成结构化指令问答样本,显著提升了样本质量和生成效率。
Need to check novelty before this filing date? Find Prior Art