An industrial large model unstructured data governance and adaptive fine-tuning method

By acquiring multi-source unstructured data from industrial sites, performing multimodal parsing and semantic standardization, and combining industrial knowledge graphs for hardware-aware quantitative optimization, the challenges of large industrial models in unstructured data processing and fine-tuning are solved. This enables efficient and automated model building and adaptive iteration, making it suitable for various intelligent industrial applications.

CN122412879APending Publication Date: 2026-07-17BEIJING PULIAN JICHUANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In the industrial sector, the preprocessing of unstructured data makes it difficult to effectively extract semantically consistent structured information. Furthermore, the general large model faces a contradiction between computing power adaptability and professional reliability during fine-tuning, resulting in unstable inference in root cause analysis and process optimization tasks. This makes it difficult to achieve efficient and high-precision model optimization under limited resources.

Method used

By acquiring multi-source unstructured data from industrial sites, data type identification and source labeling are performed. Semantic standardization is carried out using multimodal parsing and industrial entity recognition to construct structured sample data. Hardware perception quantization and memory optimization are combined with industrial knowledge graphs, and the attention layer of a pre-set large model is injected to perform multi-objective joint training for language modeling and knowledge consistency verification.

Benefits of technology

It achieves full automation from unstructured data to model fine-tuning, adapts to dynamic changes in industry, shortens the model development cycle, is applicable to a variety of intelligent industrial scenarios, has good versatility and transferability, and avoids high retraining costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122412879A_ABST
    Figure CN122412879A_ABST
Patent Text Reader

Abstract

本发明公开了一种工业大模型非结构化数据治理与自适应微调方法,涉及人工智能与工业大数据处理技术领域,包括:基于工业现场多源非结构化数据识别类型并标记来源,得到原始数据对象;通过多模态解析提取文本及结构关联,得到结构化解析结果;经工业实体识别与语义标准化处理,得到标准语料;通过映射规则构建输入输出对应,形成结构化样本数据;基于样本数据与知识图谱,通过硬件感知确定量化与显存优化,将图谱向量约束注入注意力层,以多目标联合训练得到微调后的工业大模型。通过多模态解析与工业语义治理构建标准化语料,并基于硬件感知及知识图谱约束的多目标微调,实现工业非结构化数据高效治理与大模型自适应优化。
Need to check novelty before this filing date? Find Prior Art