Large model driven multi-source heterogeneous mold knowledge graph construction method and system

CN122673370APending Publication Date: 2026-09-01BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611008942.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

这种分散化存储格局导致企业知识复用效率低下、经验传承断层问题突出,是制约模具行业向智能制造转型升级的重要瓶颈

Benefits of technology

采用七类异构数据源的统一纳管打破了模具企业长期存在的信息孤岛;大语言模型驱动的双路径半自动化Schema构建大幅减少人工工作量,同时保证了Schema与数据知识结构的高度契合;从历史最佳交付件中系统提炼的经验规律实现了隐性知识的显性化与结构化传承;七类智能服务将知识图谱的价值直接嵌入模具设计工作流的关键决策节点,从被动查询升级为主动赋能,为模具行业智能化转型提供了系统性技术支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122673370A_ABST
    Figure CN122673370A_ABST
Patent Text Reader

Abstract

The application discloses a large model driven multi-source heterogeneous mold knowledge graph construction method and system, acquires multi-source heterogeneous original data of a mold full business domain; pre-processes the multi-source heterogeneous original data of the mold full business domain, generates text block data which is complete in semantics and adapts to field ontology annotation requirements; based on the pre-processed text block data, generates a standardized field Schema and a mold design concept graph framework; completes directional knowledge triple extraction, alignment, cleaning, completion and fusion with the Schema as a hard constraint; constructs a heterogeneous modal knowledge network and completes cross-modal alignment through three types of learning methods, simultaneously mines implicit knowledge in historical delivery items; constructs a collaborative knowledge graph in four dimensions and completes persistent storage and incremental iteration; and builds seven types of core intelligent service to support mold full business domain collaborative design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of industrial knowledge engineering and artificial intelligence technology, specifically involving a method and system for constructing a multi-source heterogeneous mold knowledge graph driven by a large model. Background Technology

[0002] The design and manufacture of stamping dies is a typical knowledge-intensive engineering activity, involving highly specialized business processes such as structural design, material selection, precision machining, assembly, and debugging. In actual production, core knowledge of dies is scattered across various media, including technical standard documents, process specification manuals, CAD models, engineering drawings, daily production reports, and equipment maintenance records, exhibiting a clear multi-source heterogeneous characteristic. This decentralized storage pattern leads to low efficiency in knowledge reuse and significant gaps in experience transmission, posing a major bottleneck to the transformation and upgrading of the die industry towards intelligent manufacturing.

[0003] At the knowledge acquisition technology level, some progress has been made in the automated information extraction of industrial documents. However, the unique complexity of the industrial mold field makes it difficult to directly apply general solutions. Complex PDF layouts (spreading columns and pages, mixed text and images), two-dimensional mechanical engineering drawings with rich geometric semantics, three-dimensional assembly models stored in 3DXML or STEP format, and unstructured fault record texts differ fundamentally in information density, coding standards, and semantic expression. The same concept often has multiple equivalent expressions in different contexts. General named entity recognition models experience a sharp drop in accuracy when dealing with highly specialized corpora, and a large number of implicit design constraints and process rules are even more difficult to obtain directly through text extraction.

[0004] At the knowledge organization technology level, existing ontology modeling methods for discrete manufacturing generally adopt a purely manual definition approach. Schema construction relies entirely on domain experts who spend a significant amount of time and effort sorting out the conceptual system and relational constraints. This is not only time-consuming and costly, but the constructed results often deviate from the objectively existing knowledge structure in the data. Existing research is mostly limited to single-type data processing, with text knowledge graphs and geometric feature databases operating in isolation, lacking effective semantic bridging mechanisms, and having insufficient capabilities in cross-modal relational reasoning and multi-dimensional comprehensive querying.

[0005] At the knowledge application technology level, existing systems generally remain at the level of passive document retrieval, unable to provide proactive intelligent decision support such as automated verification of design specifications, intelligent parameter recommendation, and defect cause tracing. This is significantly different from the actual needs of the mold industry, which relies heavily on experience-based judgment. How to establish a complete link from multimodal raw data to a reasonable and sustainably evolving structured knowledge graph is a core technological challenge that urgently needs to be addressed in the field. Summary of the Invention

[0006] To address the problems existing in the prior art, this invention provides a method and system for constructing a multi-source heterogeneous mold knowledge graph driven by a large model. It organically integrates the semantic understanding and code generation capabilities of a large language model, multimodal feature alignment technology, and the flexible storage and efficient retrieval capabilities of a graph database, thus constructing a complete technical chain from massive heterogeneous mold data to a high-quality structured knowledge graph, and then to intelligent services across all business domains.

[0007] To achieve the above objectives, the present invention provides the following solution: A method for constructing a multi-source heterogeneous mold knowledge graph driven by a large model, including: Step S1: Obtain multi-source heterogeneous raw data from the entire mold business domain; among which, multi-source heterogeneous raw data includes: mold-related graphic materials, NX API and Python OCC help documents, CAD-related 3D files, complex layout PDF documents, mechanical engineering drawings, production control data, and operation and maintenance service data; Step S2: Preprocess the multi-source heterogeneous raw data of the entire mold business domain to generate semantically complete text block data that meets the domain ontology annotation requirements; Step S3: Based on the preprocessed text block data, generate a standardized domain schema and mold design concept map framework; Step S4: Using the standardized domain schema as a hard constraint, complete the targeted knowledge triple extraction of the full multimodal data, simultaneously perform knowledge alignment of terms, referential relationships, and multimodal content, and perform cleaning, completion, and fusion processing on the extraction results to generate standardized knowledge objects with high confidence. Step S5: Based on standardized knowledge objects with high confidence, construct a heterogeneous modal knowledge network with cross-dimensional association between text, image and 3D model, complete cross-modal feature extraction and semantic alignment, and at the same time, based on the mold design full-dimensional feature classification system, mine and supplement the implicit experience knowledge of mold design from the enterprise's historical design deliverables. Step S6: Classify standardized knowledge objects and tacit knowledge by dimension, construct a four-dimensional collaborative mold knowledge graph based on the mold design concept graph framework, and complete persistent storage through the iDME engine. At the same time, based on the closed-loop test verification process and incremental data, the four-dimensional collaborative mold knowledge graph is obtained. Among them, based on the four-dimensional collaborative mold knowledge graph, a mold full business domain service interface is built to output the full-process intelligent service capability corresponding to mold collaborative design.

[0008] As a preferred option, in step S1, the data is divided according to modal type and business process based on the multi-source heterogeneous raw data of the entire mold business domain to obtain a unified knowledge source pool for the mold domain.

[0009] Preferably, in step S2, a differentiated parsing strategy is executed for the original data of different modalities, the effective content of each modality is extracted and converted into a standardized intermediate format, and then semantically sliced ​​to generate text block data that is semantically complete and meets the requirements of domain ontology annotation.

[0010] As a preferred option, in step S3, based on the preprocessed text block data, through the linkage of multimodal large model and large language model, with the five elements of mold domain ontology as the core, the semi-automatic ontology schema construction, terminology alignment and verification optimization are completed, and a standardized domain schema and mold design concept map framework are generated.

[0011] This invention also provides a large-model-driven multi-source heterogeneous mold knowledge graph construction system, comprising: The first processing module is used to acquire multi-source heterogeneous raw data from the entire mold business domain. The multi-source heterogeneous raw data includes: mold-related graphic materials, NX API and Python OCC help documents, CAD-related 3D files, complex PDF documents, mechanical engineering drawings, production control data, and operation and maintenance service data. The second processing module is used to preprocess the multi-source heterogeneous raw data of the entire mold business domain to generate semantically complete text block data that meets the requirements of domain ontology annotation. The third processing module is used to generate a standardized domain schema and mold design concept map framework based on the preprocessed text block data. The fourth processing module is used to extract targeted knowledge triples from the full amount of multimodal data with standardized domain schema as hard constraint, simultaneously perform knowledge alignment of terms, referential relations and multimodal content, and perform cleaning, completion and fusion processing on the extraction results to generate standardized knowledge objects with high confidence. The fifth processing module is used to construct a heterogeneous modal knowledge network that links text, images, and 3D models across dimensions based on standardized knowledge objects with high confidence, and to complete cross-modal feature extraction and semantic alignment. At the same time, based on the full-dimensional feature classification system of mold design, it mines and supplements the implicit experience knowledge of mold design from the enterprise's historical design deliverables. The sixth processing module is used to classify standardized knowledge objects and tacit knowledge by dimension, construct a four-dimensional collaborative mold knowledge graph based on the mold design concept graph framework, and complete persistent storage through the iDME engine. At the same time, based on the closed-loop test verification process and incremental data, the four-dimensional collaborative mold knowledge graph is obtained. Among them, based on the four-dimensional collaborative mold knowledge graph, a full business domain service interface for molds is built, and the full-process intelligent service capability corresponding to mold collaborative design is output.

[0012] As a preferred approach, the first processing module divides the multi-source heterogeneous raw data from the entire mold business domain into modal types and business processes to obtain a unified knowledge source pool for the mold domain.

[0013] Preferably, the second processing module performs differentiated parsing strategies on the raw data of different modalities, extracts the effective content of each modality and converts it into a standardized intermediate format, and generates semantically complete text block data that meets the requirements of domain ontology annotation through semantic slicing.

[0014] As a preferred option, the third processing module, based on the preprocessed text block data, uses the multimodal large model and the large language model in conjunction to complete the construction of a semi-automatic ontology schema, terminology alignment and validation optimization, with the five elements of the mold domain ontology as the core, and generates a standardized domain schema and mold design concept map framework.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: The unified management of seven types of heterogeneous data sources breaks down the information silos that have long existed in mold enterprises; the dual-path semi-automated schema construction driven by the large language model greatly reduces the amount of manual work, while ensuring a high degree of fit between the schema and the data knowledge structure; the experience and rules extracted from the best historical deliverables realize the explicit and structured inheritance of tacit knowledge; and the seven types of intelligent services directly embed the value of knowledge graphs into the key decision nodes of the mold design workflow, upgrading from passive querying to proactive empowerment, providing systematic technical support for the intelligent transformation of the mold industry. Attached Figure Description

[0016] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of the method for constructing a multi-source heterogeneous mold knowledge graph driven by a large model, as described in an embodiment of the present invention. Figure 2 This is a schematic diagram of multimodal knowledge extraction; it includes LLM-driven semi-automated schema construction and knowledge transformation technology. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] Example 1 like Figure 1 As shown, this invention provides a method for constructing a multi-source heterogeneous mold knowledge graph driven by a large model, including: Step S1: Obtain multi-source heterogeneous raw data from the entire mold business domain; the multi-source heterogeneous raw data includes: mold-related graphic materials, NX API and Python OCC help documents, CAD-related 3D files, complex layout PDF documents, mechanical engineering drawings, production control data, and operation and maintenance service data; based on the multi-source heterogeneous raw data from the entire mold business domain, the data is divided according to modal type and business link to obtain a unified knowledge source pool for the mold field; Step S2: Preprocess the multi-source heterogeneous raw data of the entire business domain of molds, execute differentiated parsing strategies for raw data of different modalities, extract the effective content of each modality and convert it into a standardized intermediate format, and generate semantically complete text block data that meets the requirements of domain ontology annotation through semantic slicing. Step S3: Based on the preprocessed text block data, through the linkage of multimodal large model and large language model, with the five elements of mold domain ontology as the core, complete the construction of semi-automatic ontology schema, terminology alignment and verification optimization, and generate standardized domain schema and mold design concept map framework. Step S4: Using the standardized domain schema as a hard constraint, complete the targeted knowledge triple extraction of the full multimodal data, simultaneously perform knowledge alignment of terms, referential relationships, and multimodal content, and perform cleaning, completion, and fusion processing on the extraction results to generate standardized knowledge objects with high confidence. Step S5: Based on standardized knowledge objects with high confidence, construct a heterogeneous modal knowledge network with cross-dimensional association between text, image and 3D model, complete cross-modal feature extraction and semantic alignment, and at the same time, based on the mold design full-dimensional feature classification system, mine and supplement the implicit experience knowledge of mold design from the enterprise's historical design deliverables. Step S6: Classify standardized knowledge objects and tacit knowledge by dimension, construct a four-dimensional collaborative mold knowledge graph based on the mold design concept graph framework, and complete persistent storage through the iDME engine. At the same time, based on the closed-loop test verification process and incremental data, the four-dimensional collaborative mold knowledge graph is obtained. Among them, based on the four-dimensional collaborative mold knowledge graph, a mold full business domain service interface is built to output the full-process intelligent service capability corresponding to mold collaborative design.

[0021] As one embodiment of the present invention, step S1 includes: Step 1.1: Fully cover the core raw data of the entire end-to-end business chain of mold making, including at least seven categories: mold-related graphic materials, NX API and Python OCC help documents, CAD-related 3D files, complex-format PDF documents, mechanical engineering drawings, production control data, and operation and maintenance service data. Among them, mold-related graphic materials cover mold design drawings, production process flow charts, material performance curves, and defect analysis report images, while complex-format PDF documents cover mold design standards, specifications, guidelines, design manuals, and product manuals. Step 1.2: Divide the raw data into four categories according to modal type: text, image, 3D model, and table / formula. At the same time, divide it into eleven business domains according to business process: customer inquiry, quotation and feasibility analysis, process design, structural design, drawing and material purchase, manufacturing and processing, assembly and welding, machine debugging, product inspection, mold shipment, and master control management.

[0022] As one embodiment of the present invention, step S2 includes: Step 2.1: Perform a four-step parsing process for complex PDF documents. Step 1: Identify all elements in the PDF page using a layout detection model. Step 2: Divide the page layout into levels and determine the reading order of block-level elements. Step 3: Remove interfering elements such as headers, footers, and footnotes. Step 4: Complete the content rearrangement and sorting, extract and restore the original text, images, tables, and formulas, and output standardized JSON format intermediate data with page numbers, business process markers, and type markers. Step 2.2: Perform two-dimensional analysis on the mechanical engineering drawings. Use high-precision OCR to identify text, symbols, dimensions and tolerance information in the drawings. Use geometric understanding algorithms to analyze lines, shapes and spatial topological relationships, and output structured graph database nodes, edges and corresponding query statements. Step 2.3: For CAD-related 3D files, use the Python OCC deconstruction tool to parse the 3D model data, assembly structure, and geometric topology information, and extract part dimensions, assembly relationships, tolerance annotations, and example design parameters from the digital model operation history; Step 2.4: For all parsed text content, perform content slicing based on semantics and business processes through the semantic segmentation model to generate text block data that is semantically complete, contextually coherent, and adapted to the domain ontology annotation requirements, providing basic data for subsequent knowledge extraction and ontology construction.

[0023] As one embodiment of the present invention, step S3 is used for the closed-loop construction and optimization of a semi-automatic ontology schema in the mold domain. The five elements of the mold domain ontology are concept, relationship, attribute, instance, and axiom, specifically including: Step 3.1: Execute the knowledge system initialization process, complete the selection of representative materials for the topic, confirmation of business topics, sorting out the business of the domain superior, sorting out the business of the topic level, sorting out the object classification, sorting out the knowledge classification, and after expert review, complete the knowledge system initialization and storage, and build the basic knowledge framework; Step 3.2: Input the background prompts for the mold domain knowledge extraction task, industry standard constraints and preprocessed text block data into the multimodal large model MLLM, and extract the knowledge triples related to the task through MLLM to output the initial knowledge triple set; Step 3.3: Using the Large Language Model (LLM), the initial knowledge triple set is sorted and summarized based on the five elements of domain ontology, including entity concepts, relationships, attribute definitions, instance specifications, and axiom constraints, and an initial schema framework is automatically generated. Step 3.4: Execute the domain ontology annotation process, complete document knowledge classification annotation, directory division verification, knowledge point division and classification annotation, knowledge point content word segmentation and domain ontology annotation, perform type-specific annotation for result-type content, product standard content, method standard content, and product manual content, and output the annotation results after expert review; Step 3.5: Based on terminology alignment, complete the fusion and optimization of the initial schema and ontology annotation results, submit them for verification by mold domain experts, and combine mold domain ontology constraints. Based on FewShot, CoT thinking chain and analogy hint methods, correct entity and relationship definitions that do not conform to industry standards, and supplement mold domain-specific axiom constraints. Step 3.6: Output the final standardized mold domain schema and simultaneously generate a mold design concept map framework, which serves as a unified constraint specification for subsequent knowledge extraction throughout the entire process.

[0024] As one embodiment of the present invention, step S4 is used for multimodal knowledge extraction and fusion processing under schema constraints, specifically including: Step 4.1: Using the standardized domain schema as a hard constraint, perform targeted knowledge triple extraction for different types of data sources. Extract specification knowledge triples for mold design specifications, standards, and guidelines; extract design instance triples for mold digital models and 3D files; extract API operation triples for NX interface documents, API descriptions, and modeling operation scripts; and extract conceptual knowledge triples for mold terminology, materials, and equipment information. Step 4.2: Simultaneously perform multimodal knowledge alignment to complete the unification of terminology, resolution of referential relationships, and cross-modal semantic alignment of extracted content, thereby resolving the problem of differences in the representation of the same entity in different data sources and different modalities; Step 4.3: Perform knowledge cleaning on the extracted triples to filter redundant content, remove erroneous content, and standardize the format of technical terms; Step 4.4: Perform knowledge completion on the cleaned triples, and complete the missing entity attributes, relationships and instance information based on the domain ontology constraints; Step 4.5: Perform knowledge fusion and conflict resolution on the processed triples. Through confidence calculation and cross-validation, process the conflict information of multi-source data, resolve content conflicts, and finally generate high-confidence, standardized Python data objects or JSON format knowledge objects, and complete temporary persistent storage.

[0025] As one embodiment of the present invention, step S5 is used for cross-modal knowledge alignment and latent knowledge mining in mold design, specifically including: Step 5.1: Perform unified multimodal feature extraction. For image data, use a large visual model to extract visual features through semantic segmentation, object detection, and image classification. For 3D models, use a sample image deconstruction tool to extract structural and geometric features. For document data, use a large language model to extract semantic and textual features. Step 5.2: Integrate mold type, structure, geometric features, API description, topological relationship, process constraints, and business process attribute information to build a heterogeneous modal knowledge network with multi-dimensional associations of text, images, and 3D models; Step 5.3: Using heterogeneous graph contrastive learning, multimodal classification, and self-supervised learning methods, the feature extraction model is optimized through multimodal contrastive loss and cross-modal reconstruction loss to achieve semantic association and feature alignment of cross-modal knowledge, and output a modality-aligned mold knowledge vector library; Step 5.4: Establish a full-dimensional feature classification system for mold design, and define four major categories of core feature elements, including basic geometric features, function / performance related features, assembly and structural features, and precision related features; Step 5.5: Based on the full-dimensional feature classification system, extract features and summarize patterns from the company's best historical design deliverables. Assign corresponding confidence levels to the patterns according to the sample size and statistical results, mine the experiential tacit knowledge of non-standard mold design, and structure it to supplement the standardized knowledge objects.

[0026] As one embodiment of the present invention, step S6 is used for the construction and iterative optimization of the four-dimensional collaborative mold knowledge graph, specifically including: Step 6.1: Classify and map the standardized knowledge objects and the mined tacit knowledge according to four dimensions: concepts, norms, instances, and tools; Step 6.2: Based on the mold design concept graph framework, four sub-graphs are constructed to form a four-dimensional collaborative architecture: The mold design concept graph includes mold type, design stage, core terminology, and business process, presenting the adaptation of different types to products and the logic before and after each stage, forming a unified design knowledge framework; the mold design common knowledge graph covers mold frame structure and main functional components, summarizing the structural parameters and assembly rules of different mold types; the historical best practice graph organizes material properties and key factors affecting the forming process, combining common defects with corresponding design adjustment schemes; the 3D code generation knowledge graph maps 3D geometric features to forming processes and associates them with corresponding processing code templates; Step 6.3: Import the knowledge nodes, relationships, and attribute information of the four subgraphs into the graph database, establish semantic relationship links across subgraphs, and complete the persistent storage of the full graph data through the iDME engine to form a complete four-dimensional collaborative mold knowledge graph. Step 6.4: Execute the closed-loop testing and verification process, complete the determination of test objectives, design of usability / standardization / effectiveness evaluation indicators, formulation of test cases and plans, test execution and feedback adjustment, and complete the graph optimization after exit review; Step 6.5: Based on the incrementally accessed multimodal data, repeat the process from Step 2 to Step 5 to import the newly added knowledge objects into the graph database. Combine incremental learning with expert rule verification to achieve continuous dynamic iteration and self-improvement of the graph.

[0027] As one embodiment of the present invention, step S7 is used to output the collaborative design service capability of the entire mold business domain based on knowledge graph, specifically including: Step 7.1: Output intelligent retrieval service for mold design knowledge. Based on a modally aligned mold knowledge vector library, it supports natural language retrieval and enables accurate querying and result return of design specifications, cases, parameters, and API usage. Step 7.2: Output automated design specification verification service. Based on the common knowledge graph of mold design, automatically verify whether the non-standard mold design scheme conforms to industry standards and enterprise specifications, and output verification report and actionable modification suggestions; Step 7.3: Output intelligent recommendation service for design parameters. Based on the historical best example map, combined with the current design requirements and constraints, intelligently recommend suitable mold frame parameters, material selection, and tolerance setting schemes. Step 7.4: Output 3D machining code automatic generation service, generate knowledge graph based on 3D code, match the geometric features and process requirements of the design model, and automatically generate corresponding CNC machining code; Step 7.5: Output intelligent mold defect diagnosis service. Based on defect analysis cases and process knowledge, quickly locate the cause of mold forming defects and recommend corresponding design adjustments and mold repair solutions. Step 7.6: Output the automatic generation service of molding structure mold design scheme. Understand the user's design requirements through digital model analysis, retrieve matching standards, specifications and parameters from the atlas, and automatically generate a molding structure mold design scheme that meets the requirements. Step 7.7: Output multi-role, full-business-domain collaborative design support services, providing a unified knowledge base for multiple roles such as design, process, production, quality inspection, and operation and maintenance, supporting cross-position collaborative design and knowledge sharing of the entire business process of mold making from customer inquiry to mastering management.

[0028] Example 2 This invention also provides a method for constructing a multi-source heterogeneous mold knowledge graph driven by a large model, including: Step S1: Obtain multi-source heterogeneous raw data from the entire business domain of mold making; Step S2: Preprocess the multi-source heterogeneous raw data of the entire mold business domain to generate semantically complete text block data that meets the domain ontology annotation requirements; Step S3: Based on the preprocessed text block data, generate a standardized domain schema and mold design concept map framework; Step S4: Using the standardized domain schema as a hard constraint, complete the targeted knowledge triple extraction of the full multimodal data, simultaneously perform knowledge alignment of terms, referential relationships, and multimodal content, and perform cleaning, completion, and fusion processing on the extraction results to generate standardized knowledge objects with high confidence. Step S5: Based on standardized knowledge objects with high confidence, construct a heterogeneous modal knowledge network with cross-dimensional association between text, image and 3D model, complete cross-modal feature extraction and semantic alignment, and at the same time, based on the mold design full-dimensional feature classification system, mine and supplement the implicit experience knowledge of mold design from the enterprise's historical design deliverables. Step S6: Classify standardized knowledge objects and tacit knowledge by dimension, construct a four-dimensional collaborative mold knowledge graph based on the mold design concept graph framework, and complete persistent storage through the iDME engine. At the same time, based on the closed-loop test verification process and incremental data, the four-dimensional collaborative mold knowledge graph is obtained. Among them, based on the four-dimensional collaborative mold knowledge graph, a mold full business domain service interface is built to output the full-process intelligent service capability corresponding to mold collaborative design.

[0029] As one implementation of this invention, the goal of step S1 is to establish a unified knowledge source pool covering the entire business domain of mold manufacturing. The core knowledge assets of mold manufacturing enterprises are scattered across seven types of carriers. This step clarifies the access scope for each type and completes systematic management through a two-dimensional classification framework.

[0030] The seven types of raw data are as follows: The first category consists of mold-related graphic and textual materials. This includes mold design drawings (general assembly drawings, detailed drawings of cavity and core parts, cooling system layout diagrams, etc.), production process flow charts (process cards, process node diagrams), material property curves (stress-strain curves, heat treatment curves, fatigue life curves, etc.), and defect analysis report images (macroscopic photographs and microscopic metallographic images). This type of material constitutes the main carrier of mold knowledge, providing industry background knowledge for the semantic understanding of large-scale models and serving as the core source of material for constructing a historical best-in-class example atlas.

[0031] The second category includes NX API and Python OCC help documentation, which integrates official NX function interface documentation (including NXOpen object library descriptions, command parameter formats, and typical call examples) and Python OCC geometric modeling interface documentation. This type of documentation records complete technical specifications from 3D model geometric operations to parametric design, serving as a key source for 3D code generation knowledge graphs. The third category includes CAD-related 3D files, integrating 3D assembly and part model files in formats such as 3DXML and STEP. This extracts part geometric topology, assembly hierarchy relationships, and historical data of digital model operations, providing accurate 3D geometric quantification data for the instance-level knowledge graph.

[0032] The fourth category consists of complex-format PDF documents, covering national standards for mold design (GB / T mold terminology standards, etc.), internal enterprise design specifications, operating instructions, material manuals, and product manuals. These documents are authoritative and highly standardized, serving as a core source of standardized knowledge in the common knowledge graph of mold design, and are also the most difficult data type to analyze. The fifth category comprises mechanical engineering drawings, i.e., two-dimensional engineering drawings exported from or scanned from CAD systems. These drawings include dimensioning, geometric tolerances, surface roughness requirements, and technical specifications, and are crucial for the subsequent design verification functions of the knowledge graph.

[0033] The sixth category is production control data, which includes production plans, process quality records, equipment operating status data, and material consumption records. It carries measured data fed back from the workshop manufacturing execution layer to the design layer and is an important source for dynamically updating historical instance maps.

[0034] The seventh category is operation and maintenance service data, which includes maintenance records after mold shipment, regular maintenance logs, failure mode analysis reports and customer feedback. It records the design flaws and improvement opportunities exposed by the mold under actual working conditions, and has irreplaceable reference value for the continuous optimization of the defect diagnosis knowledge base.

[0035] Two-dimensional classification management After completing the full access of seven data categories, a systematic classification was implemented based on two orthogonal dimensions: modality type and business process. The modality type dimension was divided into four main categories: text, image, 3D model, and table / formula. Each data category was assigned a different specialized parsing engine in step 2 to ensure parsing accuracy. The business process dimension tagged each piece of data with metadata based on its business scenario, dividing it into eleven business domains: customer inquiry, quotation and feasibility analysis, process design, structural design, drawing and material procurement, manufacturing and processing, assembly and welding, machine debugging, product inspection, mold shipment, and mastering management. These business domain tags played a crucial guiding role in the graph dimension construction in step 6 and the business scenario-oriented service in step 7, ensuring a high degree of match between knowledge retrieval results and the user's current business scenario.

[0036] As one implementation of this invention, step S2, based on the knowledge source pool established in step S1, calls the corresponding specialized parsing engine to perform content extraction according to data type, and then uniformly performs semantic slicing on all parsing results, outputting standardized intermediate format text blocks with business domain tags. For example... Figure 2 As shown, the multimodal knowledge extraction process consists of data preprocessing, schema construction, and knowledge transformation.

[0037] Four Steps to Analyze Complex PDF Layouts Industrial mold documentation often features complex PDF formats: cross-column and cross-page layouts are common, scanned documents are mixed with vector text, and technical terminology is dense, making it easily fragmented by common segmentation algorithms. This invention performs the following four-step specialized analysis on such documents. First, a layout element detection model pre-trained on an engineering document dataset is invoked to identify and locate text blocks, image areas, table areas, formula areas, annotation lines, and headers / footers on each page, outputting a list of page elements including category labels and bounding box coordinates. Second, based on the multi-column layout structure, the linear reading order of each content block is determined according to the logic of top-to-bottom within a column and left-to-right between columns. Third, headers / footers, page numbers, watermarks, and formulaic paragraphs are filtered, and redundant content such as repeated copyright notices is removed. Fourth, consecutive text paragraphs spanning columns and pages are combined to restore complete semantic units, which, along with image, table, and formula content objects, are encapsulated into standardized JSON format intermediate data according to business domain tags and page number indexes. Each content object is categorized into one of four types: text, image, table, or formula.

[0038] Two-Dimensional Analysis of Mechanical Engineering Drawings The analysis of mechanical engineering drawings requires processing both textual and geometric dimensions simultaneously. In the textual dimension, a high-precision OCR engine specifically trained for engineering drawing scenarios is used to recognize part names, technical requirements, geometric tolerance symbols (flatness, coaxiality, position, etc.), surface roughness markings, and dimensional values ​​(including tolerance deviations). The OCR recognition results are then validated for numerical reasonableness using regular expressions, filtering out obvious recognition errors. In the geometric dimension, a geometric understanding algorithm analyzes the line types (solid lines / dashed lines / center lines / section lines), basic geometric shapes, and their spatial topological relationships in the drawings, extracting spatial constraints such as tangency, perpendicularity, and coaxiality. The results from both dimensions are then integrated and output as a structured description of nodes / edges in the graph database.

[0039] Structured parsing of CAD 3D files The 3D model data is analyzed in three levels using the Python OCC deconstruction tool: The assembly structure layer extracts component identifiers, type labels, and parent-child hierarchical relationships for each node, outputting a tree-like component list; the geometric parameter layer extracts the position coordinates, bounding box dimensions, volume, surface area, and topological features of each component; and the assembly constraint layer identifies relationships such as coaxial constraints, face-fitting constraints, distance constraints, and symmetry constraints, outputting relationship triples in the form of (component A, constraint type, component B, constraint parameter value). Furthermore, the co-occurrence probability of components is statistically analyzed across multiple historical design cases, quantifying the interchangeability of adjacent components. Implicit relationships are appended to the analysis results as weighted edges, enriching the semantic density of the 3D model knowledge.

[0040] Semantic slicing and intermediate format unification For all types of parsing results, a unified content slicing process based on semantics and business processes is applied: for text content, the semantic similarity between adjacent blocks is used as the boundary, and the length of each chunk is controlled between 128 and 512 tokens; for mixed text and image content, the integrity of the image and its nearest text description is maintained, and no forced segmentation is performed between text and image pairs; for table content, the complete table is used as the smallest slice unit; for CAD structured data, the individual component and its direct relationships are used as the smallest slice unit. All chunks are appended with source filenames, page numbers, modal types, content types, and business process tags, and are uniformly encapsulated into a standardized intermediate format for schema construction in step 3 and knowledge extraction in step 4.

[0041] As one implementation of this invention, step S3 generates a standardized domain schema in a semi-automatic manner through a closed-loop process of knowledge system initialization → MLLM preliminary extraction → LLM automatic schema induction → domain ontology annotation → expert fusion verification. Simultaneously, it generates the concept graph framework required for subsequent knowledge graph construction, fundamentally solving the pain points of long cycle and limited coverage of the traditional purely manual definition method.

[0042] Knowledge system initialization The knowledge system initialization phase involves six tasks: selecting authoritative and comprehensive representative documents from the full dataset in Step 1 as initialization input; defining the knowledge map's coverage as the entire business domain of stamping dies; clarifying the knowledge boundaries between the die field and related fields such as materials science and processing technology; completing hierarchical decomposition according to design, process, manufacturing, quality inspection, and operation and maintenance categories and their corresponding secondary sub-businesses; identifying the core knowledge objects involved in each business module; and defining a knowledge system of four types: declarative knowledge, procedural knowledge, conditional knowledge, and strategic knowledge. After these six tasks are reviewed and approved by domain experts, a knowledge system initialization document is generated and stored in the database, serving as the top-level constraint benchmark for schema construction.

[0043] MLLM-driven initial triple extraction Three types of input are fed into the Multimodal Large Model (MLLM): a task background prompt explicitly stating that the extraction target is knowledge triples structured according to the five elements of ontology; industry constraints injected from the GB / T standard terminology system and enterprise specifications; and text block data generated in step 2 (including images, tables, and formula-related content). To improve coverage breadth and diversity, MLLM is called multiple times on the same batch of chunks with different random seeds. Highly consistent triples are retained through a majority voting mechanism, while rare triples that appear at least once are marked as "low-frequency candidates" and saved separately. Both types of triples enter the schema automatic induction process together, ensuring that high-frequency consensus knowledge and long-tail professional knowledge are simultaneously included in the schema coverage. Each triple is accompanied by a confidence score and the original sentence of the extraction basis.

[0044] LLM automatically generates an initial schema. The Large Language Model (LLM) is used to analyze and summarize the initial set of triples from the five dimensions of domain ontology: at the entity concept level, all entity types and their classification levels are identified; at the relation level, predicate relations are semantically categorized to form a list of relation types; at the attribute definition level, attribute keys and value range constraints for each entity type are identified; at the instance specification level, naming rules and format specifications for specific instances are summarized; and at the axiomatic constraint level, hard constraint rules (such as "cavity hardness must not be lower than HRC48") are identified and formalized into axiomatic expressions. Based on the above analysis, LLM automatically outputs a structured initial schema framework in JSON-LD format.

[0045] Domain ontology annotation and schema fusion optimization Systematic annotation was performed on all text blocks according to the initial schema framework: document knowledge classification annotation assigned domain ontology concept category labels to each knowledge object; knowledge point word segmentation and ontology annotation performed part-of-speech tagging and entity boundary identification for domain terms. Based on the nature of the content, categorized special annotations were performed: for deliverables, entity instances and quantitative parameter attributes were highlighted; for product standards, normative constraints and axiomatic relationships were highlighted; for methodological standards, the sequence of steps and conditional relationships of procedural knowledge were highlighted; and for product manuals, performance parameters and scope of application attributes were highlighted. All annotation results were sampled and reviewed by domain experts, and the annotation consistency had to reach 0.80 or higher before proceeding to the next stage. Subsequently, the initial schema and ontology annotation results were merged based on terminology alignment, submitted for expert verification, and non-compliant definitions were corrected using three types of prompting engineering methods: FewShot, CoT (Coding of Thought), and analogy prompts. Specific axiomatic constraints were added, and finally, a standardized schema was output in JSON-LD format, simultaneously generating a mold design concept map framework.

[0046] As one embodiment of the present invention, step S4 performs targeted knowledge triple extraction on the full multimodal data under the constraints of the standardized domain schema output in step S3, and sequentially performs knowledge alignment, cleaning, completion and fusion processing to produce standardized knowledge objects with high confidence.

[0047] Category-based targeted triplet extraction Based on the knowledge attributes of different data sources, targeted extraction was performed along four paths: For mold design specifications, standards, and guidelines, the focus was on extracting triples of specification knowledge, such as design constraints (minimum wall thickness requirements, lower limit of draft angle, etc.), material selection specifications, and process operation procedures; for mold digital models and 3D files, the focus was on extracting triples of design examples, such as component hierarchy relationships, geometric parameter attributes, and assembly constraint relationships; for NX interface documents, API descriptions, and modeling operation scripts, the focus was on extracting API operation triples, such as modeling function input / output parameter formats, operation commands, and geometric object association relationships; and for mold terminology, material, and equipment information, the focus was on extracting conceptual knowledge triples, such as terminology definitions and hierarchical relationships, material performance parameters, and the applicable process range of equipment. Each path used the relational constraints of the corresponding knowledge type in the schema as the core prompt, and combined FewShot and CoT (Coding on Trace) methods to ensure the stability and accuracy of the extraction.

[0048] Multimodal knowledge alignment, cleaning and completion In terms of alignment, three tasks are performed: terminology standardization, constructing an equivalence mapping table through joint matching of string similarity and semantic embedding similarity; denotation of referential relationships, determining the specific entities referred to by pronouns and abbreviations through context analysis; and cross-modal semantic alignment, identifying cross-modal entity pairs with a cosine similarity exceeding the threshold τ_align=0.80 as equivalent entities and merging their attribute information. In terms of cleaning, redundant content filtering is performed (retaining the most confident identical triples and merging near-synonymous triples), erroneous content removal (marking abnormal entities for manual review), and standardization of professional terminology format (unifying the writing format of unit symbols, numerical precision, and entity naming conventions). In terms of completion, implicit triples are derived based on axiomatic constraints, and missing attribute values ​​are inferred using analogy completion as a feature of similar entities (marking "inferred values" to distinguish them from "measured values"). Search completion is used to specifically search for evidence of missing attributes in the original data source.

[0049] Knowledge integration and conflict resolution The overall confidence score of the processed triples is calculated as follows: conf_final = γ1·conf_llm + γ2·conf_model + γ3·freq_support, where conf_llm is the confidence score of the large model output, conf_model is the predicted probability of the specific extraction model, and freq_support is the normalized frequency of the triple obtaining textual support in the full chunk. The weighting coefficients are γ1=0.4, γ2=0.3, and γ3=0.3. Numerical conflicts of the same attribute for the same entity are resolved according to priority: values ​​supported by clear standard documents take precedence over non-standard sources; values ​​from precise 3D model analysis take precedence over textual descriptive statements; and values ​​with higher confidence scores take precedence. Conflicts that cannot be automatically resolved are marked as "awaiting manual adjudication" and added to the expert review queue. After conflict resolution, a high-confidence, standardized JSON-formatted knowledge object is output and temporarily persisted.

[0050] As one embodiment of the present invention, step S5, based on the standardized knowledge object output in step S4, performs unified extraction of multimodal features, construction of heterogeneous modal knowledge network and cross-modal feature alignment, while systematically mining implicit experience knowledge in historical best deliverables, and producing a modally aligned mold knowledge vector library and a structured implicit knowledge supplementary set.

[0051] Multimodal feature extraction and heterogeneous knowledge network construction For image data, a large visual model is used to perform semantic segmentation (segmenting engineering images into regions with semantic labels), object detection (identifying key objects such as gate locations and cooling hole locations), and image classification, extracting visual features. For 3D models, the Python OCC tool is used to extract component geometric feature vectors and assembly structure feature vectors. For document data, a large language model is used to extract semantic embedding vectors and entity context attribute descriptions. The feature representations of the three modalities are processed by their respective specialized encoders and then uniformly projected into a 512-dimensional embedding space. Using the three modal knowledge objects produced in step S4 as nodes, mold type, part structure, geometric features, API descriptions, topological relationships, process constraints, and business link attribute information are integrated to construct a heterogeneous modal knowledge network containing intramodal homogeneous edges (connecting semantically similar entity pairs within the same modality, with edge weights determined by cosine similarity) and cross-modal heterogeneous edges (connecting equivalent entity pairs across different modalities).

[0052] Cross-modal feature alignment Three learning methods are jointly employed to optimize node representations: heterogeneous graph contrastive learning uses known equivalent cross-modal entity pairs as positive samples and random non-equivalent entity pairs as negative samples, driving the convergence of equivalent entity embeddings through a multimodal contrastive loss function; multimodal classification uses entity categories defined in the schema as labels to train a multimodal classifier that ensures semantic consistency across modalities; and self-supervised learning uses the structural similarity of the k-order neighborhood subgraphs of a node as an unsupervised signal, combined with a cross-modal reconstruction loss function to further improve semantic consistency. After the three methods are jointly trained, a modality-aligned schema knowledge vector base is output, supporting near-nearest neighbor retrieval in step 7 milliseconds.

[0053] Multi-dimensional feature classification system and tacit knowledge mining This invention constructs a comprehensive feature classification system for mold design, encompassing four core categories: basic geometric features (outer contour and boundary, center features, solid geometry, solid projection); function and performance-related features (hole features, boss / flange features, groove features, and array features); assembly and structural features (assembly relationships, connection features, and positioning features); and precision-related features (dimensional annotations, geometric tolerances, and surface quality). Based on this system, feature value records of various types are extracted in batches from historical best-design deliverables to form a feature data matrix. Statistical analysis is performed to summarize high-frequency combination patterns, and the patterns are structured and expressed as "the optimal value range of feature Y under condition X is Z." Three levels of confidence are assigned based on sample size: low (<5 cases), medium (5-20 cases), and high (>20 cases with a narrowed confidence interval). High-confidence patterns are written into the historical best-case scenario atlas, medium-confidence patterns are saved as reference patterns, and low-confidence patterns are recorded in the analysis log. The results of tacit knowledge mining are supplemented to the knowledge object output in step S4 in the form of "feature constraint rule" triples, realizing the explicit and structured inheritance of experiential knowledge.

[0054] As one implementation of the present invention, step S6 maps the full amount of knowledge produced in steps S4 and S5 into four-dimensional categories, constructs four sub-graphs with the concept graph framework of step 3 as the skeleton, completes persistent storage through the iDME engine, and ensures the continuous evolution quality of the graph through closed-loop testing and incremental learning mechanisms.

[0055] Four-dimensional classification mapping and subgraph construction

[0056] The entire set of knowledge objects and tacit knowledge are classified and mapped according to four dimensions: concepts (basic cognitive knowledge such as mold type definition, design stage division, core terminology interpretation and business process nodes), specifications (design specification knowledge such as mold frame structure specifications, functional component assembly rules, structural parameter constraints and industry standard requirements), examples (example experience knowledge such as historical best design scheme parameters, material selection records, defect cases and adjustment measures), and tools (tool operation knowledge such as the mapping relationship between 3D geometric features and forming process and NX modeling API operation triples). This allows a knowledge object to belong to multiple dimensions at the same time.

[0057] The mold design concept map uses a domain ontology concept hierarchy as its framework, filling in mold type nodes, design stage nodes, and core terminology nodes. It establishes the adaptation relationships between different mold types and product characteristics, as well as the logical dependencies between each design stage, providing the team with a unified professional cognitive benchmark. The mold design common knowledge map covers mold frame structure and main functional component specifications, summarizing common rules for parameters such as sprue location and ejector pin diameter arrangement, as well as assembly rules such as guide pillar and guide sleeve fitting accuracy levels. This provides a complete rule knowledge base for automated verification of design specifications in step 7. The historical best practice example map system organizes performance comparison data for various grades of mold steel and the influence of wall thickness uniformity on warpage deformation. Combined with verified and effective adjustment schemes for forming defects such as short shots, flash, warpage, and shrinkage marks, it constructs an example experience library for proactive risk prediction in the design stage and rapid diagnosis in the trial molding stage. The 3D code generation knowledge map structurally maps 3D geometric features to forming process requirements, associating them with corresponding NXOpen parametric modeling code templates, achieving end-to-end automatic connection from 3D model feature semantics to executable code.

[0058] Persistent storage, closed-loop verification, and incremental iteration The nodes, relationships, and attribute information of the four sub-graphs are imported into the graph database. Cross-graph reference edges are established for nodes belonging to different sub-graphs but describing the same real-world object, ensuring four-dimensional collaborative retrieval and reasoning. The iDME engine is used for persistent storage of all data, and a FAISS-IVF approximate nearest neighbor index is constructed simultaneously to support millisecond-level semantic similarity retrieval. Before the graph goes live, closed-loop testing is performed: an evaluation system is designed around four core indicators: entity completeness, relationship accuracy, axiomatic constraint compliance, and query performance. After testing, the root causes of non-compliant indicators are identified. The graph is officially launched only after all indicators have been jointly reviewed and approved by domain experts and system engineers. After launch, it continuously evolves through an incremental update mechanism: when new data is imported, steps S2 to S5 are repeated, replacing full retraining with an incremental learning strategy (fine-tuning the weights based on the previous training). Combined with expert rules, new logical contradictions are periodically scanned, achieving a virtuous cycle where the graph becomes more complete with use.

[0059] As one implementation of this invention, step S7 uses the four-dimensional collaborative mold knowledge graph constructed in step S6 as the underlying support to build a service interface layer for the entire business domain of molds, outputting seven types of core intelligent services. The knowledge supply for various services is guaranteed by a joint synergistic mechanism of RAG retrieval-enhanced generation and LoRA fine-tuning: RAG accurately retrieves relevant vertical knowledge from the graph and injects it into the large model generation process; LoRA fine-tuning strengthens the large model's deep understanding of domain concepts and reasoning logic, taking into account both professional depth and the ability to dynamically update knowledge.

[0060] Intelligent knowledge retrieval and automated verification of design specifications The knowledge intelligent retrieval service constructs a hybrid recall mechanism based on a modally aligned mold knowledge vector library: FAISS-IVF vector retrieval handles fuzzy queries; BM25 sparse retrieval handles precise technical terminology queries; and multi-hop traversal (maximum 5 hops) of the knowledge graph retrieves related knowledge nodes. The three results are fused and sorted by the Re-ranking module, returning knowledge content such as design specifications, historical cases, process parameters, and API usage, achieving millisecond-level accurate retrieval. The automated design specification verification service uses standard constraints in a common knowledge graph as its knowledge source, comparing design parameters with specification constraints item by item, automatically identifying non-compliant configurations, and outputting a verification report containing violation entries, specification basis, and recommendations for the closest compliant parameters, distinguishing between "confirmed violation" and "suspected violation" cases.

[0061] Parameter recommendation, code generation, and defect diagnosis The intelligent design parameter recommendation service is based on a historical best-in-class example graph. It embeds design requirements into a vector search to find the most similar historical high-quality case set, statistically analyzes high-frequency parameter configurations compatible with current constraints as recommended values, and includes historical trial results and quality feedback data to assist engineers in making informed decisions regarding mold parameters, material selection, and tolerance settings. The automatic 3D machining code generation service is based on a 3D code generation knowledge graph. It receives the geometric feature description and process requirements of the design model, accurately matches the NX API operation triple sequence through the mapping relationship from geometric features to modeling operations in the graph, automatically assembles the corresponding NXOpen parametric modeling code snippets, and generates a code draft that can be directly executed after review. The intelligent mold defect diagnosis service starts with the defect phenomenon entity and performs a multi-hop causal path search along the "potentially possible" relationship chain to identify possible upstream design factors such as uneven wall thickness, improper gate location, and insufficient cooling. It returns a structured cause chain with historical support case counts and simultaneously retrieves and recommends verified and effective design adjustment solutions ranked by historical success rate.

[0062] Automatic solution generation and multi-role collaborative design support The automatic generation service for mold design schemes analyzes the product digital models submitted by users, combines process information and functional requirements, and retrieves the most similar historical schemes through a two-way fusion of graph matching neural network structural similarity matching and large language model semantic similarity matching. For new requirements without high similarity references, a large language model combined with KAG (Knowledge Augmentation) prompts the engineering process, generating parametric design schemes according to a hierarchical strategy from local to global under the constraints of design specification knowledge. After geometric constraint verification and compliance checks, the schemes are output, and engineers can make multiple rounds of interactive adjustments via natural language commands throughout the process. The multi-role, full-business-domain collaborative design support service uses a four-dimensional collaborative mold knowledge graph as a unified knowledge base. Design engineers obtain specifications and recommendations through concept graphs and common knowledge graphs; process engineers obtain process parameters and code templates through historical instance graphs and code generation graphs; and production, quality inspection, and operation and maintenance personnel obtain the required information through their respective knowledge domains. The unified knowledge base breaks down information barriers between roles, enabling design intent to be accurately transmitted to the manufacturing execution layer. Feedback knowledge from production and operation and maintenance ends continuously flows back to the design library, forming a knowledge loop covering the entire business process.

[0063] This invention has been verified in an engineering project at a stamping die company. It has achieved remarkable results in terms of the breadth of knowledge integration, extraction accuracy, digitization rate of tacit knowledge, and service capability coverage. It has successfully transformed mold-related knowledge scattered in various carriers into four-dimensional collaborative digital assets that can be called upon, linked, and continuously evolved, which strongly supports the implementation of the company's dual goals of standardization and intelligentization in mold design.

[0064] Example 3 This invention also provides a large-model-driven multi-source heterogeneous mold knowledge graph construction system, comprising: The first processing module is used to acquire multi-source heterogeneous raw data from the entire mold business domain. The multi-source heterogeneous raw data includes: mold-related graphic materials, NX API and Python OCC help documents, CAD-related 3D files, complex PDF documents, mechanical engineering drawings, production control data, and operation and maintenance service data. The second processing module is used to preprocess the multi-source heterogeneous raw data of the entire mold business domain to generate semantically complete text block data that meets the requirements of domain ontology annotation. The third processing module is used to generate a standardized domain schema and mold design concept map framework based on the preprocessed text block data. The fourth processing module is used to extract targeted knowledge triples from the full amount of multimodal data with standardized domain schema as hard constraint, simultaneously perform knowledge alignment of terms, referential relations and multimodal content, and perform cleaning, completion and fusion processing on the extraction results to generate standardized knowledge objects with high confidence. The fifth processing module is used to construct a heterogeneous modal knowledge network that links text, images, and 3D models across dimensions based on standardized knowledge objects with high confidence, and to complete cross-modal feature extraction and semantic alignment. At the same time, based on the full-dimensional feature classification system of mold design, it mines and supplements the implicit experience knowledge of mold design from the enterprise's historical design deliverables. The sixth processing module is used to classify standardized knowledge objects and tacit knowledge by dimension, construct a four-dimensional collaborative mold knowledge graph based on the mold design concept graph framework, and complete persistent storage through the iDME engine. At the same time, based on the closed-loop test verification process and incremental data, the four-dimensional collaborative mold knowledge graph is obtained. Among them, based on the four-dimensional collaborative mold knowledge graph, a full business domain service interface for molds is built, and the full-process intelligent service capability corresponding to mold collaborative design is output.

[0065] As one embodiment of the present invention, the first processing module divides the data according to modal type and business link based on the multi-source heterogeneous raw data of the entire mold business domain, and obtains a unified knowledge source pool in the mold domain.

[0066] As one embodiment of the present invention, the second processing module performs a differentiated parsing strategy on the raw data of different modalities, extracts the effective content of each modality and converts it into a standardized intermediate format, and generates semantically complete text block data that meets the requirements of domain ontology annotation through semantic slicing processing.

[0067] As one embodiment of the present invention, the third processing module, based on the preprocessed text block data, uses the linkage between the multimodal large model and the large language model, with the five elements of the mold domain ontology as the core, to complete the construction of a semi-automatic ontology schema, terminology alignment and verification optimization, and generate a standardized domain schema and mold design concept map framework.

[0068] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for constructing a multi-source heterogeneous mold knowledge graph driven by a large model, characterized in that, include: Step S1: Obtain multi-source heterogeneous raw data from the entire mold business domain; among which, multi-source heterogeneous raw data includes: mold-related graphic materials, NX API and Python OCC help documents, CAD-related 3D files, complex layout PDF documents, mechanical engineering drawings, production control data, and operation and maintenance service data; Step S2: Preprocess the multi-source heterogeneous raw data of the entire mold business domain to generate semantically complete text block data that meets the domain ontology annotation requirements; Step S3: Based on the preprocessed text block data, generate a standardized domain schema and mold design concept map framework; Step S4: Using the standardized domain schema as a hard constraint, complete the targeted knowledge triple extraction of the full multimodal data, simultaneously perform knowledge alignment of terms, referential relationships, and multimodal content, and perform cleaning, completion, and fusion processing on the extraction results to generate standardized knowledge objects with high confidence. Step S5: Based on standardized knowledge objects with high confidence, construct a heterogeneous modal knowledge network with cross-dimensional association between text, image and 3D model, complete cross-modal feature extraction and semantic alignment, and at the same time, based on the mold design full-dimensional feature classification system, mine and supplement the implicit experience knowledge of mold design from the enterprise's historical design deliverables. Step S6: Classify standardized knowledge objects and tacit knowledge by dimension, construct a four-dimensional collaborative mold knowledge graph based on the mold design concept graph framework, and complete persistent storage through the iDME engine. At the same time, based on the closed-loop test verification process and incremental data, the four-dimensional collaborative mold knowledge graph is obtained. Among them, based on the four-dimensional collaborative mold knowledge graph, a mold full business domain service interface is built to output the full-process intelligent service capability corresponding to mold collaborative design.

2. The method for constructing a multi-source heterogeneous mold knowledge graph driven by a large model as described in claim 1, characterized in that, In step S1, based on the multi-source heterogeneous raw data of the entire mold business domain, the data is divided according to modal type and business link to obtain a unified knowledge source pool for the mold domain.

3. The method for constructing a large-model-driven multi-source heterogeneous mold knowledge graph as described in claim 2, characterized in that, In step S2, a differentiated parsing strategy is executed for the raw data of different modalities to extract the effective content of each modality and convert it into a standardized intermediate format. After semantic slicing, text block data with semantic integrity and adapted to the domain ontology annotation requirements is generated.

4. The method for constructing a large-model-driven multi-source heterogeneous mold knowledge graph as described in claim 3, characterized in that, In step S3, based on the preprocessed text block data, through the linkage of multimodal large model and large language model, with the five elements of mold domain ontology as the core, the semi-automatic ontology schema construction, terminology alignment and verification optimization are completed, and a standardized domain schema and mold design concept map framework are generated.

5. A large-model-driven multi-source heterogeneous mold knowledge graph construction system, characterized in that, include: The first processing module is used to acquire multi-source heterogeneous raw data from the entire mold business domain. The multi-source heterogeneous raw data includes: mold-related graphic materials, NX API and Python OCC help documents, CAD-related 3D files, complex PDF documents, mechanical engineering drawings, production control data, and operation and maintenance service data. The second processing module is used to preprocess the multi-source heterogeneous raw data of the entire mold business domain to generate semantically complete text block data that meets the requirements of domain ontology annotation. The third processing module is used to generate a standardized domain schema and mold design concept map framework based on the preprocessed text block data. The fourth processing module is used to extract targeted knowledge triples from the full amount of multimodal data with standardized domain schema as hard constraint, simultaneously perform knowledge alignment of terms, referential relations and multimodal content, and perform cleaning, completion and fusion processing on the extraction results to generate standardized knowledge objects with high confidence. The fifth processing module is used to construct a heterogeneous modal knowledge network that links text, images, and 3D models across dimensions based on standardized knowledge objects with high confidence, and to complete cross-modal feature extraction and semantic alignment. At the same time, based on the full-dimensional feature classification system of mold design, it mines and supplements the implicit experience knowledge of mold design from the enterprise's historical design deliverables. The sixth processing module is used to classify standardized knowledge objects and tacit knowledge by dimension, construct a four-dimensional collaborative mold knowledge graph based on the mold design concept graph framework, and complete persistent storage through the iDME engine. At the same time, based on the closed-loop test verification process and incremental data, the four-dimensional collaborative mold knowledge graph is obtained. Among them, based on the four-dimensional collaborative mold knowledge graph, a full business domain service interface for molds is built, and the full-process intelligent service capability corresponding to mold collaborative design is output.

6. The large-model-driven multi-source heterogeneous mold knowledge graph construction system as described in claim 5, characterized in that, The first processing module divides the multi-source heterogeneous raw data from the entire mold business domain into modal types and business processes to obtain a unified knowledge source pool for the mold domain.

7. The large-model-driven multi-source heterogeneous mold knowledge graph construction system as described in claim 6, characterized in that, The second processing module executes differentiated parsing strategies for the raw data of different modalities, extracts the effective content of each modality and converts it into a standardized intermediate format, and generates semantically complete text block data that meets the requirements of domain ontology annotation through semantic slicing.

8. The large-model-driven multi-source heterogeneous mold knowledge graph construction system as described in claim 7, characterized in that, The third processing module, based on the preprocessed text block data, uses the multimodal large model and the large language model in conjunction to complete the construction of a semi-automatic ontology schema, terminology alignment and validation optimization, with the five elements of the mold domain ontology as the core, and generates a standardized domain schema and mold design concept map framework.