Operation method of intelligent atlas analysis and application system based on fusion of resin-based heat-proof material data and large language model

By combining heterogeneous graphs and large language models, the problem of data silos and sparse data utilization in small and medium-sized laboratories is solved, efficient management and intelligent application of material data is realized, and material research and development efficiency and innovation capabilities are improved.

CN120561312APending Publication Date: 2025-08-29SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510630468.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-29

Smart Images

  • Figure CN120561312A_ABST
    Figure CN120561312A_ABST
Patent Text Reader

Abstract

The invention discloses an operation method of an intelligent atlas analysis and application system based on fusion of resin-based heat-proof material data and a large language model, and the operation method specifically comprises the following steps: (1) managing a unified data base and a multi-modal material data text, the data management module is used for efficiently managing laboratory material data texts in various modes; (2) constructing and representing a heterograph of the laboratory material data texts: abstracting the laboratory material data texts of different sources and types into a uniform heterograph structure; and (3) carrying out material data text vectorization and knowledge migration based on graph calculation and a large language model: carrying out high-quality vector embedding representation on nodes in the knowledge graph by utilizing an external pre-training embedding model. According to the method, the problem that the sparse material data text is difficult to utilize by the AI model is effectively solved, and the sparse material data text can be efficiently processed by a large language model and a downstream AI task through knowledge migration from the knowledge graph to the data graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the application of artificial intelligence technology in the field of materials science, and in particular to a method and system for intelligent analysis, efficient retrieval, knowledge discovery and application of material data by integrating laboratory material data text, large-scale language model (LLM) and graph computing technology. Background Art

[0002] The rapid development of artificial intelligence (AI), particularly large language models, is bringing unprecedented transformative opportunities to various industries. The field of materials science has accumulated a vast amount of experimental data, literature, and composition-property relationships. While this data holds immense value, its complex structure and diverse modalities make efficient use of this data for new materials research and development, performance prediction, and knowledge discovery a key challenge in this field.

[0003] Currently, the industry's application of graph data is primarily focused on knowledge graphs, such as leveraging knowledge graphs to provide enhanced contextual information for large language models (e.g., GraphRAG). The advantage of these knowledge graphs is that their nodes and edges often contain rich textual or semantic information, which can be well integrated with text-based LLMs. However, many companies and laboratories possess large amounts of data graphs, such as the "user-destination" graphs of automotive service companies and the "user-product" graphs of internet sales companies. These data graphs are characterized by sparse internal features of nodes and edges, low information density, and difficulty in direct textualization. Their core value lies in the relationships between data. Traditional methods for processing data graphs, such as machine learning-based graph embedding or graph neural networks (GNNs), can enhance node and edge representations through representation learning. However, these advanced technologies present numerous challenges for small and medium-sized laboratories, including the requirement for large amounts of data for model training but limited business volume; high deployment and application costs with uncertain return on investment; and the need for specialized algorithm teams, resulting in a high technical barrier to entry. This makes the technology more suitable for large enterprises or laboratories with greater resource investment capabilities. The increasingly significant differences in application and technical approaches between "knowledge graphs" and "data graphs" make it difficult for laboratories to fully utilize the value of graph data. This is especially true for small and medium-sized laboratories with limited resources, who struggle to reap the benefits of combining graph technology with large language models.

[0004] Traditional methods for managing and analyzing data on resin-based thermal insulation materials typically rely on structured databases and manual experience, making them difficult to cope with the growing scale and complexity of data. In recent years, some research has begun to apply machine learning to materials science, such as predicting material properties through supervised learning or extracting information from literature using natural language processing techniques. However, these methods often suffer from the following shortcomings:

[0005] (1) The data island problem is significant: laboratory resin-based heat-resistant material data (such as experimental records, formula parameters, performance characterization data, literature abstracts, material safety data sheets, etc.) from different sources and in different formats are stored in a scattered manner, lacking a unified management and effective correlation analysis mechanism.

[0006] (2) Feature engineering is complex and relies on expert knowledge: Traditional machine learning methods are highly dependent on manually designed features. For complex material systems, feature extraction is difficult and requires deep domain knowledge.

[0007] Limited ability to process unstructured and semi-structured data: A large amount of laboratory resin-based heat-resistant material information exists in unstructured or semi-structured forms such as text and charts. Traditional methods find it difficult to directly utilize the rich semantics contained in this information.

[0008] (3) Insufficient model generalization and knowledge transfer capabilities: Models trained for specific material systems are often difficult to directly apply to new material systems, and the efficiency of knowledge transfer and reuse is low.

[0009] (4) The integration with large language models is still immature: Although large language models have performed well in text understanding and generation, how to deeply integrate them with professional laboratory material data texts to fully realize their potential in materials research and development is still in the exploratory stage. In particular, for those material data texts with sparse internal features and relatively low information density but complex associations (for example, records containing only experimental numbers, sample batches, test conditions, and numerical results), LLMs have difficulty in directly and effectively processing and understanding their deep meaning and internal connections. Summary of the Invention

[0010] The present invention aims to overcome existing issues such as low utilization of laboratory material data, difficulty integrating different types of data, and difficulty directly applying large language models to sparse data in specialized fields. The present invention provides an operating method for an intelligent graph analysis and application system based on the integration of resin-based heat-resistant material data with a large language model. Specifically, the present invention aims to achieve the following goals:

[0011] (1) Unified management and association of laboratory resin-based heat-resistant material data: Build a unified platform to integrate, clean, structure and associate laboratory resin-based heat-resistant material data texts from different experiments, different characterization methods and different recording formats, break down data silos and form an interconnected material data network.

[0012] (2) Improving the representation capability of sparse material data: For a large number of laboratory material data texts with sparse internal features and little direct text information (such as experimental records containing only numbers, parameters, and results), innovative association and transfer learning mechanisms are used to enhance the semantic richness of their vector representations so that they can be effectively utilized by large language models and downstream AI tasks.

[0013] (3) Realize the deep integration of large language models and professional material data: enable large language models to directly retrieve laboratory material data through learned enhanced vector representations, perform complex queries, analysis, reasoning, and content generation without relying on complex traditional query languages ​​or manual rule construction, and simplify the application of LLM in the field of materials science.

[0014] (4) Enabling intelligent material R&D applications: Based on the integrated data and model capabilities, it provides material scientists and engineers with intelligent material performance prediction, new material design, experimental plan recommendation, scientific research literature analysis, knowledge question and answer and other application services to improve R&D efficiency and accelerate the innovation process.

[0015] To achieve the above objectives, the present invention proposes an operating method of an intelligent graph analysis and application system based on the fusion of resin-based heat-resistant material data and a large language model, which specifically includes the following steps:

[0016] (1) Management of unified data base and multimodal material data text: The system adopts a unified data base for efficient management of multimodal laboratory material data text. The base is composed of multiple types of databases and a unified data interaction interface OpenDAL;

[0017] (2) Construct and represent the heterogeneous graph of laboratory material data text: abstract the laboratory material data text of different sources and types into a unified heterogeneous graph structure, which includes two parts: knowledge graph nodes and data graph. A possible heterogeneous graph process is as follows: Figure 2 shown.

[0018] (3) Material data text vectorization and knowledge transfer based on graph computing and large language models: Aims to enhance the semantic representation of "associative data text". In the knowledge graph, high-quality vector embedding representations of nodes in the knowledge graph are performed using external pre-trained embedding models (such as Word2Vec, BERT's material domain fine-tuned version, or directly using the embedding interface of a powerful general large language model). These vectors can capture their rich semantic information and are stored in a vector database on a unified data base.

[0019] As a preferred embodiment, the mechanism of knowledge migration is: using graph algorithms to realize the migration of vector representations of knowledge graph nodes to data graph nodes, and by learning the complex connection relationships between different types of nodes in heterogeneous graphs, the embeddings on the knowledge graph nodes are propagated along the structure of the graph and generalized to the data graph nodes connected to them, which themselves lack detailed text descriptions.

[0020] As a preferred embodiment, a specific data graph node, although its own description may be simple (such as "sample ID: S001"), is connected to the material component node, experimental method concept node and performance indicator concept node with rich semantics through edges. Through the graph algorithm, the semantic information of these connected nodes can be effectively aggregated and transmitted to the sample node, thereby generating an enhanced, context-rich vector representation. This process can be iterated so that all nodes in the entire laboratory material data text heterogeneous graph obtain high-quality vector representations. The migrated node embedding vectors are also stored in the vector database for subsequent efficient vector recall and retrieval.

[0021] As a preferred embodiment, the knowledge transfer graph includes the following three algorithms:

[0022] Assume that the heterogeneous graph data is A knowledge graph node set The external embedding model has been assigned, the dimension size is d, and the embedding vector is Now we need to set its adjacent data graph nodes Conduct knowledge transfer:

[0023] (1) Graph Neural Network uses message passing algorithm for supervised embedding learning, which is suitable for When a target feature is present, the embedding representation associated with the target feature can be obtained. The formula is as follows:

[0024]

[0025] represents the node i at the lth layer. If l = 0, it represents the initial node, and if l = n, it represents the final embedding, where n is the total number of message passing layers. i,j represents the edge from node i to node j, Represents node i in the knowledge graph node set Neighbor node embedding in, MSG (Message), AGG (Aggregate), COM (Combine) are permutation and other variable operations responsible for message construction, message aggregation and feature update operations respectively. After the message is passed, the model uses Predict the target properties of the node;

[0026] (2) Graph Embedding uses the random walk method to unsupervisedly learn knowledge graph node embedding, which can ensure that the knowledge graph and the data graph are in the same vector space. In the subgraph under the relationship, walking sampling is performed. Assume that the positive sampling sequence obtained by walking path sampling is in is a node in the knowledge graph, whose node embedding has been previously obtained from an external pre-trained embedding model, It is a node in the data graph. Its node embedding can be represented by multi-layer nonlinear mapping using existing features (implemented using a multi-layer perceptron (MLP)) or by embedding its number using an embedding layer. In addition, negative sampling is required, that is, selecting 2m sequences that are not on the same path. The training goal of graph embedding is to make the embeddings of positively sampled sequences as similar as possible and the embeddings of negatively sampled sequences as dissimilar as possible. Cosine similarity is used to measure embedding similarity.

[0027] (3) The method based on the preset formula uses the characteristics of vectors, that is, for a set of vectors in a given vector space, its arithmetic mean vector (that is, the mean vector of each dimensional component) is the center point with the smallest sum of squared Euclidean distances to all vectors in the set. So directly The embedding vector of is assigned as:

[0028]

[0029] Where M(i) represents the node Neighbor nodes, ω j To associate relevance, it can be set in advance. The advantage of this method is that it is fast and can calculate billions of knowledge graph embeddings in a few days.

[0030] As a preferred embodiment, the data text in step (1) includes: structured data, semi-structured data, unstructured text data, vector data, graph data, and system metadata. The specific contents of these data are shown in Table 1.

[0031] As a preferred embodiment, the framework of the system in step (1) provides users with an AI application service layer for deploying service applications such as Agent; the system framework also opens up the interface calls of external knowledge bases and external large models to facilitate data retrieval and the use of large models. The framework architecture diagram is shown in FIG. Figure 1 shown.

[0032] As a preferred embodiment, in step (2), the data text parts with relatively rich semantic information (for example, detailed information extracted from literature abstracts, material description texts, patents, and material concept nodes and relationships containing rich text descriptions obtained through manual annotation or LLM preliminary processing) are regarded as the system's knowledge graph, while those parts with sparse internal features and less direct text information (for example, experimental record nodes, sample nodes, pure performance numerical nodes, etc. that only contain IDs, numerical values ​​and simple labels) are regarded as data graphs. These two types of data texts together constitute the laboratory material heterogeneous graph.

[0033] Compared with the existing technology, this invention can produce the following beneficial effects by deeply integrating laboratory material data text with a large language model and introducing an innovative knowledge transfer mechanism:

[0034] (1) Significantly improve the utilization efficiency and value density of laboratory material data text: By constructing a unified heterogeneous graph and realizing semantically enhanced vectorized representation, the originally isolated and sparse laboratory data text can be effectively associated and deeply mined, releasing its potential value.

[0035] (2) Effectively solve the problem that sparse material data text is difficult to be used by AI models: By migrating knowledge from knowledge graphs to data graphs, even experimental records and sample data that lack direct text descriptions can obtain high-quality semantic representations, which can be efficiently processed by large language models and downstream AI tasks.

[0036] (3) Significantly lower the threshold for the application of large language models in professional fields: enabling large language models to interact and reason with complex laboratory material data texts directly through vector retrieval and calculation, avoiding the need to build complex parsing rules or train dedicated small models for specific tasks, and accelerating the application of large language models in the field of materials science.

[0037] (4) Significantly improve the intelligence level and efficiency of material research and development: The functions provided, such as intelligent retrieval, performance prediction, new material design, and experimental scheme optimization, can assist scientific researchers in obtaining information, making decisions, and discovering innovations more quickly, thereby shortening the R&D cycle and reducing R&D costs.

[0038] (5) Good scalability and adaptability: The system architecture is flexible and can easily access new data sources, integrate new AI models and algorithms, and expand new application scenarios to meet the needs of the rapid development of materials science. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a schematic diagram of the architecture structure of the intelligent graph analysis and application system based on the fusion of resin-based heat-proof material data and a large language model.

[0040] Figure 2 This is a flow chart of a heterogeneous graph implementation in an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The following is a detailed description of an embodiment of the present invention in conjunction with the accompanying drawings: This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and specific operation process are given, but the protection scope of the present invention is not limited to the following embodiment.

[0042] An operating method of an intelligent graph analysis and application system based on the fusion of resin-based heat-resistant material data and a large language model, specifically comprising the following steps:

[0043] (1) Manage unified data base and multimodal material data text: The system adopts a unified data base for efficient management of laboratory material data text of multiple modalities. The base is composed of multiple types of databases and a unified data interaction interface OpenDAL; the data text in step (1) includes: structured data, semi-structured data, unstructured text data, vector data, graph data, and system metadata. The specific contents of these data are shown in Table 1. The framework of the system in step (1) provides users with an AI application service layer for deploying service applications such as Agent; the system framework also opens up the interface call of external knowledge bases and external large models to facilitate data retrieval and the use of large models. The framework architecture diagram is shown in the figure below. Figure 1 shown.

[0044] Table 1

[0045]

[0046] (2) Construct and represent a heterogeneous graph of laboratory material data texts: Laboratory material data texts of different sources and types are abstracted into a unified heterogeneous graph structure, which includes two parts: knowledge graph nodes and data graphs. In step (2), the data text parts with relatively rich semantic information (for example, detailed information extracted from literature abstracts, material description texts, patents, and material concept nodes and relationships containing rich text descriptions obtained through manual annotation or LLM preliminary processing) are regarded as the knowledge graph of the system, while those parts with sparse internal features and less direct text information (for example, experimental record nodes, sample nodes, pure performance numerical nodes, etc. that only contain IDs, numerical values ​​and simple labels) are regarded as data graphs. These two types of data texts together constitute the laboratory material heterogeneous graph. A possible heterogeneous graph process is as follows: Figure 2 shown.

[0047] (3) Material data text vectorization and knowledge transfer based on graph computing and large language models: Aims to enhance the semantic representation of "associative data text". In the knowledge graph, high-quality vector embedding representations of nodes in the knowledge graph are performed using external pre-trained embedding models (such as Word2Vec, BERT's material domain fine-tuned version, or directly using the embedding interface of a powerful general large language model). These vectors can capture their rich semantic information and are stored in a vector database on a unified data base.

[0048] Example 1: Screening and performance prediction of new catalyst materials:

[0049] Data preparation and heterogeneous graph construction: Collect experimental data texts on different metal oxide catalysts in the laboratory, including:

[0050] Knowledge graph: Descriptive text on the types, crystal structures, surface properties, and catalytic mechanisms of metal oxides in relevant published literature; detailed descriptions of the physicochemical properties of known efficient catalysts (such as TiO2, ZnO).

[0051] Data Map: This data contains historical laboratory experimental records, including sample numbers, the precise ratios of the metal oxide components (e.g., 0.5 mol% CuO, 1.0 mol% CeO2 loaded on Al2O3), preparation process parameters (calcination temperature and time), catalytic reaction conditions (reactants, temperature, pressure), and the corresponding catalytic activity and selectivity results. This data is constructed into a heterogeneous graph, with nodes including metal oxide entities, samples, process parameters, reaction conditions, performance indicators (activity, selectivity), and literature concepts.

[0052] Vectorization and knowledge transfer: Use the text-embedding-3-small model in ChatGPT to vectorize literature description text and description text of known efficient catalysts to generate initial vectors for knowledge graph nodes.

[0053] Using graph neural networks as a knowledge transfer module, the knowledge of these core semantic vectors is propagated along the edges of the heterogeneous graph ("such as sample X contains component A", "sample X is prepared under process Y", "performance Z is measured by sample X under condition W") to specific experimental sample nodes, ratio parameter nodes and performance result nodes.

[0054] Model training and application:

[0055] The embedding of data graphs can be used for both property prediction in machine learning and vector retrieval in question-answering systems.

[0056] Based on the enhanced sample node vector (containing rich semantic and association information) and the corresponding catalytic performance values, a deep neural network model is trained to predict catalytic activity and selectivity under different component ratios and process parameters.

[0057] A vector-based question-answering system could also be built, allowing researchers to use natural language queries such as: "Which cerium-based oxide catalysts exhibit high activity for CO oxidation at low temperatures (<200°C) and are reported in the literature to be resistant to sulfur poisoning?" The agent system would automatically find knowledge graph nodes matching the query and search the vector database for the data graph nodes closest to the knowledge graph node. Alternatively, the agent system could modify the user's initial prompt word briefly, directly vectorize it using the embedding model, and then query the vector database.

[0058] Through knowledge transfer, this example enables sample nodes of novel catalyst formulations with only a small number of experimental data points to obtain a superior semantic representation, improving the accuracy and generalization of the performance prediction model. The intelligent question-answering system can quickly filter catalyst information that meets complex criteria from a large amount of internal experimental data and related literature, assisting researchers in quickly identifying promising candidate materials and reducing the number of blind experiments. Querying a database of tens of millions of vectors takes less than 100 milliseconds, meeting the time requirements for rapid responses.

[0059] Example 2: Polymer material literature knowledge mining and auxiliary formulation optimization:

[0060] Data preparation and heterogeneous graph construction:

[0061] Knowledge graph: collects a large number of academic paper abstracts in the field of polymer science, descriptive texts about monomers, initiators, additives, polymerization processes, and material properties (such as tensile strength, elongation at break, and glass transition temperature Tg) in patent literature; and physical property manuals of common polymer materials (such as polypropylene (PP) and polyvinyl chloride (PVC).

[0062] Data Map: Internal polymer modification experiment records, including the types and amounts of various additives (plasticizers, antioxidants, fillers), blending process parameters (temperature, screw speed), and corresponding product performance test data. Construct a heterogeneous graph with nodes such as monomers, initiators, additives, polymers, process parameters, performance indicators, and literature / patent IDs.

[0063] Vectorization and knowledge transfer:

[0064] Use the BERT model pre-trained on material text data to perform high-quality vectorization of core semantic data text.

[0065] Knowledge transfer is achieved through graph embedding, which transfers the deep semantic knowledge in literature and physical property manuals to specific internal experimental formula nodes and performance nodes.

[0066] Model training and application:

[0067] Train a multi-objective optimization model, input the desired product performance range (such as tensile strength > 50MPa, Tg at 80-100℃), and output the recommended additive combination and dosage range.

[0068] Develop a literature knowledge discovery tool. When researchers enter a new adjuvant name or encounter a performance problem, the system can automatically retrieve relevant mechanism analysis, solutions or potential synergistic / antagonistic effect information from the associated literature vector.

[0069] The recipe optimization model in this example can recommend optimized recipes that are difficult to discover through traditional experience based on complex interactions learned from large amounts of data and text. This literature knowledge discovery tool significantly improves researchers' efficiency in accessing and understanding literature. When an experiment produces an abnormal result, simply inputting a description of the abnormal phenomenon will quickly link to relevant literature explaining similar phenomena or analyzing potential causes, assisting in rapid diagnosis of the problem.

[0070] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. An operating method of an intelligent graph analysis and application system based on the fusion of resin-based heat-resistant material data and a large language model, characterized in that: The specific steps include: (1) Management of unified data base and multimodal material data text: The system adopts a unified data base for efficient management of multimodal laboratory material data text. The base is composed of multiple types of databases and a unified data interaction interface OpenDAL; (2) Constructing and representing a heterogeneous graph of laboratory material data text: abstracting laboratory material data texts of different sources and types into a unified heterogeneous graph structure, which includes two parts: knowledge graph nodes and data graph; (3) Material data text vectorization and knowledge transfer based on graph computing and large language models: Aims to enhance the semantic representation of "associative data text". In the knowledge graph, an external pre-trained embedding model is used to perform high-quality vector embedding representation of the nodes in the knowledge graph. These vectors can capture their rich semantic information and are stored in the vector database of the unified data base.

2. The method for operating an intelligent graph analysis and application system based on the fusion of resin-based heat-proof material data and a large language model according to claim 1, characterized in that: The mechanism of knowledge transfer is as follows: using graph algorithms to realize the migration of vector representations of knowledge graph nodes to data graph nodes. By learning the complex connection relationships between different types of nodes in heterogeneous graphs, the embeddings on knowledge graph nodes are propagated along the structure of the graph and generalized to the data graph nodes connected to them that lack detailed text descriptions.

3. The method for operating an intelligent graph analysis and application system based on the fusion of resin-based heat-resistant material data and a large language model according to claim 1, characterized in that: Although the description of a specific data graph node itself may be simple, it is connected to the material component node, experimental method concept node and performance indicator concept node with rich semantics through edges. Through graph algorithms, the semantic information of these connected nodes can be effectively aggregated and transmitted to the sample node, thereby generating an enhanced, context-rich vector representation. This process can be iterated so that all nodes in the entire laboratory material data text heterogeneous graph obtain high-quality vector representations. The migrated node embedding vectors are also stored in the vector database for subsequent efficient vector recall and retrieval.

4. The method for operating an intelligent graph analysis and application system based on the fusion of resin-based heat-resistant material data and a large language model according to claim 2, characterized in that: The knowledge transfer graph includes the following three algorithms: Assume that the heterogeneous graph data is A knowledge graph node set The external embedding model has been assigned, the dimension size is d, and the embedding vector is Now we need to set its adjacent data graph nodes Conduct knowledge transfer: (1) Graph Neural Network uses message passing algorithm for supervised embedding learning, which is suitable for When a target feature is present, the embedding representation associated with the target feature can be obtained. The formula is as follows: represents the node i at the lth layer. If l = 0, it represents the initial node, and if l = n, it represents the final embedding, where n is the total number of message passing layers. i,j represents the edge from node i to node j, Represents node i in the knowledge graph node set Neighbor node embedding in, MSG, AGG, COM are permutation and other variant operations responsible for message construction, message aggregation and feature update operations respectively. After message passing, the model uses Predict the target properties of the node; (2) The graph embedding algorithm uses the random walk method to unsupervisedly learn the knowledge graph node embedding, which can ensure that the knowledge graph and the data graph are in the same vector space. In the subgraph under the relationship, walking sampling is performed. Assume that the positive sampling sequence obtained by walking path sampling is in is a node in the knowledge graph, whose node embedding has been previously obtained from an external pre-trained embedding model, It is a node in the data graph. Its node embedding can be represented by multi-layer nonlinear mapping using existing features or by embedding its number using an embedding layer. In addition, negative sampling is required, that is, 2m sequences that are not on the same path are selected. The training goal of graph embedding is to make the embeddings of positive sampling sequences as similar as possible and the embeddings of negative sampling sequences as dissimilar as possible. Cosine similarity is used to measure the embedding similarity; (3) The method based on the preset formula uses the characteristics of vectors, that is, for a set of vectors in a given vector space, its arithmetic mean vector is the center point with the smallest sum of squared Euclidean distances to all vectors in the set, so directly The embedding vector of is assigned as: Where M(i) represents the node Neighbor nodes, ω j To associate relevance, it can be set in advance. The advantage of this method is that it is fast and can calculate billions of knowledge graph embeddings in a few days.

5. The method for operating an intelligent graph analysis and application system based on the fusion of resin-based heat-proof material data and a large language model according to claim 1, characterized in that: The data text in step (1) includes: structured data, semi-structured data, unstructured text data, vector data, graph data, and system metadata.

6. The method for operating an intelligent graph analysis and application system based on the fusion of resin-based heat-proof material data and a large language model according to claim 1, characterized in that: The system framework in step (1) provides users with an AI application service layer for deploying service applications such as Agent; the system framework also opens up interface calls to external knowledge bases and external large models to facilitate data retrieval and the use of large models.

7. The method for operating an intelligent graph analysis and application system based on the fusion of resin-based heat-proof material data and a large language model according to claim 1, characterized in that: In step (2), the data text part with relatively rich semantic information is regarded as the knowledge graph of the system, while the part with sparse internal features of the nodes and less direct text information is regarded as the data graph. These two types of data text together constitute the laboratory material heterogeneous graph.