Fiber master batch formula under intelligent agent technology and product development method

By building a localized intelligent agent system based on the RAGFlow framework, combining multimodal input embedding modules and large models, and integrating the enterprise knowledge base, the problems of low efficiency, insufficient accuracy and poor data security in traditional fiber masterbatch formula development have been solved, and efficient and secure formula generation and optimization have been achieved.

CN120690341APending Publication Date: 2025-09-23POLY PLASTIC MASTERBATCH SUZHOU
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510735870.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The traditional fiber masterbatch formula development model is inefficient, has high trial-and-error costs, is difficult to integrate corporate knowledge resources, has insufficient accuracy in general AI models, has poor data security, and is unable to meet application requirements in highly sensitive scenarios.

Method used

Build a localized intelligent agent system based on the RAGFlow framework, combine multimodal input embedding modules and localized large models, integrate enterprise knowledge bases, and use knowledge base retrieval enhancement and large model dynamic collaboration to ensure data security and accuracy.

Benefits of technology

It significantly improves the efficiency and accuracy of fiber masterbatch formula development, ensures data security, generates formulas that meet actual production needs, reduces R&D costs, and enhances the scalability and flexibility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690341A_ABST
    Figure CN120690341A_ABST
Patent Text Reader

Abstract

The invention relates to the crossing field of artificial intelligence and new material research and development, and provides a fiber master batch formula and product development method based on an intelligent agent technology, and the method depends on a large model deployed locally, constructs an RAGFlow framework, accesses a structured internal professional knowledge base, and improves the development efficiency of a fiber master batch. And the automatic whole process of intelligent optimization, performance prediction and product research and development of the fiber master batch formula is realized. According to the system, multi-modal heterogeneous information such as material formula composition, processing technology parameters, experimental performance data, historical production records and market demands is fused, and knowledge calling efficiency is improved through vectorized storage and enhanced retrieval strategies; the retrieval and generation process is optimized based on a multi-path recall and dynamic reordering technology, so that the masterbatch formula recommendation is more accurate, and the reliability of output content is ensured by means of a generation constraint mechanism; the development efficiency, the optimization precision and the data security of a master batch formula are remarkably improved, meanwhile, the trial and error cost is reduced, and the method is suitable for intelligent development of high polymer materials and fiber master batches and meets the requirements of high-sensitivity industries for data privacy protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the intersection of artificial intelligence and new material research and development, and specifically relates to a method for intelligently generating fiber masterbatch formulas based on a RAGFlow framework built on a local large model. Background Art

[0002] Fiber masterbatch is a core functional material in chemical fiber production. Made by melt-blending a base resin (such as polyester or polypropylene), functional additives (colorants, flame retardants, conductive agents, etc.), and a dispersant, it is widely used in textiles and apparel, medical protective equipment, industrial filtration, and other fields. Its formulation directly impacts the mechanical properties, functional characteristics, and processing feasibility of the final fiber. With the continuous escalation of market demand, the need for masterbatch development is becoming increasingly urgent. However, the traditional human-based development model faces multiple technical bottlenecks, resulting in inefficient development and escalating costs.

[0003] First, the trial-and-error cost of formula development is extremely high. Every new formula development requires multiple steps such as raw material screening, blending experiments, and spinning verification. The entire process often takes several days, which not only prolongs the development cycle but also increases resource consumption. In addition, when designing a formula, engineers must consider multiple factors such as functional performance, processing feasibility, and cost constraints. However, due to the trade-offs between these goals, it is difficult to find a true Pareto optimal solution. At the same time, the experimental data and production experience accumulated by enterprises over a long period of time are scattered and stored in multiple independent systems. The lack of a unified structured knowledge graph leads to low utilization of historical data and inability to effectively integrate and use it to optimize formula development.

[0004] At the same time, the implementation of intelligent applications also faces numerous challenges. Because general AI models lack specialized training for fiber masterbatches, their error rate in formula prediction tasks can reach as high as 32%, far exceeding the accuracy requirements for engineering applications. Furthermore, data security issues have become a key obstacle to the promotion of intelligent applications. Companies' core data (including raw material ratios and process parameters) is subject to strict privacy protection restrictions, making it difficult to directly use it for model training, further increasing the difficulty of intelligent R&D.

[0005] To address these issues, a locally deployed intelligent agent system is urgently needed. While ensuring data security, it enhances the efficiency and accuracy of fiber masterbatch formulation development through domain knowledge augmentation and multi-objective dynamic optimization. This system effectively integrates a company's knowledge resources and promotes data-driven R&D while also avoiding the privacy risks associated with cloud-based deployments and meeting the demands of highly sensitive applications. Therefore, combining a local knowledge base with a locally deployed large model to develop an intelligent agent-based fiber masterbatch formulation and product development method has significant practical application value. Summary of the Invention

[0006] The purpose of the present invention is to solve the above-mentioned existing research and development difficulties and provide a fiber masterbatch formula and product development method under intelligent body technology.

[0007] A fiber masterbatch formulation and product development method based on intelligent agent technology includes the following steps:

[0008] (a) Build a high-performance computing platform to provide computing power support for large-scale data processing and intelligent agent reasoning;

[0009] (b) building an intelligent agent system based on modular services and a locally deployed large model, wherein the intelligent agent system uses a knowledge base search enhancement and a dynamic collaboration mechanism with the locally deployed large model;

[0010] (c) Build and analyze a specialized knowledge base containing data related to fiber masterbatch formulations;

[0011] (d) The intelligent agent system retrieves and enhances the user input to achieve intelligent recommendation, optimization and product development of fiber masterbatch formula.

[0012] The locally deployed large model includes a multimodal input embedding module, which is used to uniformly convert input information of different modalities, including text descriptions, chemical structures, and numerical attributes, into a dense vector representation that can be processed by the Transformer layer of the large model.

[0013] The steps of constructing the multimodal input embedding module and fusing information include:

[0014] (a) Using a chemical structure encoder, the SMILES string of the chemical component is converted into a molecular graph representation, which is then input into a graph neural network to extract a graph-level structural embedding vector representing the entire molecule;

[0015] (b) Using a numerical attribute encoder, a vector consisting of a series of numerical molecular descriptors or physical properties of the chemical components is input into a multi-layer perceptron network to obtain an attribute embedding vector;

[0016] (c) When preparing an input sequence for the large model, the text information is first segmented and processed through a word embedding layer, and then a special tag is introduced for the chemical entity in the input sequence, and the structure embedding vector obtained in step (a) and the attribute embedding vector obtained in step (b) respectively replace the default embedding of the special tag of the corresponding chemical entity, thereby forming a sequence that integrates text embedding, structure embedding and attribute embedding, and finally position encoding is applied to the integrated sequence and sent to the first Transformer layer of the large model.

[0017] The graph neural network in the chemical structure encoder (a) is selected from at least one of a graph convolutional network (GCN), a graph attention network (GAT) or a message passing neural network (MPNN); and the step of extracting a graph-level structure embedding vector representing the entire molecule further includes applying a global pooling operation to integrate all atomic embeddings after all layers of the graph neural network have processed the molecular graph representation; the SMILES string of the chemical component in step (a) specifically includes a SMILES string of a base resin and at least one functional additive constituting the fiber masterbatch formula, and the functional additive is selected from at least one of a pigment, a dispersant or an antioxidant.

[0018] Before inputting the vector composed of a series of numerical molecular descriptors or physical properties of the chemical components into the multi-layer perceptron network, the numerical attribute encoder (b) also includes a step of preprocessing the composed vector, and the preprocessing step includes splicing the numerical features into a single vector and performing normalization or standardization on the single vector; the series of numerical molecular descriptors or physical properties of the chemical components in the step (b) specifically include at least one item selected from the molecular weight, topological polar surface area, octanol-water partition coefficient LogP, melting point, glass transition temperature, density or solubility parameter used to characterize the fiber masterbatch component.

[0019] When preparing the input sequence for the large model in step (c), the text information includes a specific description of the fiber masterbatch formula, which at least includes the type of base resin used in the fiber masterbatch formula, the type of one or more additives and their amount or percentage concentration in the formula, and the special mark corresponds to these chemical entities in the fiber masterbatch formula, so that the integrated sequence can characterize the composition and structural characteristics of the fiber masterbatch for subsequent performance prediction or formula optimization tasks.

[0020] The step of converting the SMILES string into a molecular graph representation includes: defining the nodes of the molecular graph as atoms and the edges as chemical bonds; and, the atomic features include at least one of atom type, charge and hybridization mode, and the chemical bond features include the type of chemical bond; when the graph neural network is a graph convolutional network (GCN), the GCN comprises four layers, each layer updates the vector representation of the current atom by aggregating information of neighboring atoms, and uses a ReLU activation function between layers; the dimension (d_struct) of the graph-level structure embedding vector is configured to be the same as the internal representation dimension (d_model) of the large model or can be projected to the d_model dimension through a linear layer; and, the chemical structure encoder is pre-trained on a large molecular dataset for molecular property prediction or self-supervised learning tasks; the multilayer perceptron network comprises 2-6 hidden layers and uses a ReLU activation function.

[0021] The dimension (d_numerical) of the attribute embedding vector is configured to be the same as the internal representation dimension (d_model) of the large model or to be projected to the d_model dimension via a linear layer; and the numerical attribute encoder is jointly trained end-to-end during the overall fine-tuning of the large model;

[0022] The special tags introduced for the chemical entity include a first special tag for indicating a chemical structure embedding insertion position and a second special tag for indicating a chemical attribute embedding insertion position; and the step of replacing the default embedding of the special tags corresponding to the chemical entity includes:

[0023] For the first special marker, replace its default embedding with the structure embedding vector obtained by the chemical structure encoder. If the dimension of the structure embedding vector is inconsistent with the internal representation dimension of the large model, it is first adjusted to the internal representation dimension through a linear projection layer;

[0024] For the second special tag, its default embedding is replaced with the attribute embedding vector obtained by the numerical attribute encoder. If the dimension of the attribute embedding vector is inconsistent with the internal representation dimension of the large model, it is first adjusted to the internal representation dimension through a linear projection layer.

[0025] The steps of constructing and parsing a specialized knowledge base include:

[0026] Collecting multimodal data, wherein the multimodal data includes at least chemical structure information, molecular descriptors, physical properties, formulation components and amounts, process parameters, and corresponding product performance data of chemical components;

[0027] Connecting the multimodal data to the RAGFlow framework for data parsing and structuring, including converting tabular data into natural language descriptions, using visual models for document layout analysis, using OCR recognition and semantic segmentation to generate structured text blocks, and vectorizing the text blocks and storing them together with the original text to build a hybrid index system;

[0028] The steps of building the intelligent agent system include:

[0029] Adopting Docker container management to build an open source model management framework, supporting dynamic switching of large models and parameter fine-tuning;

[0030] Locally deploy open source large language models and provide intelligent reasoning services based on the Ollama framework.

[0031] The steps of the retrieval enhancement process include:

[0032] A hybrid search strategy combining BM25 keyword search and vector semantic search is used;

[0033] Introducing a dynamic re-ranking model to optimize the ranking of search results by combining semantic relevance, time factors, and data weights;

[0034] The steps of output enhancement processing include:

[0035] Adopt a restricted output whitelist mechanism to ensure that the generated answers only reference verified material names, raw materials and process parameters;

[0036] By generating a constraint mechanism, the answer must cite the source document and its specific location, and users can click to jump for verification;

[0037] Set up professional terminology verification rules to detect terminology accuracy based on a regular rule base and automatically trigger regeneration when errors are detected;

[0038] The high-performance computing platform uses physical isolation technology and virtualization technology to ensure the integrity and privacy of data processing and storage. The beneficial effects of the present invention are:

[0039] The fiber masterbatch formulation and product development method based on the intelligent agent technology provided by this invention, through the local deployment of large models and the RAGFlow framework, has significant advantages over traditional development models, which are specifically reflected in the following aspects:

[0040] (1) Improved development efficiency and accuracy: This invention significantly shortens the development cycle of fiber masterbatch formulas through the intelligent collaboration of the local knowledge base and the local large model. The system can quickly automate the process of raw material screening, formula optimization, and process parameter adjustment, avoiding the tedious trial-and-error process, effectively reducing R&D costs and improving development accuracy.

[0041] (2) The model’s ability to process chemical information is expected to be significantly enhanced, and multimodal input enables it to directly perceive and understand the structure and inherent properties of chemical substances.

[0042] (3) Improve data utilization and security: By integrating historical enterprise data and field experience, the system builds a comprehensive knowledge base that can efficiently process and store multimodal data. At the same time, the use of physical isolation and virtualization technologies ensures data security and privacy protection, avoiding the risk of data leakage.

[0043] (4) Avoid output “hallucinations”: The original large model is trained with extensive data and is prone to outputting erroneous information when generating fiber masterbatch formulas, such as fabricating raw material names, chemical structures, or properties, which makes the formula unfeasible. By combining the industry knowledge base and introducing a generation constraint mechanism and a mandatory output whitelist mechanism, a regular rule library can be used to ensure the consistency and professionalism of terminology, and only verified raw materials and processes are allowed to be recommended, so that the final output information has a verified basis, enhances reliability, and provides a formula that conforms to production reality.

[0044] (5) Process, production capacity and cost constraints: The production of fiber masterbatch involves not only formula improvement, but also various constraints such as actual production factors. When generating formulas, ordinary large models often find it difficult to take into account the actual factors in the production process, such as raw material supply, special requirements of the production process and production cost constraints. This makes it impossible to fully consider the actual feasibility and economy when generating formulas, resulting in the generated formula process not being consistent with the actual situation and difficult to implement or the raw material cost being too high. By integrating actual information such as specific production environment into the knowledge base, the system can combine real-time production information and industry knowledge for reasoning and adjust the generated formula to ensure that it not only meets functional requirements but also adapts to actual production conditions. In this way, a formula that is more in line with production needs and has operability can be generated, while effectively controlling costs and improving production efficiency.

[0045] (6) Enhanced scalability and flexibility: Through containerized deployment and local API interfaces, the system can flexibly switch and optimize large model configurations according to business needs, support large-scale data processing, and quickly adapt to market and technological changes.

[0046] The method is applied to the intelligent development of polymer materials and chemical fiber masterbatch products, and can be expanded to fields such as textiles, medical treatment, and industrial filter materials, while meeting data privacy protection needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 The specific structure of the present invention;

[0048] Figure 2This is the parsing process of the knowledge base in the present invention;

[0049] Figure 3 This is the specific analysis and retrieval process after inputting questions on the RAGFlow framework in the present invention;

[0050] Figure 4 Targeted answers for the large model agent connected to the knowledge base in the present invention;

[0051] Figure 5 Answer for ordinary large models. DETAILED DESCRIPTION

[0052] This paper provides a fiber masterbatch formulation and product development method based on intelligent agent technology. This method is based on a locally deployed large model, constructs a RAGFlow framework, and accesses an enterprise knowledge base to achieve intelligent optimization of fiber masterbatch formulations and product development. The specific workflow is as follows:

[0053] The first step is to build a high-performance computing platform that combines advanced hardware architecture to provide powerful computing support, including:

[0054] (1) The computing platform uses physical isolation technology and virtualization technology to ensure the integrity and privacy of data processing and storage, as well as the data security between different computing nodes, to avoid data leakage and unauthorized access;

[0055] (2) Based on a high-performance computing platform and Linux operating system, combined with advanced parallel computing hardware architecture, it provides powerful computing power support for large-scale data processing and intelligent agent reasoning, ensuring system operation efficiency;

[0056] The second step is to build the RAGFlow framework and construct an intelligent agent system based on modular services and local large models. The RAGFlow framework is used to achieve dynamic collaboration between knowledge base knowledge retrieval enhancement and local large models. The implementation steps are as follows:

[0057] (1) Construct an open source model management framework, comprehensively considering the flexibility of model management and the scalability of large models; adopt Docker container management to support dynamic switching and parameter fine-tuning of large models. (2) Multi-architecture compatible model library design, support dynamic loading of large models with different parameter scales (such as LLaMA, DeepSeek, Qwen, etc.), and select the optimal model according to task requirements, and provide unified services through RESTful interfaces; (3) Local deployment of large models, adopt open source large language models and provide intelligent reasoning services based on the Ollama framework, get rid of the traditional cloud dependency model, support large model reasoning with local computing power, ensure reasoning speed, and at the same time ensure data security and avoid sensitive data leakage; (4) Model version snapshot mechanism and management, retain historical fine-tuning parameters when switching or upgrading the model, ensure business continuity, avoid performance degradation caused by model changes, ensure that the new version model inherits the existing optimization, and ensure the continuity and stability of the system; (5) Support dynamic switching and parameter fine-tuning of large models to improve model adaptability, including: (a) model hot loading, It allows seamless switching of large models of different sizes at runtime to adapt to different application scenarios; (b) parameter fine-tuning mechanism, which supports users to fine-tune large models for specific tasks in the local environment to enhance their performance in the subdivided field; (6) supports large model inference optimization, automatically adjusts model parameters in different computing environments, and improves inference efficiency, including: (a) low-latency inference mode, for tasks with high real-time requirements, uses lightweight models (such as DeepSeek-6B) for fast inference to reduce response time; (b) high-precision inference mode, in application scenarios requiring higher precision, uses LLaMA3.2 with a larger parameter scale for deep optimization to ensure the accuracy of the generated results; (c) adaptive computing resource management and dynamic scheduling mechanism, the system can dynamically adjust between lightweight inference and high-precision inference modes according to the computing resource load to optimize computing resource allocation;

[0058] The third step is to build a professional knowledge base and analyze:

[0059] (1) Multimodal data collection, integrating production-related data, including production plans, raw material inventory, order data, transaction records, product preferences, after-sales service logs, production line parameters, production parameters, experimental records, product performance test data and market demand information, and manual voice recording of experience notes of developers and field engineers, etc., to form a comprehensive data foundation and further enrich the knowledge base; further, it is used to enable LLM to not only process text information, but also understand and utilize the microstructure and physical and chemical properties of molecules, so as to more accurately predict the performance of formulas, especially compatibility, dispersion effect and final product performance, and further enhance learning based on the characterization of chemical structure and material properties. By collecting fiber masterbatch formula data from internal enterprises and public literature, each data not only contains the component name and amount, but also contains: chemical structure information: SMILES (Simplified Molecular Input Line Entries) of each component (base resin, key additives such as pigments, dispersants, antioxidants, etc.) System) string or InChIKey; then, using cheminformatics tools (e.g., RDKit, Mordred), a series of molecular descriptors are calculated for each key component, such as molecular weight (Mw), topological polar surface area (TPSA), LogP (octanol-water partition coefficient), number of hydrogen bond donors / acceptors, number of rotatable bonds, number of aromatic rings, etc.; experimentally determined or literature-reported physical properties of key components are incorporated into the input data, such as melting point (Tm), glass transition temperature (Tg), density, solubility parameter, etc.; and corresponding experimentally measured properties of fiber masterbatch or final fiber product are also included, such as tensile strength, MI value, color difference ΔE, dispersion uniformity rating (which can be quantified by electron microscopy images), thermal stability index (such as weight loss temperature in TGA data), etc. Existing material chemistry knowledge graph fragments are reconstructed or utilized, and their entity components must at least include chemical substances (e.g., "polypropylene," "carbon black N330," "zinc stearate"), chemical groups (e.g., "-OH," "-COOH"), and properties (e.g., "high melting point," "good dispersibility"). The About section must contain at least: "is_a" (e.g., "Carbon Black N330 is_a pigment"), "has_property" (e.g., "Polypropylene has_property hydrophobicity"), "may_interact_with" (e.g., "Zinc stearate may_interact_with polyethylene matrix" to improve dispersion), and "functional_group_of" (e.g., "-COOH functional_group_of stearic acid"). A molecular information-enhanced dataset containing at least 500 complete records, including at least 100 different chemical substances and their structures and descriptor information, was constructed. The knowledge graph contained at least 2,000 triples.(2) Connecting data to RAGFlow for data parsing and structural processing involves the following core steps: (a) Encrypting and storing all user-uploaded documents through MinIO and triggering the DeepDoc parsing engine for deep parsing; (b) Converting Excel / PDF tables and other content into natural language descriptions for tabular data; (c) Using visual models to analyze the layout of the document's logical structure to ensure that the document data is complete and logical; (d) Using OCR recognition, layout analysis, and semantic segmentation methods to generate structured text blocks with location information to improve indexing efficiency; (e) Vectorizing data storage, vectorizing the optimized text blocks and storing them together with the original text in ElasticSearch to build a hybrid indexing system. (3) Hierarchical classification of knowledge base content allows manual selection of knowledge bases when calling to search specific fields and enhance local retrieval capabilities. (4) In the process of structural data standardization, ensure the normalization of SMILES strings, normalize molecular descriptors and physical property parameters (to the range of 0-1) or standardize (mean is 0, standard deviation is 1) to eliminate the dimension effect; in the process of fiber masterbatch formula data, all the above information (SMILES, descriptors, physical properties, formula, process, performance) needs to be converted into a unified text sequence format for LLM processing. For example: "Input: base material SMILES: CC(C)C(=O)O[*]C(C)C(=O)O[*]...; base material descriptor: Mw=150000, TPSA=50.2, LogP=...; additive A_SMILES: c1ccccc1; additive A dosage: 5%; ... target performance tensile strength:?"; output: "30MPa". Convert the knowledge graph triples into natural language sentences, such as "Material A is a polymer. Material A has property X."

[0060] The locally deployed large language model described in this invention (such as DeepSeek-R1:32B) is selected as the infrastructure, and its input representation layer and mechanism are optimized. A multimodal input embedding module is used, including a MILES / chemical structure encoder: a molecular encoder based on a graph neural network (GNN, such as a graph convolutional network (GCN) or a graph network (GAT)). The SMILES string is first converted into a molecular graph structure, and then input into the GNN to extract the molecular-level structural feature vector. This vector will serve as the structural embedding of the chemical substance. It also includes a molecular descriptor / property encoder: for numerical molecular descriptors and physical property parameters, they are spliced ​​into a vector and then projected into the same dimensional space as the text and structure embedding through a small multi-layer perceptron (MLP) network; and a text encoder, which uses the LLM's own original word embedding layer to process the text description part of the recipe (such as the natural language description of the component name, the text description of the process conditions, etc.); in addition, the embedding vectors from the SMILES encoder (or GNN), descriptor / property encoder and text encoder are fused into a unified, chemical information-rich initial representation through a splicing or gating mechanism through a fusion layer, and then input into the main Transformer layer of the LLM.

[0061] During fine-tuning, the input is a recipe description containing the chemical structure information, descriptors, physical properties, and dosage of each component, and the corresponding key performance indicators (such as mechanical, thermal, and optical properties) are output. For continuous value prediction, the mean squared error (MSE) or mean absolute error (MAE) can be used as the loss function. Knowledge graph triples are also introduced, allowing the model to determine the correctness of triples or predict missing entities or relations. This helps the model learn deep semantic connections between entities. A cross-entropy loss can be used as the loss function.

[0062] The fourth step is to perform retrieval and output enhancement after specific input: (1) Optimize the retrieval process and improve the efficiency of information matching, including the following steps: (a) Use BM25 keyword retrieval (accounting for 30%-40%) to match relevant document fragments based on user query content; (b) Perform vector semantic retrieval (accounting for 60%-70%), calculate the similarity between the question vector and the knowledge base document fragment, and measure the semantic matching degree for accurate retrieval; (c) Introduce a dynamic reranking model, combine semantic relevance, time factors (timeliness) and data weights to optimize the ranking of retrieval results and improve the quality of retrieval results; (d) Start a hybrid weight dynamic adjustment mechanism to automatically adjust the weights of BM25 and vector retrieval according to different question types to enhance adaptability; (e) Set a time decay factor (weight + 30%) to ensure that the recommended solution conforms to the latest technology trends and improve the timeliness of retrieval results. (2) Optimize the output generation process to ensure the accuracy and reliability of the content, including the following steps: (a) Splice the relevant fragments retrieved by Top-K (e.g., K=40) into prompt words as the context and reference for the large model to generate answers, and use Top-P sampling (e.g., P=0.9) and set temperature coefficient (e.g., 0.7) to optimize the text generation effect; (b) Restrict the output whitelist mechanism (force output whitelist mechanism) to ensure that the generated answers only quote verified material names, raw materials, and process parameters to prevent the generation of false information and ensure the reliability of the generated content; (c) Through a strict generation constraint mechanism, require that the answer must quote the source document and its specific location, and support users to click to jump for verification to enhance traceability; (d) Set professional terminology verification rules, detect the accuracy of terms based on the regular rule library, and automatically trigger regeneration in case of errors to ensure that professional terminology meets industry standards and avoid errors or hallucinations. (3) Support adjustment of various parameters to obtain output that better meets the needs.

[0063] Through the above four steps, this system realizes the complete process from efficient computing power support, intelligent architecture construction, knowledge base optimization to retrieval and output enhancement. It combines knowledge-enhanced retrieval technology to realize masterbatch formula recommendation and optimization, greatly improving the intelligence level of fiber masterbatch formula research and development, and improving data security, retrieval accuracy and the credibility of generated content, providing the industry with a set of efficient, safe and accurate intelligent R&D solutions.

[0064] In the specific implementation process, the specific steps are as follows:

[0065] (1) Build a localized computing power platform:

[0066] The computing environment is built based on the Ubuntu 22.04 operating system, using the NVIDIA GeForce RTX 4090 graphics card as the main computing resource, making full use of its powerful parallel computing capabilities and graphics processing performance, and can provide strong support for complex multimodal data analysis and model reasoning.

[0067] (2) Localized knowledge base:

[0068] In this invention, the knowledge base is constructed based on various relevant papers, online patents, and related technical materials as the core data source. Due to the limited data, the system screens, classifies, and organizes data from different sources, extracting key technical information, material ratios, performance parameters, and process methods, forming only a preliminary knowledge base framework. Later, through the expansion and structuring of this data, it ensures that the knowledge base covers basic theories, mature formulas, experimental data, and the latest patents, providing data support for subsequent intelligent retrieval and formula optimization.

[0069] (3) Local deployment of large models:

[0070] During the implementation of this invention, the system implemented a RAGFlow-based agent architecture using the Ollama framework, successfully deploying multiple large local models to achieve efficient knowledge retrieval and dynamic collaboration among these models. Specifically, these models include the DeepSeek-R1:32B chat model, the Mxbai-Embed-Large:Latest embedding model, and the LLaMA3.2-Vision image-to-text conversion model, responsible for semantic understanding, text vectorization, and image parsing, respectively. All models provide services through a RESTful API.

[0071] (4) Build the RAGFlow framework and connect the knowledge base with the large model to analyze the knowledge base:

[0072] (4.1)RAGFlow deployment and model connection

[0073] a) Use Docker container to deploy RAGFlow and use local port 80 for service management;

[0074] b) Set the URL of the local large model to host.docker.internal:11434 so that RAGFlow can access the local model through the RESTful API.

[0075] (4.2) Chat Assistant Optimization Configuration

[0076] a) Set the temperature coefficient to 0.7 to balance the diversity and stability of the answers;

[0077] b) The maximum generated length is set to 4096 to ensure that the generated content is complete and not limited in length;

[0078] c) Top-K sampling (K=40) and Top-P sampling (P=0.9) were used to improve the consistency and rationality of the answers.

[0079] (4.3) Improvement of LLM model architecture

[0080] In order to improve the understanding and reasoning ability of the LLM model in the field of chemistry, a multimodal input embedding module is constructed to accept and integrate diverse information such as text, chemical structure and numerical attributes.

[0081] The multimodal input embedding module is constructed to uniformly convert input information from different modalities, including text descriptions, chemical structures, and related numerical attributes, into a dense vector representation that can be processed by the LLM's Transformer layer. The original LLM input stage is modified to prepare specialized encoders for different types of non-text input. This includes a chemical structure encoder, which is typically built based on a graph neural network (GNN). Its input is a SMILES string of chemical components. During the preprocessing step, the SMILES string is converted into a molecular graph representation, where nodes represent atoms and edges represent chemical bonds. Atomic features can include atom type, charge, and hybridization, while chemical bond features include bond type. The GNN architecture uses a graph convolutional network (GCN). Taking a typical four-layer GCN as an example, each layer updates the vector representation of the current atom by aggregating information from neighboring atoms, and ReLU activation functions are used between layers. After all GNN layers are processed, a global pooling operation is applied to combine all atom embeddings to obtain a fixed-size graph-level embedding vector representing the entire molecule. This GNN encoder is designed or fine-tuned to output an embedding vector of dimension d_struct, ideally the same as the LLM model's internal representation dimension d_model. During training, this GNN module is pre-trained on large molecular datasets (such as the QM9 dataset) for molecular property prediction or self-supervised learning tasks. The numerical attribute encoder is typically implemented using a multi-layer perceptron (MLP). Its input is a set of numerical molecular descriptors or physical properties of chemical components, such as molecular weight (Mw), topological polar surface area (TPSA), lipid-water partition coefficient (LogP), melting point (Tm), and glass transition temperature (Tg). These values ​​are first aggregated into a single vector. Preprocessing steps include concatenating these numerical features into a single vector and performing any necessary normalization or standardization. The MLP architecture consists of three hidden layers using the ReLU activation function. The size of the input layer corresponds to the number of numerical features, while the output layer has the dimension d_numerical, which is also expected to be consistent with or projectable onto d_model. The numerical attribute encoder is jointly trained end-to-end during the overall fine-tuning of the LLM.

[0082] Preparing the integrated input sequence for LLM requires integrating text, chemical structure, and numerical property information. For example, consider a formulation design task: "Design a PP-based masterbatch containing 5% additive X (SMILES: CCC) and 0.2% additive Y (SMILES: C1=CC=C(C=C1)O) to improve UV resistance." Assume we have the SMILES representations of PP, additives X, and Y, along with their associated numerical descriptors. The text portion is first processed by the tokenizer like regular LLM input, yielding, for example: [CLS]Design a PP-based masterbatch containing 5% [ADD_X_NAME] and 0.2% [ADD_Y_NAME] to improve UV resistance. [SEP] Here, specific chemical names such as "[ADD_X_NAME]" serve as placeholders, whose corresponding structure and property information will be injected in subsequent steps. For each chemical component (such as additive X), its SMILES string (CCC) and numerical descriptor vector ([desc_x1,desc_x2,...]) need to be extracted. Similar processing is performed for additive Y and matrix PP.

[0083] The LLM's input embedding layer is then modified to incorporate multimodal information. The goal is for the first Transformer layer of the LLM to receive a rich embedding that incorporates all the information. This is accomplished by replacing the embeddings with special tokens. First, new special tokens are added to the LLM's vocabulary. For example, [CHEM_STRUCT_EMB] indicates where to insert chemical structure embeddings, and [CHEM_PROP_EMB] indicates where to insert chemical property embeddings. Then, when constructing the sequence input to the LLM, these special tokens are inserted after the textual mention of the corresponding chemical entity, resulting in a sequence such as ...5%[ADD_X_NAME][CHEM_STRUCT_EMB][CHEM_PROP_EMB]... When generating the final input embedding, the LLM's existing word embedding lookup mechanism is used for normal text tokens. When encountering the token [CHEM_STRUCT_EMB] for Additive X, the system feeds the SMILES string of Additive X into the previously prepared chemical structure encoder (GNN) to obtain its structural embedding, struct_emb_X. If the embedding dimension d_struct does not match the LLM's d_model, it is resized to the d_model dimension through a linear projection layer. This processed structural embedding vector replaces the default embedding of the [CHEM_STRUCT_EMB] token itself. Similarly, when the [CHEM_PROP_EMB] token is encountered, the numerical descriptor vector of the additive X is input into the numerical property encoder (MLP) to obtain the property embedding prop_emb_X, which is also optionally projected and replaces the default embedding of [CHEM_PROP_EMB]. This operation is performed for all chemical entities in the input sequence. Finally, standard positional encoding is applied to the entire sequence containing text, structure, and property embeddings, and this highly information-rich embedding sequence is then fed into the first Transformer layer of the LLM.

[0084] (4.4) Knowledge base retrieval strategy optimization

[0085] a) Setting the similarity threshold to 0.8 ensures that only search results with high semantic matching are included, improving the accuracy of the answers;

[0086] b) Adopt a hybrid search strategy, combining BM25 keyword search (accounting for 30%-40%) and vector semantic search (accounting for 60%-70%);

[0087] c) BM25 search matches high-frequency keywords, achieving initial recall and increasing the breadth of search;

[0088] d) Vector semantic retrieval calculates the semantic similarity between user questions and knowledge fragments to improve retrieval accuracy. A 20% validation set is reserved from the dataset to ensure that the formulations in the validation set have not appeared in the training set and include samples of different chemical structure types and performance ranges.

[0089] The model output is evaluated by the following indicators: For quantitative performance prediction (such as tensile strength, MI value):

[0090] Mean Absolute Error (MAE):

[0091] Root Mean Squared Error (RMSE):

[0092] Coefficient of determination (R-squared, R 2 ): Measures the ability of the model to explain variance.

[0093] For qualitative prediction or classification tasks (such as good / medium / poor compatibility, whether the color meets the standard):

[0094] Accuracy

[0095] Precision

[0096] Recall

[0097] F1-Score

[0098] Comparison of test results on the validation set before and after the improvement

[0099] The improved models (e.g., DeepSeek-R1:32B+GNN_Encoder) were compared with the baseline model (standard DeepSeek-R1:32B, which makes predictions based only on text names and RAGs) on the -Validation set enhanced with molecular information.

[0100]

[0101] By incorporating a learning enhancement mechanism that characterizes chemical structure and material properties, LLM is able to understand the properties of each fiber masterbatch component and their interactions within the formulation at a more fundamental level. This allows the model to not only "remember" existing formulation-performance relationships but also to perform a degree of "reasoning" and "generalization" based on chemical principles. This results in more accurate predictions of the properties of known systems and demonstrates stronger predictive power and innovative potential for new components or formulation combinations. This improvement significantly enhances the practical value of LLM in fiber masterbatch formulation design and optimization.

Claims

1. A fiber masterbatch formula and product development method based on intelligent technology, characterized in that: The following steps are involved: (a) Build a high-performance computing platform to provide computing power support for large-scale data processing and intelligent agent reasoning; (b) building an intelligent agent system based on modular services and a locally deployed large model, wherein the intelligent agent system uses a knowledge base search enhancement and a dynamic collaboration mechanism with the locally deployed large model; (c) Build and analyze a specialized knowledge base containing data related to fiber masterbatch formulations; (d) The intelligent agent system retrieves and enhances the user input to achieve intelligent recommendation, optimization and product development of fiber masterbatch formula.

2. The method according to claim 1, characterized in that The locally deployed large model includes a multimodal input embedding module, which is used to uniformly convert input information of different modalities, including text descriptions, chemical structures, and numerical attributes, into a dense vector representation that can be processed by the Transformer layer of the large model.

3. The method according to claim 2, characterized in that The steps of constructing the multimodal input embedding module and fusing information include: (a) Using a chemical structure encoder, the SMILES string of the chemical component is converted into a molecular graph representation, which is then input into a graph neural network to extract a graph-level structural embedding vector representing the entire molecule; (b) Using a numerical attribute encoder, a vector consisting of a series of numerical molecular descriptors or physical properties of the chemical components is input into a multi-layer perceptron network to obtain an attribute embedding vector; (c) When preparing an input sequence for the large model, the text information is first segmented and processed through a word embedding layer, and then a special tag is introduced for the chemical entity in the input sequence, and the structure embedding vector obtained in step (a) and the attribute embedding vector obtained in step (b) respectively replace the default embedding of the special tag of the corresponding chemical entity, thereby forming a sequence that integrates text embedding, structure embedding and attribute embedding, and finally position encoding is applied to the integrated sequence and sent to the first Transformer layer of the large model.

4. The method according to claim 3, characterized in that The graph neural network in the chemical structure encoder (a) is selected from at least one of a graph convolutional network (GCN), a graph attention network (GAT) or a message passing neural network (MPNN); and the step of extracting a graph-level structure embedding vector representing the entire molecule further includes applying a global pooling operation to integrate all atomic embeddings after all layers of the graph neural network have processed the molecular graph representation; the SMILES string of the chemical component in step (a) specifically includes a SMILES string of a base resin and at least one functional additive constituting the fiber masterbatch formula, and the functional additive is selected from at least one of a pigment, a dispersant or an antioxidant.

5. The method according to claim 3, characterized in that Before inputting the vector composed of a series of numerical molecular descriptors or physical properties of the chemical components into the multi-layer perceptron network, the numerical attribute encoder (b) also includes a step of preprocessing the composed vector, and the preprocessing step includes splicing the numerical features into a single vector and performing normalization or standardization on the single vector; the series of numerical molecular descriptors or physical properties of the chemical components in the step (b) specifically include at least one item selected from the molecular weight, topological polar surface area, octanol-water partition coefficient LogP, melting point, glass transition temperature, density or solubility parameter used to characterize the fiber masterbatch component.

6. The method according to claim 3, characterized in that When preparing the input sequence for the large model in step (c), the text information includes a specific description of the fiber masterbatch formula, which at least includes the type of base resin used in the fiber masterbatch formula, the type of one or more additives and their amount or percentage concentration in the formula, and the special mark corresponds to these chemical entities in the fiber masterbatch formula, so that the integrated sequence can characterize the composition and structural characteristics of the fiber masterbatch for subsequent performance prediction or formula optimization tasks.

7. The method according to claim 4, characterized in that The step of converting the SMILES string into a molecular graph representation includes: defining the nodes of the molecular graph as atoms and the edges as chemical bonds; and, the atomic features include at least one of atom type, charge and hybridization mode, and the chemical bond features include the type of chemical bond; when the graph neural network is a graph convolutional network (GCN), the GCN comprises four layers, each layer updates the vector representation of the current atom by aggregating information of neighboring atoms, and uses a ReLU activation function between layers; the dimension (d_struct) of the graph-level structure embedding vector is configured to be the same as the internal representation dimension (d_model) of the large model or can be projected to the d_model dimension through a linear layer; and, the chemical structure encoder is pre-trained on a large molecular dataset for molecular property prediction or self-supervised learning tasks; the multilayer perceptron network comprises 2-6 hidden layers and uses a ReLU activation function.

8. The method according to claim 3 or 6, characterized in that The dimension (d_numerical) of the attribute embedding vector is configured to be the same as the internal representation dimension (d_model) of the large model or to be projected to the d_model dimension via a linear layer; and the numerical attribute encoder is jointly trained end-to-end during the overall fine-tuning of the large model; The special tags introduced for the chemical entity include a first special tag for indicating the position where the chemical structure is embedded and a second special tag for indicating the position where the chemical attribute is embedded; Furthermore, the step of replacing the default embedding of the special tag corresponding to the chemical entity comprises: For the first special marker, replace its default embedding with the structure embedding vector obtained by the chemical structure encoder. If the dimension of the structure embedding vector is inconsistent with the internal representation dimension of the large model, it is first adjusted to the internal representation dimension through a linear projection layer; For the second special tag, its default embedding is replaced with the attribute embedding vector obtained by the numerical attribute encoder. If the dimension of the attribute embedding vector is inconsistent with the internal representation dimension of the large model, it is first adjusted to the internal representation dimension through a linear projection layer.

9. The method according to claim 1, characterized in that The steps of constructing and parsing a specialized knowledge base include: Collecting multimodal data, wherein the multimodal data includes at least chemical structure information, molecular descriptors, physical properties, formulation components and amounts, process parameters, and corresponding product performance data of chemical components; Connecting the multimodal data to the RAGFlow framework for data parsing and structuring, including converting tabular data into natural language descriptions, using visual models for document layout analysis, using OCR recognition and semantic segmentation to generate structured text blocks, and vectorizing the text blocks and storing them together with the original text to build a hybrid index system; The steps of building the intelligent agent system include: Adopting Docker container management to build an open source model management framework, supporting dynamic switching of large models and parameter fine-tuning; Locally deploy open source large language models and provide intelligent reasoning services based on the Ollama framework.

10. The method according to claim 9, characterized in that The steps of the retrieval enhancement process include: A hybrid search strategy combining BM25 keyword search and vector semantic search is used; Introducing a dynamic re-ranking model to optimize the ranking of search results by combining semantic relevance, time factors, and data weights; The steps of output enhancement processing include: Adopt a restricted output whitelist mechanism to ensure that the generated answers only reference verified material names, raw materials and process parameters; By generating a constraint mechanism, the answer must cite the source document and its specific location, and users can click to jump for verification; Professional terminology verification rules are set to detect terminology accuracy based on a regular rule library and automatically trigger regeneration when errors are detected; the high-performance computing platform adopts physical isolation technology and virtualization technology to ensure the integrity and privacy of data processing and storage.

Citation Information

Patent Citations

  • Semi-supervised learning fiber master batch electron microscope image agglomeration structure area identification method

    CN118135566A

  • Intelligent process generation method and system based on data driving

    CN118297275A

  • High polymer material modification optimization method and system based on RAG

    CN119230021A

  • Intelligent question answering system based on large language model

    CN119623646A

  • Power field knowledge question and answer optimization system based on large model retrieval enhancement generation and instruction supervision fine tuning

    CN119961388A