Process design standard intelligent question and answer system and method based on multi-modal large model and graph rag
By constructing a knowledge graph using a multimodal large model and a GraphRAG model, the problems of low efficiency and poor accuracy in traditional process design are solved, enabling efficient and accurate process standard query and design suggestions, and supporting dynamic updates.
Patent Information
- Application Number
- CN202511186903.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-08-25
AI Technical Summary
In traditional process design, designers need to manually consult a large number of scattered standard documents, which is inefficient and prone to errors. Traditional knowledge graphs cannot process unstructured image data, and RAG retrieval is not sensitive to the range of process parameters and lacks multimodal data fusion capabilities, resulting in low information utilization.
By employing a multimodal large model and a GraphRAG model, a knowledge graph is constructed to achieve intelligent recognition and information extraction from process standard images. An interval-sensitive retrieval strategy is designed, and through multimodal data collection, recognition, knowledge graph construction, user intent parsing, interval-sensitive retrieval, and intelligent suggestion generation, accurate intelligent question answering services are provided.
It improves the efficiency and accuracy of process design, increases retrieval speed by 80%, achieves a numerical range matching accuracy of 97.3%, supports the dynamic addition of new national standards without retraining the model, has a system response time of less than 2 seconds, and provides accurate design suggestions.
Smart Images

Figure CN120725156B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a process design standard intelligent question and answer system and method based on a multi-modal large model and GraphRAG, belonging to the field of artificial intelligence and process design models. BACKGROUND
[0002] Under the background of the digital transformation of today's manufacturing industry, process design, as a key link in product production, puts forward higher requirements for efficiency and accuracy. The traditional process design process relies on manual consultation of a large number of standard files, which not only is inefficient, but also is prone to human error.
[0003] Traditional process design pain points:
[0004] In the traditional process design process, designers face great challenges. Process design standards are scattered in numerous standard files, which are not only large in quantity, but also diverse in format, including paper documents, electronic documents, etc. Designers need to manually review these files to find relevant standard information for the current design task. For example, in the field of mechanical manufacturing, standards related to material selection, heat treatment process, machining precision, etc. Designers need to find relevant content in different standard files one by one. This manual retrieval method is extremely inefficient, and according to statistics, designers spend 30%-40% of the entire design cycle on finding standard information. Moreover, due to human negligence, it is easy to miss key information, resulting in defects in the design scheme, affecting product quality and production progress.
[0005] Limitations of existing models:
[0006] Traditional knowledge graph cannot handle unstructured picture data: Traditional knowledge graphs are mainly based on structured data construction, and can effectively organize and represent structured data such as text and tables. However, process design standards contain a large number of GB two-dimensional table pictures, and the information in these pictures is unstructured, which is difficult for traditional knowledge graphs to handle directly. For example, the information in the picture, such as the GB number, GB name, and process design standard two-dimensional table, cannot be directly recognized and utilized by traditional knowledge graphs, limiting the application range of knowledge graphs in the field of process design.
[0007] RAG retrieval is not sensitive to the interval range of process parameters: the traditional retrieval augmented generation (RAG) model is not sensitive to the interval range of process parameters when processing process standard retrieval. In process design, many parameters have specific interval requirements, such as quenching temperature, hardness range, etc. When querying the standard, the designer often pays attention to whether the specific value is within the reasonable interval. However, the traditional RAG retrieval cannot accurately judge the matching relationship between the user input value and the interval range in the process standard, resulting in the inability to normally recall the corresponding interval range of the process design standard. For example, when the designer queries "what is the cooling medium for 870℃ quenching", the traditional RAG retrieval may not accurately match the standard information with quenching temperature in the 860-880℃ interval, so as to provide accurate answers.
[0008] Lack of multi-modal data fusion capability:
[0009] The data in the process design field has the characteristics of multi-modal, including text, pictures, charts, etc. However, the existing model lacks effective multi-modal data fusion capability, and cannot integrate and utilize these different modal data. For example, when processing process standards, the information in the picture cannot be associated and analyzed with the text description, resulting in low utilization of information and inability to provide comprehensive and accurate design suggestions for designers.
[0010] Based on this background, the present application utilizes a multi-modal large model to process unstructured process standard pictures, combines a GraphRAG model to construct a knowledge graph and realize efficient retrieval, and provides accurate and fast intelligent question and answer services for process designers. SUMMARY
[0011] In order to solve the above problems, the present application discloses a process design standard intelligent question and answer system and method based on a multi-modal large model and GraphRAG, and the specific model scheme is as follows:
[0012] The process design standard intelligent question and answer system based on a multi-modal large model and GraphRAG comprises:
[0013] A multi-modal data acquisition module is used to collect GB two-dimensional table pictures of process standard design, wherein the GB two-dimensional table pictures include national standard number, national standard name, national standard introduction, and process design standard two-dimensional table corresponding to each national standard;
[0014] A multi-modal recognition module is integrated with a multi-modal large model to realize recognition of the GB two-dimensional table pictures, recognize the national standard number, national standard name, national standard introduction in the pictures, and convert the process design standard two-dimensional table corresponding to each national standard in the GB two-dimensional table pictures into a JSON string format array;
[0015] Knowledge graph construction module: Integrates graph database, saves the JSON string format array into graph database, where each national standard is an entity node, the name of the entity node is the national standard name, each row of data in the two-dimensional table forms a path, each attribute is a node of the path, the name of the node is the name + value of the attribute, and so on to construct knowledge graphs of multiple standards;
[0016] User intent parsing module: Based on a large language model, it performs named entity recognition on user questions and extracts the entities, attributes and numbers in the user questions;
[0017] The interval-sensitive retrieval module accesses the graph database for the extracted entities and attributes, matches similar paths by node name, and if a match is found, it further extracts the numerical range of the node path and determines whether the number in the user's question is within the numerical range of the node path. If yes, the path satisfies both the name and numerical range requirements and is used as input for the context generation module; otherwise, although the path name matches, the numerical range does not meet the requirements and the path is ignored.
[0018] Context generation module: After multiple rounds of recall, multiple paths are recalled, then deduplication is performed on the recalled paths, and finally they are converted into plain text format context.
[0019] Intelligent suggestion generation module: Provides context and prompt words to the large language model, which then generates final design suggestions for process standards.
[0020] Furthermore, the knowledge graph construction module stores entity nodes and paths using the Neo4j graph database. The Neo4j graph database stores data in the form of nodes, relationships, and attributes. Nodes store entity data, relationships store edge data between nodes and include directionality and attributes, and attributes are stored as key-value pairs on nodes or relationships, supporting graph traversal and attribute queries. Both nodes and relationships contain attribute information, and the attribute information is stored as attribute name + attribute value, i.e., a key-value pair pattern.
[0021] Furthermore, the interval-sensitive retrieval module employs a path matching algorithm based on cosine similarity to calculate the cosine similarity between entities and attributes in the user's question and the node paths in the graph database. Combining this with the numerical range of the node paths, it determines whether the numbers in the user's question fall within the numerical range of the node paths, thereby determining whether to recall the path.
[0022] Furthermore, the intelligent suggestion generation module guides the large language model to output structured suggestions through prompt word engineering. Based on the user's question, it recalls the set of paths with the highest relevance to the user's question through a knowledge graph and transforms the set of paths into context, thereby guiding the large language model to generate structured suggestions that meet the design requirements of the process standard.
[0023] The intelligent question-answering method based on the above-mentioned intelligent question-answering system for process design standards using multimodal large models and GraphRAG includes the following steps:
[0024] Step 1. Collect images of the GB two-dimensional table for process standard design;
[0025] Step 2. By integrating a multimodal large model, the GB two-dimensional table image is identified to recognize the national standard number, national standard name, national standard description in the image, and the corresponding process design standard two-dimensional table in the GB two-dimensional table image is converted into a JSON string format array;
[0026] Step 3. Store the JSON string array into a graph database, where each national standard is treated as an entity node, and the name of the node is the national standard name. Each row of data in the two-dimensional table forms a path, and each attribute is treated as a node in the path, with the name of the node being the name and value of the attribute. Construct a knowledge graph of multiple standards.
[0027] Step 4. Based on the large language model, perform named entity recognition on the user's question to extract the entity, attribute, and number of the user's question;
[0028] Step 5. For the extracted entities and attributes, access the graph database and match similar paths by node name. If a match is found, further extract the numerical range of the node path and determine whether the number in the user's question is within the numerical range of the node path.
[0029] Step 6. After multiple rounds of recall, multiple paths are recalled, and then deduplication is performed on the recalled paths, and finally they are converted into plain text format context.
[0030] Step 7. Provide the context and prompt words to the large language model, and the large language model will generate the final design suggestions for the process standard.
[0031] Furthermore, step 2, multimodal processing, specifically involves: using a multimodal large model combined with an OCR model to extract the table text from the image; the OCR model recognizing the text information in the image and converting it into an editable text format; and using the multimodal large model to convert the extracted two-dimensional table text information into JSON format data for subsequent storage and processing.
[0032] Furthermore, step 3 includes constructing a three-layer graph structure, specifically as follows:
[0033] National Standard Entity Layer: Includes metadata such as national standard number, name, and scope of application. Each national standard is an independent entity node, and its metadata is stored as the node's attributes.
[0034] Process parameter layer: Attributes are used as nodes, and attribute nodes are connected to nodes in the national standard entity layer through edges to indicate the relationship between them;
[0035] Numerical range layer: By using B+ tree indexes to define ranges, B+ tree indexes can be used for range queries, improving the retrieval efficiency of numerical parameters.
[0036] Furthermore, step 5 specifically involves designing a two-stage retrieval strategy. First, candidate paths are recalled through entity matching. Based on the entity information extracted from the user's question, matching entity nodes are searched in the knowledge graph, and paths related to the entity node are recalled. Then, numerical range filtering is used. For the recalled candidate paths, based on the numerical information in the user's question, it is determined whether they are within the range of numerical range nodes in the path, and non-matching paths are filtered out.
[0037] Furthermore, step 6 specifically includes:
[0038] 6.1 Range-sensitive search:
[0039] Path matching: The cosine similarity algorithm is used to calculate the matching degree between entity attributes and node paths. The core idea is to convert the source string and the target string into two vectors and evaluate the similarity by calculating the cosine value of the angle between the two vectors. The value range is [-1, 1]. The closer the value is to 1, the more similar the directions are. The closer the value is to -1, the more opposite the directions are. 0 means orthogonal and no correlation.
[0040] In a knowledge graph, the entity attribute strings in a user's question and the node path strings in the graph database are converted into two vectors. The cosine similarity is calculated to determine the degree of matching between the user's question and the node path. If the similarity is >= 0.9, the matching requirement is met.
[0041] Numerical comparison: Convert the user input value into a floating-point number and perform numerical operations with the interval endpoints in the knowledge graph to determine whether the user input value is within the interval, thereby determining whether to recall the corresponding path;
[0042] 6.2 Context Generation Strategy:
[0043] Deduplication algorithm: Path similarity judgment based on the longest common subsequence (LCS). The LCS algorithm measures the similarity between two sequences by finding the longest common subsequence. Among the recalled multiple paths, the LCS algorithm is used to judge the similarity between paths and remove paths with a similarity of ≥95% to reduce redundant information.
[0044] Text formatting: Converts JSON paths into natural language descriptions while preserving key parameters.
[0045] In this invention, GraphRAG refers to a knowledge graph-based retrieval enhancement generative model. The multimodal large model is the Qwen-VL-14B multimodal large model. GB is an abbreviation for 'national standard'.
[0046] The working principle of this invention is:
[0047] This invention is the first to propose a collaborative architecture between a multimodal large model and GraphRAG: This invention innovatively combines a multimodal large model with a GraphRAG model to construct a novel intelligent question-answering system architecture. The multimodal large model is responsible for processing unstructured image data, transforming it into structured information; the GraphRAG model is used to construct the knowledge graph and achieve efficient retrieval. The two work collaboratively, fully leveraging their respective advantages to improve the system's intelligence level and retrieval efficiency.
[0048] This invention enables the dynamic construction of a knowledge graph from process standard images. By introducing a multimodal large model, it achieves intelligent recognition and information extraction from process standard images, dynamically integrating this information into a knowledge graph. Unlike traditional static knowledge graphs, this invention's knowledge graph can be updated in real time, adapting to the continuous changes and updates in process standards, providing designers with the latest standard information.
[0049] To improve the matching accuracy of numerical parameters, this invention designs a range-sensitive retrieval strategy to address the problem that traditional RAG retrieval is insensitive to the range of process parameters. This strategy can accurately determine the matching relationship between the user-input numerical value and the range of process parameters in the knowledge graph, thereby improving the matching accuracy of numerical parameters and ensuring the retrieval of accurate process design standard information.
[0050] The beneficial effects of this invention are:
[0051] 1) Improved Efficiency: Retrieval speed is increased by 80% compared to traditional methods, with an average response time of less than 2 seconds. Through optimized multimodal processing, efficient knowledge graph construction, and fast retrieval algorithms, the system can quickly respond to user query requests, greatly improving the work efficiency of process designers. For example, in traditional process standard retrieval, it may take several minutes to find relevant information, while the system of this invention can return accurate results within 2 seconds.
[0052] 2) Accuracy Optimization: The numerical range matching accuracy reaches 97.3%, an improvement of 15 percentage points compared to traditional RAG. The range-sensitive retrieval strategy and precise numerical comparison algorithm enable the system to accurately match the numerical values in the user's problem with the ranges in the process standards, providing more accurate design suggestions. For example, when processing queries involving numerical ranges, traditional RAG may result in matching errors, while the system of this invention can accurately determine whether the value is within the range, improving the reliability of the retrieval results.
[0053] 3) High scalability: Supports dynamic addition of new national standards without retraining the model. When a new national standard for a process is released, the system can directly incorporate it into the knowledge graph. Support for the new standard can be achieved through a simple update operation, without retraining the entire model. For example, when a new national standard for heat treatment processes appears, simply process the relevant GB image and add it to the knowledge graph, and the system can provide users with intelligent question-and-answer services regarding the new standard. Attached Figure Description
[0054] Figure 1 This is a schematic diagram of the overall system module architecture in this invention.
[0055] Figure 2 This is a schematic diagram illustrating an example of how the present invention is implemented. Detailed Implementation
[0056] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.
[0057] The system framework of this invention is as follows:
[0058] 1) Multimodal Processing Layer: Utilizing a multimodal large model (Qwen-VL-14B) combined with an advanced OCR (Optical Character Recognition) model, the layer extracts text from GB 2D table images. The OCR model accurately identifies text information in the image and converts it into an editable text format. The multimodal large model then converts the extracted 2D table text information into JSON format data for subsequent storage and processing. For example, for a GB 2D table image containing heat treatment process standards, the multimodal processing layer can identify information such as the national standard number, national standard name, material grade, and quenching temperature, and organize it into a JSON format array, facilitating the construction of a knowledge graph.
[0059] 2) Knowledge Graph Layer: Construct a three-layer knowledge graph structure.
[0060] National Standard Entity Layer: Contains metadata such as national standard number, name, and scope of application. Each national standard is an independent entity node, and its metadata is stored as node attributes. For example, the national standard "GB / T 18211-2012 Terminology for Heat Treatment Process Materials" has "Terminology for Heat Treatment Process Materials" as the node name in the national standard entity layer, and the national standard number "GB / T 18211-2012", scope of application, etc., are used as node attributes.
[0061] Process parameter layer: This layer uses attributes such as material grade, quenching temperature, and cooling medium as nodes. These attribute nodes are connected to nodes in the national standard entity layer via edges, indicating their relationships. For example, in the national standard GB / T 18211-2012, "35 steel" is used as a material grade node, which is connected to the national standard node "Terminology for Heat Treatment Process Materials," indicating that "35 steel" is a material under this national standard.
[0062] Numerical Range Layer: This layer stores the ranges of parameters such as temperature and hardness using a B+ tree index. The B+ tree index enables efficient range queries, improving the retrieval efficiency of numerical parameters. For example, for the quenching temperature parameter, its range [860, 880] is stored in a B+ tree index, allowing users to quickly locate the corresponding range when querying relevant temperature information.
[0063] 3) Retrieval Enhancement Layer: A two-stage retrieval strategy is designed. First, candidate paths are recalled through entity matching. Based on the entity information extracted from the user's question, matching entity nodes are searched in the knowledge graph, and paths related to those entity nodes are recalled. Second, numerical range filtering is used. For the recalled candidate paths, based on the numerical information in the user's question, it is determined whether they fall within the range of numerical range nodes in the path, filtering out mismatched paths to improve the accuracy of the retrieval results.
[0064] The user question retrieval method of this invention is specifically as follows:
[0065] 1) Interval-sensitive search algorithm:
[0066] Path matching: The cosine similarity algorithm is used to calculate the matching degree between entity attributes and node paths. The core idea is to convert the source string and the target string into two vectors and evaluate the similarity by calculating the cosine value of the angle between the two vectors. The value range is [-1, 1]. The closer the value is to 1, the more similar the directions are. The closer the value is to -1, the more opposite the directions are. 0 means orthogonal and no correlation.
[0067] In the knowledge graph, the entity attribute strings in the user's question and the node path strings in the graph database are converted into two vectors. Cosine similarity is calculated to determine the degree of matching between the user's question and the node path; a similarity >= 0.9 indicates a successful match. Numerical comparison: The user's input value is converted to floating-point number and compared with the endpoints of the interval in the knowledge graph. For example, when a user queries "What is the cooling medium for quenching at 870℃?", 870 is converted to floating-point number and compared with the endpoints of the quenching temperature interval [860, 880] in the knowledge graph to determine if 870 is within that interval, thus determining whether to recall the corresponding path.
[0068] 2) Context generation strategy:
[0069] Deduplication Algorithm: Path Similarity Assessment Based on Longest Common Subsequence (LCS). The LCS algorithm measures the similarity between two sequences by finding the longest common subsequence. Among the recalled multiple paths, the LCS algorithm is used to determine the similarity between paths, removing paths with excessively high similarity (>=95%) to reduce redundant information.
[0070] Text formatting: Convert JSON paths into natural language descriptions while preserving key parameters. For example, convert JSON path data retrieved from a knowledge graph into a natural language description such as "According to GB / T 18211-2012, the quenching temperature of No. 35 steel is 860-880℃, and the cooling medium is water," making it easier for users to understand and use.
[0071] The specific execution process of this invention is described below; see [link / reference]. Figure 1 :
[0072] Step 1: Multimodal data acquisition, collecting GB two-dimensional table images of process standard design. The GB two-dimensional table images contain the national standard number, national standard name, national standard introduction, and the corresponding process design standard two-dimensional table. Refer to Table 1. Table 1 is GB / T 18211-2012-heat treatment process material terminology in the embodiment.
[0073] Table 1. GB / T 18211-2012 - Terminology for Heat Treatment Process Materials
[0074]
[0075] Step 2: By integrating a multimodal large model, all GB two-dimensional table images are identified, including the national standard number, national standard name, national standard description, and the two-dimensional tables are converted into JSON string arrays. Each element of the JSON string array contains attributes such as material grade, required hardness range, quenching temperature, cooling medium, hardness after quenching, tempering temperature range, national standard number, national standard name, and scope of application of the national standard. The specific JSON format is as follows: [
[0077] {
[0078] Material grade: 35# steel
[0079] Required hardness range: Greater than 50
[0080] Quenching temperature: 860~880
[0081] Cooling medium: Water
[0082] Hardness after quenching: ≥50
[0083] Tempering temperature range: Below 200℃
[0084] National Standard Number: GB / T 18211-2012
[0085] National Standard Name: "Terminology for Heat Treatment Process Materials"
[0086] "Scope of Application": "This standard specifies the main terms, definitions, and English translations for heat treatment process materials, including heating media, quenching cooling media, chemical heat treatment infiltration agents, and heat treatment protective coatings. This standard applies to relevant standards and model files for heat treatment professional models."
[0087] },
[0088] {
[0089] Material grade: 35# steel
[0090] Required hardness range: 46~50
[0091] Quenching temperature: 860~880
[0092] Cooling medium: Water
[0093] Hardness after quenching: ≥50
[0094] Tempering temperature range: 240~270
[0095] National Standard Number: GB / T 18211-2012
[0096] National Standard Name: "Terminology for Heat Treatment Process Materials"
[0097] "Scope of Application": "This standard specifies the main terms, definitions, and English translations for heat treatment process materials, including heating media, quenching cooling media, chemical heat treatment infiltration agents, and heat treatment protective coatings. This standard applies to relevant standards and model files for heat treatment professional models."
[0098] },...]
[0099] Step 3: Save the JSON string array to the graph database, where each national standard is treated as an entity node, and the name of the node is the national standard name. Each row of data in the two-dimensional table forms a path, and each attribute is treated as a node of the path, with the name of the node being the name and value of the attribute. Construct a knowledge graph of multiple standards.
[0100] Step 4: Based on the large language model, perform named entity recognition on the user's question to extract the entities, attributes, and numbers. See the example code below:
[0101] def extract_entities(self, text):
[0102] # Step 1: Basic Entity Recognition
[0103] ner_results = self.ner_pipe(text)
[0104] # Step 2: Merge stop word entities
[0105] merged_entities = self._merge_stop_entities(text, ner_results)
[0106] # Step 3: Extracting Numbers and Attributes
[0107] results = self._extract_properties(merged_entities, text)
[0108] return results
[0109] def _merge_stop_entities(self, text, entities):
[0110] # Create a text tag mapping table
[0111] text_length = len(text)
[0112] mask = [0] * text_length
[0113] merged = []
[0114] # Prioritize handling custom stop words
[0115] for stop_word in self.stop_entities:
[0116] start = 0
[0117] while True:
[0118] idx = text.find(stop_word, start)
[0119] if idx == -1:
[0120] break
[0121] # Mark matched regions
[0122] for i in range(idx, idx+len(stop_word)):
[0123] if i <text_length:
[0124] mask[i] = 1
[0125] merged.append({
[0126] "word": stop_word,
[0127] "start": idx,
[0128] "end": idx + len(stop_word),
[0129] "entity_group": "MATERIAL"
[0130] })
[0131] start = idx + len(stop_word)
[0132] # Merge model identification results (excluding already labeled regions)
[0133] for ent in entities:
[0134] overlap = sum(mask[ent['start']:ent['end']])
[0135] if overlap == 0:
[0136] merged.append(ent)
[0137] # Update marker
[0138] for i in range(ent['start'], ent['end']):
[0139] if i <text_length:
[0140] mask[i] = 1
[0141] return sorted(merged, key=lambda x: x['start'])
[0142] def _extract_properties(self, entities, text):
[0143] # Extract Numbers
[0144] numbers = [m.group() for m in re.finditer(r'\d+\.?\d*',text)]
[0145] # Extract attributes (including composite attributes)
[0146] properties = []
[0147] current_prop = []
[0148] for ent in entities:
[0149] word = ent["word"]
[0150] # Check property keywords
[0151] if word in self.property_keywords:
[0152] current_prop.append(word)
[0153] elif current_prop:
[0154] # Process compound properties (such as "hardness after quenching")
[0155] compound_prop = "".join(current_prop + [word])
[0156] if any(kw in compound_prop for kw in self.property_keywords):
[0157] properties.append(compound_prop)
[0158] current_prop = []
[0159] else:[[ID=-28]]
[0160] current_prop.append(word)
[0161] else:
[0162] current_prop = []
[0163] # Combine the final results
[0164] result = []
[0165] for ent in entities:
[0166] if ent["entity_group"] in ["MATERIAL", "ORG", "PRODUCT"]:
[0167] result.append(ent["word"])
[0168] result += numbers
[0169] result += [prop for prop in properties if prop]
[0170] return list(set(result)) # Remove duplicates
[0171] Step 5: For the extracted entities, attributes, and numbers, access the graph database and match similar paths by node name. If a match is found, further extract the numerical range of the node path and determine if the number in the user's question falls within this range. If it does, use it as the target path and as input for Step 6. If it doesn't, the path is ignored because it doesn't meet the range requirement. See the example code below:
[0172] class BPlusTreeNode:
[0173] def __init__(self, is_leaf=False):
[0174] self.is_leaf = is_leaf
[0175] self.keys = [] # Starting value of the storage range (sorted)
[0176] self.ranges = [] # Stores complete range tuples (new feature)
[0177] self.children = [] # Pointers to the child nodes of non-leaf nodes
[0178] self.next = None # Linked list pointer to the leaf node
[0179] class BPlusTree:
[0180] def __init__(self, order=3):
[0181] self.root = BPlusTreeNode(is_leaf=True)
[0182] self.order = order # Maximum node capacity
[0183] def _find_leaf(self, value):
[0184] Locate the leaf node containing the target value.
[0185] node = self.root
[0186] while not node.is_leaf:
[0187] # Binary search to locate child nodes
[0188] idx = self._bisect_right(node.keys, value)
[0189] node = node.children[idx]
[0190] return node
[0191] def _split_leaf(self, node):
[0192] """Leaf node splitting"""
[0193] mid = self.order / / 2
[0194] new_node = BPlusTreeNode(is_leaf=True)
[0195] new_node.keys = node.keys[mid:]
[0196] new_node.ranges = node.ranges[mid:]
[0197] node.keys = node.keys[:mid]
[0198] node.ranges = node.ranges[:mid]
[0199] new_node.next = node.next
[0200] node.next = new_node
[0201] self._insert_parent(node, new_node.keys[0], new_node)
[0202] def insert_range(self, start, end):
[0203] """Insert range (auto-sort)""
[0204] leaf = self._find_leaf(start)
[0205] idx = self._bisect_left(leaf.keys, start)
[0206] leaf.keys.insert(idx, start)
[0207] leaf.ranges.insert(idx, (start, end)) # Store the complete range
[0208] if len(leaf.keys) > self.order: # Node splitting
[0209] self._split_leaf(leaf)
[0210] def query_value(self, value):
[0211] """Determines if the value is within the storage range"""
[0212] leaf = self._find_leaf(value)
[0213] for i, key in enumerate(leaf.keys):
[0214] start, end = leaf.ranges[i]
[0215] if value>= start and value<= end:
[0216] return True
[0217] return False
[0218] Step 6: After multiple rounds of recall, multiple paths may be recalled. Then, deduplication is performed on the recalled paths, and finally, they are converted into plain text format context.
[0219] Step 7: Provide the context and prompt words to the large language model, and the large language model will generate the final design suggestions for the process standard.
[0220] Specific implementation examples are as follows:
[0221] Example: Quenching process parameter lookup
[0222] The user entered: "What is the required hardness of 35 steel after quenching?"
[0223] Processing flow:
[0224] Named Entity Recognition: A large language model is used to perform named entity recognition on the user's question, extracting "35 steel" (material grade) and "hardness after quenching" (attribute). Through semantic analysis of the question text, the large language model can accurately identify key entities and attribute information.
[0225] Knowledge graph retrieval: Based on the extracted entity information of "No. 35 steel", the "No. 35 steel" node is matched in the knowledge graph, and paths related to hardness are traversed. The knowledge graph, with its structured storage method, can quickly locate process parameter paths related to "No. 35 steel".
[0226] Numerical judgment: The system queries the hardness range [50, ∞). Since the user's question does not contain a specific numerical value, the system directly retrieves the path. The knowledge graph already stores the range information for "hardness of 35 steel after quenching," allowing the system to quickly query and determine whether to retrieve the path.
[0227] Result Generation: The recalled paths are converted into context and combined with prompts, which are then provided to the large model. The large model outputs, "According to GB / T 18211-2012, the hardness of No. 35 steel after quenching should be ≥50HRC." Based on the context and prompts, the large model generates accurate and clear process standard design suggestions to meet the user's query needs. (See effect for reference.) Figure 2 .
[0228] Example: Temperature Range Matching Verification
[0229] The user entered: "What is the cooling medium for quenching at 870℃?"
[0230] Processing flow:
[0231] Entity extraction: The user's question is analyzed using a large language model to extract "quenching temperature" (attribute) and "870℃" (numerical value). The large language model can accurately understand the semantics of the question and extract key attribute and numerical information.
[0232] Knowledge graph retrieval: Matching paths within the quenching temperature range [860, 880]℃ in the knowledge graph. The index structure and query algorithm of the knowledge graph can quickly locate paths related to this temperature range.
[0233] Numerical comparison: The system compares the user-inputted 870℃ with the quenching temperature range [860, 880]℃ in the graph, determines that 870℃ falls within the range, and recalls the corresponding path. The system uses a precise numerical comparison algorithm to determine the matching relationship between the user-inputted value and the range in the graph.
[0234] Result Generation: The recalled path is converted into context and combined with prompts, provided to the large model. The large model outputs, "When quenching at 870℃, water is recommended as the cooling medium (GB / T 18211-2012)." Based on the context information, the large model generates design suggestions that conform to process standards and explicitly cites national standards, providing users with reliable references. (Effectiveness reference follows.) Figure 2 .
[0235] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including modeling and scientific terms) have the same meaning as commonly understood by those skilled in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have a meaning consistent with their meaning in the context of existing models, and should not be interpreted in an idealized or overly formal sense unless defined as herein.
[0236] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. An intelligent question-answering method for a process design standard intelligent question-answering system based on multimodal large models and GraphRAG, characterized in that, Includes the following steps: Step 1. Collect images of the GB two-dimensional table for process standard design; Step 2. By integrating a multimodal large model, the GB two-dimensional table image is identified, including the national standard number, national standard name, national standard description, and the corresponding process design standard two-dimensional table in the GB two-dimensional table image is converted into a JSON string format array; Step 3. Store the JSON string array into a graph database, where each national standard is treated as an entity node, and the name of the node is the national standard name. Each row of data in the two-dimensional table forms a path, and each attribute is treated as a node of the path, with the name of the node being the name + value of the attribute. Construct a knowledge graph of multiple standards. The knowledge graph has a three-layer graph structure, specifically: National Standard Entity Layer: Includes metadata such as national standard number, name, and scope of application. Each national standard is an independent entity node, and its metadata is stored as the node's attributes. Process parameter layer: Attributes are used as nodes, and attribute nodes are connected to nodes in the national standard entity layer through edges to indicate the relationship between them; Numerical range layer: By using B+ tree indexes to define ranges, B+ tree indexes can be used for range queries, improving the retrieval efficiency of numerical parameters; Step 4. Based on the large language model, perform named entity recognition on the user's question to extract the entity, attribute, and number of the user's question; Step 5. For the extracted entities and attributes, access the graph database and match similar paths by node name. If a match is found, further extract the numerical range of the node path and determine whether the number in the user's question is within the numerical range of the node path. Specifically, design a two-stage retrieval strategy: first, recall candidate paths through entity matching. Based on the entity information extracted from the user's question, search for matching entity nodes in the knowledge graph and recall paths related to that entity node; then, filter by numerical range. For the recalled candidate paths, determine whether they are within the range of the numerical range nodes in the path based on the numerical information in the user's question, and filter out unmatched paths. Step 6. After multiple rounds of recall, multiple paths are retrieved. Then, deduplication is performed on the retrieved paths, and finally, they are converted into plain text context. Specifically: 6.1 Range-sensitive search: Path matching: The cosine similarity algorithm is used to calculate the matching degree between entity attributes and node paths. The core idea is to convert the source string and the target string into two vectors and evaluate the similarity by calculating the cosine value of the angle between the two vectors. The value range is [-1, 1]. The closer the value is to 1, the more similar the directions are. The closer the value is to -1, the more opposite the directions are. 0 means orthogonal and no correlation. In a knowledge graph, the entity attribute strings in a user's question and the node path strings in the graph database are converted into two vectors. The cosine similarity is calculated to determine the degree of matching between the user's question and the node path. If the similarity is >= 0.9, the matching requirement is met. Numerical comparison: Convert the user input value into a floating-point number and perform numerical operations with the interval endpoints in the knowledge graph to determine whether the user input value is within the interval, thereby determining whether to recall the corresponding path; 6.2 Context Generation Strategy: Deduplication algorithm: Path similarity judgment based on the longest common subsequence (LCS). The LCS algorithm measures the degree of similarity by finding the longest common subsequence between two sequences. Among the recalled multiple paths, the LCS algorithm is used to judge the similarity between paths and remove paths with a similarity of ≥95% to reduce redundant information. Text formatting: Converts JSON paths into natural language descriptions; Step 7. Provide the context and prompt words to the large language model, and the large language model will generate the final design suggestions for the process standard.
2. The intelligent question-answering method of the process design standard intelligent question-answering system based on multimodal large model and GraphRAG as described in claim 1, characterized in that, Step 2, multimodal processing, specifically involves: using a multimodal large model combined with an OCR model to extract the table text from the image; the OCR model recognizing the text information in the image and converting it into an editable text format; and using the multimodal large model to convert the extracted two-dimensional table text information into JSON format data for subsequent storage and processing.
3. A standard intelligent question-answering system for process design based on multimodal large model and GraphRAG, implementing the intelligent question-answering method of claim 1 or 2, characterized in that, include: Multimodal data acquisition module: used to collect GB two-dimensional table images of process standard design. The GB two-dimensional table images include the national standard number, national standard name, national standard introduction, and the corresponding two-dimensional table of process design standards for each national standard. Multimodal recognition module: Integrates a large multimodal model to recognize GB two-dimensional table images, identify the national standard number, national standard name, national standard description in the image, and convert the two-dimensional tables of process design standards corresponding to each national standard in the GB two-dimensional table image into JSON string format arrays; Knowledge graph construction module: Integrates graph database, saves the JSON string format array to graph database, where each national standard is an entity node, the name of the entity node is the national standard name, each row of data in the two-dimensional table forms a path, each attribute is a node of the path, the name of the node is the name + value of the attribute, and so on to construct knowledge graphs of multiple standards; User intent parsing module: Based on a large language model, it performs named entity recognition on user questions and extracts the entities, attributes and numbers in the user questions; The interval-sensitive retrieval module accesses the graph database for the extracted entities and attributes, matches similar paths by node name, and if a match is found, it further extracts the numerical range of the node path and determines whether the number in the user's question is within the numerical range of the node path. If yes, the path satisfies both the name and numerical range requirements and is used as input for the context generation module; otherwise, although the path name matches, the numerical range does not meet the requirements and the path is ignored. Context generation module: After multiple rounds of recall, multiple paths are recalled, then deduplication is performed on the recalled paths, and finally they are converted into plain text format context. Intelligent suggestion generation module: Provides context and prompt words to the large language model, which then generates the final design suggestions for process standards.
4. The intelligent question-answering system for process design standards based on multimodal large models and GraphRAG as described in claim 3, characterized in that, The knowledge graph construction module stores entity nodes and paths through the Neo4j graph database. The Neo4j graph database stores data in the form of nodes, relations, and attributes. Nodes store entity data, relations store edge data between nodes and include directionality and attributes, and attributes are stored as key-value pairs on nodes or relations. It supports graph traversal and attribute query.
5. The intelligent question-answering system for process design standards based on multimodal large models and GraphRAG as described in claim 3, characterized in that, The interval-sensitive retrieval module employs a path matching algorithm based on cosine similarity to calculate the cosine similarity between entities and attributes in the user's question and the node paths in the graph database. Combining this with the numerical range of the node paths, it determines whether the numbers in the user's question fall within the numerical range of the node paths, thereby determining whether to recall the path.
6. The intelligent question-answering system for process design standards based on multimodal large models and GraphRAG as described in claim 3, characterized in that, The intelligent suggestion generation module guides the large language model to output structured suggestions through prompt word engineering. Based on the user's question, it recalls the set of paths with the highest relevance to the user's question through a knowledge graph and transforms the set of paths into context, thereby guiding the large language model to generate structured suggestions that meet the process standard design requirements.
Citation Information
Patent Citations
Aviation standard question and answer optimization method and system based on atlas and document data
CN119621894A
Cited By
Software system development method based on multi-modal AI large model
CN121657970A