Mixed knowledge retrieval enhancement generation method, device and equipment

By using a hybrid knowledge retrieval enhancement generation method, the similarity score of the knowledge graph for prefabricated building quality management is calculated using dense and sparse retrieval algorithms, and a structurally complete retrieval answer is generated. This solves the problem of relying on personal experience and isolated information in the quality inspection of prefabricated buildings, and improves the efficiency and accuracy of quality control.

CN121808032APending Publication Date: 2026-04-07UNIV OF MACAU
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Quality inspection of prefabricated buildings relies on personal experience and lacks intelligent reasoning and interactive decision support, resulting in inconsistent quality levels and isolated information, which cannot meet the efficient query needs of modern project management.

Method used

A hybrid knowledge retrieval enhancement generation method is adopted. By using dense and sparse retrieval algorithms to calculate the similarity scores of each node block in the quality management knowledge graph, a complete retrieval answer is generated, covering key relationships and forming a logical chain of evidence.

Benefits of technology

Significantly shortens problem diagnosis time, improves the efficiency and accuracy of quality control, reduces rework and cost overruns, and effectively enhances the quality management level of prefabricated building projects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808032A_ABST
    Figure CN121808032A_ABST
Patent Text Reader

Abstract

The invention provides a mixed knowledge retrieval enhancement generation method, device and equipment, and relates to the technical field of knowledge retrieval, and the method comprises the steps: obtaining a quality management knowledge query request for an assembly type prefabricated part; according to a semantic entity in the query request, adopting a preset dense retrieval algorithm to calculate a dense similarity score of each node block in a quality management knowledge graph of the fabricated prefabricated component; according to terms in the query request, calculating a sparse similarity score of each node block by adopting a preset sparse retrieval algorithm; calculating a fusion similarity score of each node block according to the dense similarity score and the sparse similarity score, and traversing the quality management knowledge graph according to the fusion similarity score to obtain a similarity score of each edge block in the quality management knowledge graph; and generating a retrieval answer corresponding to the query request according to the fusion similarity score of each node block and the similarity score of each edge block. The control level of quality management of the fabricated prefabricated parts can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of knowledge retrieval technology, and more specifically, to a hybrid knowledge retrieval enhancement generation method, apparatus, and device. Background Technology

[0002] The core of prefabricated construction lies in transferring a large amount of on-site work to factories, producing prefabricated components and assembling them on-site to complete the building. However, based on this off-site production characteristic of "production first, assembly later," if quality problems cannot be detected in time or even transported to the construction site, they will face repairs, rework, or even scrapping, resulting in serious delays and cost overruns.

[0003] Currently, the accuracy of quality inspections in prefabricated building projects relies excessively on the personal experience and professional knowledge of inspectors. The shortage and high turnover of experienced quality inspectors lead to unsatisfactory inspection results and inconsistent quality levels. Secondly, the large number, complexity, and fragmented regulations and standards involved pose significant challenges to the learning, tracking, and implementation by quality inspectors. Although some quality information query systems exist, they are mostly limited to static data records, rely on rigid predefined rules, and lack intelligent reasoning and interactive decision support capabilities for complex quality issues, failing to meet the dynamic and efficient query needs of modern project management. Summary of the Invention

[0004] This application provides a hybrid knowledge retrieval enhancement generation method, apparatus, and equipment, which can improve the quality management and control level of prefabricated components.

[0005] In a first aspect, embodiments of the present invention provide a hybrid knowledge retrieval enhancement generation method, the method comprising: Request a query for quality management knowledge related to prefabricated components; Based on the semantic entities in the quality management knowledge query request, a preset dense retrieval algorithm is used to calculate the dense similarity score of each node block in the quality management knowledge graph of the prefabricated components. Based on the terms in the quality management knowledge query request, a preset sparse retrieval algorithm is used to calculate the sparse similarity score of each node block in the quality management knowledge graph. Based on the dense similarity scores of each node and the sparse similarity scores of each node block, the fusion similarity score of each node block is calculated. Based on the fusion similarity scores of each node block, the quality management knowledge graph is traversed to obtain the similarity scores of each edge block in the quality management knowledge graph; Based on the fusion similarity score of each node block and the similarity score of each edge block, the retrieval answer corresponding to the quality management knowledge query request is generated.

[0006] Optionally, the step of calculating the dense similarity score of each node block in the quality management knowledge graph of the prefabricated component based on the semantic entities in the quality management knowledge query request and using a preset dense retrieval algorithm includes: The semantic entities are encoded into dense vectors; Calculate the cosine similarity between the dense vector of the semantic entity and the dense vector of each node block to obtain the dense similarity score of each node block.

[0007] Optionally, the step of calculating the sparse similarity score of each node block in the quality management knowledge graph based on the terms in the quality management knowledge query request using a preset sparse retrieval algorithm includes: Based on the total number of documents corresponding to the quality management knowledge graph and the number of documents in which the term appears, calculate the inverse document frequency information of the term; The weight of the term is calculated based on the inverse document frequency information of the term, the named entity recognition information of the term, and the part-of-speech tagging information; Based on the weights of the terms, the sparse similarity score of each node block is calculated.

[0008] Optionally, the step of traversing the quality management knowledge graph based on the fusion similarity scores of each node block to obtain the similarity scores of each edge block in the quality management knowledge graph includes: Based on the fusion similarity score of each node block and the preset jump distance, a preset distance decay function is used to traverse the quality management knowledge graph to obtain the similarity score of each edge block.

[0009] Optionally, generating the retrieval answer corresponding to the quality management knowledge query request based on the fusion similarity score of each node block and the similarity score of each edge block includes: Based on the fusion similarity score and importance evaluation value of each node block, semantic alignment is performed on each node block to obtain the target score of each node block; Based on the similarity score and edge weight of each edge block, semantic alignment is performed on each edge block to generate a target score for each edge block; Based on the target scores of each node block and each edge block, the target node block and target edge block with the highest scores are determined from the quality management knowledge graph. The search answer is generated based on the target node block and the target edge block.

[0010] Optionally, generating the search answer based on the target node block and the target edge block includes: Based on the target node block and the target edge block, the target community report is determined from the quality management knowledge graph; The search answer is generated based on the target community report.

[0011] Optionally, the method further includes: The quality management knowledge document for the prefabricated components is converted into text knowledge blocks. Based on the text knowledge block, a first prompt word is generated using a first preset structured indicator word template; Based on the first prompt word, a preset large model is used to generate semantic entities in the text knowledge block and the relationships between semantic entities; Based on the semantic entities in the text knowledge block and the relationships between them, construct the subgraph corresponding to the text knowledge block; The subgraphs corresponding to each of the text knowledge blocks are merged to generate the quality management knowledge graph.

[0012] Optionally, the method further includes: Entity clustering is performed on the quality management knowledge graph to obtain the knowledge communities corresponding to the quality management knowledge graph; Based on the knowledge community, a second prompt word is generated using a second preset structured indicator word template; Based on the second prompt word, a community report of the knowledge community is generated using a preset large model.

[0013] Secondly, embodiments of the present invention also provide a hybrid knowledge retrieval enhancement generation device, the device comprising: The acquisition module is used to retrieve query requests for quality management knowledge related to prefabricated components. The calculation module is used to calculate the dense similarity score of each node block in the quality management knowledge graph of the prefabricated component based on the semantic entities in the quality management knowledge query request and using a preset dense retrieval algorithm. The calculation module is also used to calculate the sparse similarity score of each node block in the quality management knowledge graph based on the terms in the quality management knowledge query request and using a preset sparse retrieval algorithm. The calculation module is also used to calculate the fusion similarity score of each node block based on the dense similarity score of each node and the sparse similarity score of each node block; The calculation module is also used to traverse the quality management knowledge graph based on the fusion similarity score of each node block to obtain the similarity score of each edge block in the quality management knowledge graph. The generation module is used to generate the retrieval answer corresponding to the quality management knowledge query request based on the fusion similarity score of each node block and the similarity score of each edge block.

[0014] Thirdly, embodiments of the present invention also provide an electronic device, including: a processor, a memory, and a bus, wherein the memory stores program instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the program instructions to perform the steps of the hybrid knowledge retrieval enhancement generation method as described in any of the first aspects. This application provides a hybrid knowledge retrieval enhancement generation method, apparatus, and equipment. First, it obtains a user's query request for quality management knowledge related to prefabricated components. Based on the semantic entities in the query request, a dense retrieval algorithm is used to calculate the dense similarity score of each node block in the quality management knowledge graph. Simultaneously, based on the terms in the query request, a sparse retrieval algorithm is used to calculate the sparse similarity score of each node block. By fusing the two types of scores, a fused similarity score for each node block is obtained. Based on this, the knowledge graph is traversed to obtain the similarity score of each edge block. Finally, based on the fused similarity score of each node block and the similarity score of each edge block, the retrieval answer corresponding to the quality management knowledge query request is generated. This hybrid knowledge retrieval enhancement generation method, by fusing dense and sparse retrieval, calculates and merges the similarity of each node block in the quality management knowledge graph at the semantic and terminological levels, obtaining the fused similarity score of the node block. This score guides the traversal of the graph, further deriving the similarity scores of relational edge blocks. The final answer not only includes the core entities in the query but also covers key relationships, forming a logical chain of evidence. This directly solves the problems of fragmented knowledge and isolated retrieval information in prefabricated building quality management. It provides structurally complete and traceable reference information for specific tasks such as quality defect diagnosis and process standard query, effectively supporting on-site decision-making. This significantly shortens problem diagnosis time, improves the efficiency and accuracy of quality control, reduces rework and cost overruns caused by information omissions or misjudgments, and effectively enhances the quality management control level of prefabricated building projects. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A flowchart illustrating a hybrid knowledge retrieval enhancement generation method provided in this application; Figure 2 A flowchart illustrating a dense retrieval algorithm for hybrid knowledge retrieval enhancement generation provided in this application; Figure 3 A flowchart illustrating a sparse retrieval algorithm in hybrid knowledge retrieval enhancement generation provided in this application; Figure 4 A schematic diagram illustrating the process of generating search answers in a hybrid knowledge retrieval enhancement generation method provided in this application; Figure 5 A schematic diagram illustrating the process of generating retrieval answers in another hybrid knowledge retrieval enhancement generation method provided in this application; Figure 6 A schematic diagram illustrating the process of constructing a quality management knowledge graph in a hybrid knowledge retrieval enhancement generation method provided in this application; Figure 7 A schematic diagram illustrating the process of generating a community report in a hybrid knowledge retrieval enhancement generation method provided in this application; Figure 8 A schematic diagram of a hybrid knowledge retrieval enhancement generation device provided in this application; Figure 9 A schematic diagram of an electronic device provided in this application.

[0017] The reference numerals are as follows: 1000, acquisition module; 2000, calculation module; 3000, generation module; 10, electronic device; 11, processor; 12, memory; 13, bus. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0019] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0020] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0021] Before providing a detailed explanation of this application, let's first introduce its application scenarios.

[0022] In the quality management phase of prefabricated building projects, quality managers frequently face the challenge of quickly diagnosing quality defects or process compliance issues in prefabricated components. For example, when a prefabricated component is found to have appearance or performance defects, the quality manager needs to trace possible causes, which may involve multiple isolated aspects such as design specifications, material supply, production line processes, transportation and storage, or on-site installation. Currently, quality managers can only manually consult scattered design drawings, material reports, process and method libraries, and acceptance standards, relying on personal experience for correlation and analysis. This process is not only inefficient but also heavily reliant on individual experience, making it difficult to guarantee the completeness and accuracy of the analysis results, and easily leading to delays in problem handling or biased decision-making. Existing quality information systems cannot automatically construct and infer complete quality retrieval answers across stages and disciplines from massive amounts of data.

[0023] Based on this, embodiments of this application provide a hybrid knowledge retrieval enhancement generation method, apparatus, and device. By fusing dense and sparse retrieval, the similarity of each node block in the quality management knowledge graph at the query semantic and terminology levels is calculated and fused to obtain the fused similarity score of the node block. This score guides the traversal of the graph to further deduce the similarity score of the relational edge blocks. This ensures that the final answer not only includes the core entity in the query but also covers the key relationships, forming a logical chain of evidence. This directly solves the problems of fragmented knowledge and isolated retrieval information in the quality management of prefabricated buildings. It can provide structurally complete and traceable reference information for specific tasks such as quality defect diagnosis and process standard query, effectively supporting on-site decision-making. This significantly shortens the problem diagnosis time, improves the efficiency and accuracy of quality control, reduces rework and cost overruns caused by information omissions or misjudgments, and effectively improves the control level of quality management in prefabricated building projects.

[0024] It should be noted that the execution subject of the hybrid knowledge retrieval enhancement generation method provided in this application is an electronic device with processing capabilities, such as a computer device, which may be a laptop, desktop computer, server, programmable logic controller (PLC) circuit, etc. This application does not impose any restrictions on this.

[0025] The following explanation, in conjunction with the accompanying drawings, uses several embodiments to illustrate the concepts. Figure 1 A flowchart illustrating a hybrid knowledge retrieval enhancement generation method provided in this application is shown below. Figure 1 As shown, this hybrid knowledge retrieval enhancement generation method includes: S101, retrieve a query request for quality management knowledge related to prefabricated components.

[0026] Among them, the "Quality Management Knowledge Inquiry Request for Precast Components" refers to inquiries made by users in natural language regarding quality phenomena, compliance issues, fault diagnosis, or process consultations arising in any stage of the design, production, transportation, stacking, installation, and acceptance of precast building components (such as composite slabs, precast columns, and exterior wall panels). Examples include: "What is the allowable deviation in the mesh size of the reinforcing steel mesh during reinforcement installation?", "What are the causes of transverse cracks in prestressed composite slabs?", or "What are the quality acceptance standards for grouting connections of reinforcing steel sleeves?".

[0027] Specifically, the system can receive user-inputted quality management knowledge query requests through a graphical user interface, an API (Application Programming Interface), or a voice interaction module. Optionally, after obtaining the quality management knowledge query request, preprocessing can be performed, such as encoding conversion, removal of irrelevant characters, and normalization into a standardized string that can be processed in subsequent steps, i.e., the original query text.

[0028] In one possible implementation, after obtaining a quality management knowledge query request, information such as user role and project stage can be associated to provide context for subsequent retrieval.

[0029] S102, based on the semantic entities in the quality management knowledge query request, use a preset dense retrieval algorithm to calculate the dense similarity score of each node block in the quality management knowledge graph of prefabricated components.

[0030] First, it is necessary to identify semantic entities from the quality management knowledge query request. A semantic entity is a word or phrase in the text that carries core semantic information. For a quality management knowledge query request, such as "What is the allowable deviation of the mesh size of the reinforcing steel mesh during rebar installation?", the semantic entities are "rebar installation", "rebar mesh", "mesh size", and "allowable deviation". Specifically, a pre-defined named entity recognition model can be used to identify the quality management knowledge query request and obtain at least one corresponding semantic entity.

[0031] Subsequently, a pre-defined dense retrieval algorithm is used to calculate the degree of association between the semantic entities extracted from the query request and the deep semantic information represented by each node in the pre-constructed quality management knowledge graph, in order to evaluate the semantic relevance between the node and the query.

[0032] The quality management knowledge graph is a pre-defined structured semantic knowledge model that organizes all quality management knowledge related to prefabricated components of prefabricated buildings in the form of a graph. The quality management knowledge graph consists of nodes and edges: nodes represent various entities in quality management, such as specific prefabricated components (e.g., "composite slabs," "prefabricated exterior walls"), materials (e.g., "grouting material," "sealant"), processes (e.g., "mold assembly," "steam curing"), quality defects (e.g., "cracks," "honeycomb surface"), and technical standard clauses; edges represent the semantic relationships between these entities, such as "leads to," "applies to," "complies with," "belongs to," and "has attributes."

[0033] In a quality management knowledge graph, each node is transformed into a node block. A node block not only represents a single entity but also contains multiple related entities and the semantic relationships between them. Each node block, as a structural unit in the graph, carries deep semantic information between nodes.

[0034] For example, in a quality management knowledge graph, "reinforcement installation" might be a node representing the process of installing reinforcement bars in precast components. When "reinforcement installation" is converted into a node block, the node block not only contains "reinforcement installation" itself, but may also include: related technological steps, such as "reinforcement binding" and "reinforcement welding"; related material nodes, such as "reinforcement bars"; related technical standards, such as "reinforcement installation standards"; and related quality defect nodes, such as "reinforcement displacement" and "reinforcement corrosion". In this way, "reinforcement installation" as a node block contains more related information, helping the system to more comprehensively understand the semantics of the query. Specifically, a dense retrieval algorithm is used to calculate the similarity between the query semantic entity and the semantic content carried by the nodes in the quality management knowledge graph. Based on the similarity calculation, a dense similarity score for each node block is obtained. The level of the dense similarity score quantifies the degree of fit between the node block and the query request at the semantic level. The higher the score, the more relevant the node block is to the query request at the semantic level. For example, even if the semantics represented by a node block in the quality management knowledge graph is "tying the steel bar framework", and the query request only mentions "steel bar installation", the dense retrieval algorithm can still recognize a high degree of relevance due to the semantic association and give a corresponding high score to this node block.

[0035] In actual implementation, by calculating the dense similarity scores between the semantic entities extracted from the query and each node in the quality management knowledge graph, node blocks that are literally different but semantically relevant can be discovered, overcoming the limitations of traditional keyword matching and improving the accuracy of retrieval.

[0036] S103. According to the terms in the quality management knowledge query request, use a preset sparse retrieval algorithm to calculate the sparse similarity scores of each node block in the quality management knowledge graph.

[0037] First, the quality management knowledge query request needs to be processed to obtain the corresponding terms. Terms refer to the keywords obtained by segmenting the query request, filtering stop words, and performing词性标注 (word - type tagging). Specifically, a domain - specific dictionary can be used for word segmentation to split the query request into independent lexical units; then stop - word filtering is applied to remove虚词 (function words) such as "de", "le", "ma", etc. that have no practical meaning; then word - type tagging and screening are carried out, with emphasis on retaining terms such as nouns, verbs, and adjectives that carry substantial information. For example, after processing the query request "What is the allowable deviation of the mesh size of the steel bar mesh tied during steel bar installation?", the terms "steel bar", "installation", "tying", "steel bar mesh", "mesh", "size", "allowance", "deviation" are obtained.

[0038] Subsequently, use a preset sparse retrieval algorithm to calculate the direct correspondence between the terms extracted from the query request and the text descriptions associated with each node block in the pre - constructed quality management knowledge graph to evaluate the relevance between the node block and the query at the term level.

[0039] Specifically, use the sparse retrieval algorithm to calculate the matching degree between the terms in the query request and the description text of the nodes in the quality management knowledge graph. Based on the matching - degree calculation, the sparse similarity scores of each node block are obtained. The level of the sparse similarity score quantifies the degree of fit between the node block and the query request at the term level. The higher the score, the more direct and precise the literal match between the node block and the query request. For example, if the term in the query request is "steel bar", and the description text of a certain node in the quality management knowledge graph contains "steel bar tying construction", the sparse retrieval algorithm will recognize a strong relevance due to the high overlap of the terms and give a corresponding high score to this node block.

[0040] In practice, by calculating the sparse similarity scores of each node block in the quality management knowledge graph, the keywords and core concepts in the user query can be aligned, avoiding retrieval bias that may be caused by semantic generalization and improving the accuracy of retrieval.

[0041] It should be noted that there is no strict sequential dependency between S102 and S103 in the execution process. They can be executed in parallel or sequentially according to actual needs. Regardless of the execution order, their technical objectives are independent of each other.

[0042] S104. Calculate the fusion similarity score of each node block based on the dense similarity score and the sparse similarity score of each node block.

[0043] To address the potential biases caused by relying on a single retrieval dimension, a dense similarity score for each node block is used. Sparse similarity scores of each node block The nodes are then fused to obtain their fusion similarity scores. This yields a comprehensive score that reflects the relevance of each node block to the query request.

[0044] Specifically, the two types of scores can first be standardized, for example, by using a normalization method to map the original scores to the [0,1] interval; then, the fusion similarity score of each node block can be calculated using a weighted linear combination formula, as shown in formula (1): Formula (1) in and Adjustable weighting coefficients (satisfying) This is used to balance the contribution of semantic relevance and term matching in the final evaluation. For example, in scenarios requiring highly precise matching of normative clauses, it is appropriate to increase the weighting. The value can be increased when performing open fault diagnosis. The value is set to enhance semantic association. In one possible implementation, it can be set... , .

[0045] In practice, by fusing dense similarity scores and sparse similarity scores, not only is the semantic level of association strength preserved, but the accuracy of term matching is also enhanced. This fusion provides an entity basis for the generation of the final retrieval answer.

[0046] S105. Based on the fusion similarity score of each node block, the quality management knowledge graph is traversed to obtain the similarity score of each edge block in the quality management knowledge graph.

[0047] In a quality management knowledge graph, each edge is transformed into an edge block. An edge block not only represents the semantic relationship between two node blocks, but also includes the semantic strength and contextual information of the relationship between the two node blocks. An edge block is not merely a simple relationship connecting nodes, but also carries semantic information about the interactions between node blocks.

[0048] For example, in a quality management knowledge graph, there exists an edge representing the relationship between rebar installation and rebar mesh binding, such as the edge "rebar installation → rebar mesh binding". This edge itself only indicates a connection and does not contain any information about the strength or context of the relationship between "rebar installation" and "rebar mesh binding". However, when this edge is converted into an edge block, the edge block becomes "rebar installation leads to rebar mesh binding". This edge block not only represents the relationship between them (such as "leads to"), but also includes the semantic information of the relationship (i.e., "rebar installation" will trigger "rebar mesh binding").

[0049] To reveal the semantic relationships between nodes in the knowledge graph that match the query request, it is necessary to traverse the quality management knowledge graph using the fusion similarity score of each node as a basis, obtaining the similarity score of each edge block. Specifically, starting from each node, the search and extension are performed along the pre-defined connections between entities in the knowledge graph. During the search, each found relational edge block is recorded, and a score representing its importance is calculated for each edge block based on the fusion similarity scores of the node blocks at both ends of the edge block and the distance of the edge from the starting entity; this score is the edge block's similarity score. The similarity score of each edge block depends not only on the importance of the connected node block itself, but also on the position of the relationship in the overall knowledge graph.

[0050] By calculating the similarity scores of edge blocks, causal, conditional, and attribute relationships between different node blocks can be identified and evaluated. For example, the query "rebar installation" may be related to multiple node blocks, and the relationships between these node blocks (such as "rebar installation leads to rebar mesh binding") will be reflected through the similarity scores of the edge blocks, thus providing a more accurate logical connection for the final answer.

[0051] In practice, by calculating the similarity score of each edge block, the final search answer can reflect logical relationships such as cause and effect, condition and attribute, which not only improves the accuracy of the search, but also makes the generated answer more professional and explanatory.

[0052] S106. Based on the fusion similarity score of each node block and the similarity score of each edge block, generate the retrieval answer corresponding to the quality management knowledge query request.

[0053] Specifically, based on a preset filtering threshold, a set of node blocks with high fusion similarity scores and a set of edge blocks with high similarity scores can be extracted from all node blocks and edge blocks. These sets are then input into the Large Language Model (LLM) along with the quality management knowledge query request. In this way, the LLM can simultaneously obtain the user's original query intent and highly relevant entities and relationships (node ​​blocks and edge blocks) extracted from the knowledge graph, achieving enhanced hybrid knowledge retrieval generation. Using these inputs, the LLM organizes the discrete highly relevant entities and relationships into a coherent, accurate, and professionally contextualized natural language answer according to their semantic associations and logical structures.

[0054] The following describes this embodiment through a specific scenario.

[0055] For example, if a user's query is "What is the allowable deviation of the mesh size of the reinforcing mesh during rebar installation?", the semantic entities "rebar installation", "rebar mesh", "mesh size", and "allowable deviation" are extracted. Then, a dense retrieval algorithm is used to calculate the correlation between these semantic entities and the semantic information of each node in the quality management knowledge graph, obtaining the dense similarity score of each node in the quality management knowledge graph. Next, the query request is segmented into words, obtaining terms such as "rebar", "installation", "tying", "rebar mesh", "mesh", "size", "allowable", and "deviation". A sparse retrieval algorithm is used to calculate the direct correspondence between these terms and the text descriptions associated with each node block in the quality management knowledge graph, obtaining the sparse similarity score of each node block. This sparse similarity score is then fused to obtain the fused similarity score of each node block. Based on the fused similarity scores of each node block, the quality management knowledge graph is traversed to calculate the similarity score of each edge block. Finally, node blocks with high fused scores and edge blocks with high similarity scores are selected and input into the large language model along with the quality management knowledge query request to generate the final answer.

[0056] In this embodiment, by fusing dense and sparse retrieval, the similarity of each node block in the quality management knowledge graph at the query semantics and terminology levels is calculated and fused to obtain the fused similarity score of the node block. This score guides the traversal of the graph to further deduce the similarity score of the relational edge blocks. This ensures that the final answer not only includes the core entity in the query but also covers the key relationships, forming a logical chain of evidence. This directly solves the problems of fragmented knowledge and isolated retrieval information in the quality management of prefabricated buildings. It can provide structurally complete and traceable reference information for specific tasks such as quality defect diagnosis and process standard query, effectively supporting on-site decision-making. This significantly shortens the problem diagnosis time, improves the efficiency and accuracy of quality control, reduces rework and cost overruns caused by information omissions or misjudgments, and effectively improves the control level of quality management in prefabricated building projects.

[0057] In the above Figure 1 Based on the corresponding embodiments, in order to more clearly demonstrate the process of the dense retrieval algorithm, this application also provides a possible implementation of the dense retrieval algorithm in hybrid knowledge retrieval enhancement generation. Figure 2 This is a flowchart illustrating a dense retrieval algorithm for hybrid knowledge retrieval enhancement generation provided in this application. Figure 2 As shown, in S102 above, based on the semantic entities in the quality management knowledge query request, a preset dense retrieval algorithm is used to calculate the dense similarity score of each node block in the quality management knowledge graph of prefabricated components, including: S210 encodes semantic entities into dense vectors.

[0058] First, the semantic entities obtained from the quality management knowledge query request need to be vectorized and encoded. Specifically, each semantic entity is input into a pre-defined text embedding model. Through multi-layer neural network transformation, each entity is mapped to a high-dimensional real-number vector of a fixed dimension, called a dense vector. The dense vector represents the complete semantic features of the entity in mathematical space, and entities with similar semantics are closer in the vector space. For example, the vector directions of "tying steel mesh" and "steel mesh sheet" will be highly similar, while the vector of "concrete strength" will show a clear angle.

[0059] S220: Calculate the cosine similarity between the dense vector of the semantic entity and the dense vector of each node block to obtain the dense similarity score of each node.

[0060] Next, the dense vectors of semantic entities are compared with the pre-calculated dense vectors of each node block in the quality management knowledge graph to calculate the similarity score of each node block. Each node block has already generated a corresponding vector representation using the same embedding model during the knowledge graph construction phase.

[0061] The dense similarity score of each node is calculated using formula (2), which evaluates semantic consistency by measuring the difference in direction between two vectors.

[0062] Formula (2) in A dense vector representing the semantic entity of the query. A dense vector representing a knowledge graph node block. It is an index variable, representing the index of... and The dimension of each element in the text. It is the dimension of the vector, that is... and Dimensions Representing vectors In the Component values ​​in each dimension Representing vectors In the Each vector has multiple dimensions, representing different semantic information.

[0063] express and The cosine similarity between nodes is used to calculate the density similarity score of that node. . The calculation result ranges from [-1, 1], and the closer the value is to 1, the more similar the semantics are.

[0064] In practice, by calculating the density similarity score of each node block, semantically relevant node blocks can be quickly identified. This method ensures that the retrieved node blocks not only match the query literally but are also semantically related to the query content, improving the quality and accuracy of the search results.

[0065] In the above Figure 1 Based on the corresponding embodiments, in order to more clearly demonstrate the process of the sparse retrieval algorithm, this application also provides a possible implementation of the sparse retrieval algorithm in hybrid knowledge retrieval enhancement generation. Figure 3 This is a flowchart illustrating a sparse retrieval algorithm for hybrid knowledge retrieval enhancement generation provided in this application. Figure 3 As shown, in S103 above, based on the terms in the quality management knowledge query request, a preset sparse retrieval algorithm is used to calculate the sparse similarity score of each node block in the quality management knowledge graph, including: S310, calculate the inverse document frequency information of the terms based on the total number of documents corresponding to the quality management knowledge graph and the number of documents in which the terms appear.

[0066] Among them, the total number of documents refers to the number of documents included in the quality management knowledge graph. Each document can be an academic article, a technical report, an industry standard, etc.; the number of documents in which a term appears : refers to the term appears in how many documents. If the term appears in many documents, it indicates that the term has a relatively high prevalence in the knowledge graph. Conversely, it means that the term is relatively rare; the inverse document frequency IDF (Inverse Document Frequency) is an important indicator for measuring the rarity of a term in information retrieval and is calculated according to formula (3): Formula (3) The constants 10 and 0.5 in the formula help smooth the weights, avoiding the IDF being zero when a term appears frequently or the IDF being too high when the term is too rare.

[0067] In actual implementation, by calculating the inverse document frequency information of the term , important and discriminatory terms in the quality management knowledge graph can be identified, avoiding the influence of high-frequency and semantically weak terms on the analysis results and improving the effectiveness of information retrieval.

[0068] S320. Calculate the weight of the term according to the inverse document frequency information of the term, the named entity recognition information of the term, and the词性标记信息 of the term.

[0069] Among them, the named entity recognition information of the term is the weight of whether the term is a named entity, usually referring to proper nouns, key terms in a specific field, etc.; if the term is a named entity, the score is relatively high, representing the core role of the term in the field; for example, "ISO 9001" is an important standard in the field of quality management and has a relatively high weight, while function words such as "of" and "and" will not be recognized as named entities.

[0070] The词性标记信息 of the term is the词性标记得分 of the term , representing that the grammatical role has an impact on the importance of the term: nouns and proper nouns usually carry more information, so a higher POSTAG score will be given;词性 such as verbs, pronouns, and conjunctions have lower scores due to less information.

[0071] The weight of the term is calculated according to its inverse document frequency , named entity recognition and词性标记 and represents the term It should be noted that "词性标记信息" and "词性标记得分" in the original text seem to be incomplete or incorrect expressions. It might be better to have more accurate and clear descriptions in the original text for a more precise translation. Also, some tags like

[0066] etc. are likely specific to a certain system or format and are left unchanged as per the requirement. The importance of querying and matching knowledge graph node blocks. Calculate the term according to formula (4). weight : Formula (4) In practice, by combining inverse document frequency information, named entity recognition, and part-of-speech tagging of terms into the weight calculation, we can avoid some terms with low semantic contribution (such as conjunctions and pronouns) from occupying too much weight, so that terms with high semantic value (such as proper nouns and domain terms) can receive more attention, which helps to improve the retrieval accuracy of terms.

[0072] S330, calculate the sparse similarity score of each node block based on the weight of the terms.

[0073] Among them, the sparse similarity score of node blocks This reflects the relevance between the terms in the query request and the nodes in the knowledge graph, calculated using formula (5): Formula (5) in, This refers to the set of terms used in a quality management knowledge query request. This represents a set of node blocks in a quality management knowledge graph. For a given node block, based on each term in the query request... Calculate its , , Fractions, then multiplied together to get the term. The weights of all terms in the query request are summed to obtain the sparse similarity score of the node block.

[0074] In practice, this method of comprehensively evaluating term weights can more accurately identify node blocks related to the query request, improving query accuracy. In particular, by combining information such as syntax and entity recognition, it can better handle complex terms and different types of node blocks in the query, improving the robustness and flexibility of the query.

[0075] In the above Figure 1 Based on the corresponding embodiments, in order to more clearly demonstrate the process of obtaining the similarity scores of each edge, optionally, in S105 above, the quality management knowledge graph is traversed according to the fused similarity scores of each node block to obtain the similarity scores of each edge block in the quality management knowledge graph, including: S510: Based on the fusion similarity score of each node block and the preset jump distance, a preset distance decay function is used to traverse the quality management knowledge graph to obtain the similarity score of each edge block.

[0076] The jump distance refers to the connection depth or path length between nodes in a quality management knowledge graph when traversing it; it is a preset value. In practical applications, some nodes may be more directly related to the query request, while others may be indirectly related to the query request through several intermediate nodes. The jump distance is used to measure the "indirectness" of the relationship between two nodes.

[0077] The preset distance decay function is used to adjust the similarity that decreases as the jump distance increases during traversal. Typically, the semantic association between node blocks gradually weakens as the jump distance increases. Therefore, a decay function is used to reduce the influence of more distant nodes, thereby more accurately evaluating node blocks relevant to the query request. The preset distance decay function is expressed by formula (6): Formula (6) in, The similarity score is the score for the edge pieces. This represents the fusion similarity score of the node blocks.

[0078] A jump distance is defined as the minimum distance between the currently explored node block and the target node block connected by the edge currently being computed. Here, the currently explored node block is the node block corresponding to a semantic entity in the query request, and the target node block is the node block related to the query request. This is to ensure that even the closest node block (i.e. This also prevents the similarity score from becoming excessively high. By employing a jump distance decay function, the influence of node blocks relatively far from the query request can be effectively reduced, thereby ensuring that the edge blocks in the retrieval results reflect the relationships most relevant to the query semantics.

[0079] The following describes this embodiment through a specific scenario.

[0080] Suppose a user's query is "What is the allowable deviation in the mesh size of the reinforcing mesh during rebar installation?", which contains the semantic entity "rebar installation." The quality management knowledge graph contains nodes "rebar installation" and "rebar mesh tying." The node "rebar mesh tying" relates to this semantic entity... In the quality management knowledge graph, the edge is "reinforcing bar installation → reinforcing mesh binding"; the jump distance from the query request entity to the reinforcing mesh binding node block is... (Indicating there is an intermediate node between them), the similarity score of this edge piece is calculated as follows:

[0081] this This represents the similarity score of the edge block "rebar installation → rebar mesh binding".

[0082] In practice, starting with the node block corresponding to the entity in the query request, the algorithm traverses the edges of the quality management knowledge graph, recording each edge block encountered. For each edge block, the fusion similarity score between the starting node block and the target node block is calculated, and the overall similarity score is further calculated based on the jump distance and decay function. This edge block similarity score calculation ensures that not only directly related node blocks can be identified, but also indirectly related node blocks can be identified based on semantic relationships within the knowledge graph. This allows the retrieval results to better reflect logical connections such as causal relationships, conditional relationships, and attribute relationships, making the retrieval results more in-depth and professional.

[0083] In the above Figure 1 Based on the corresponding embodiments, in order to more clearly demonstrate the process of generating search answers, this application also provides a possible implementation method for generating search answers in hybrid knowledge retrieval enhancement generation. Figure 4 This is a flowchart illustrating the process of generating search answers in a hybrid knowledge retrieval enhancement method provided in this application. For example... Figure 4 As shown, in S106 above, the retrieval answer corresponding to the quality management knowledge query request is generated based on the fusion similarity score of each node block and the similarity score of each edge block, including: S610, based on the fusion similarity score and importance evaluation value of each node block, perform semantic alignment on each node block to obtain the target score of each node block.

[0084] In quality management queries, the semantic similarity of node blocks and their importance in the graph determine their relevance to the query. To ensure that the most relevant nodes are found, the target score for each node block must first be calculated.

[0085] The target score for each node block was obtained by calculating its similarity score and its importance assessment value in the graph. This target score combines the relevance of a node with its position in the knowledge graph, ensuring that the subsequent retrieval process prioritizes the most relevant and important node blocks.

[0086] Among them, the importance assessment value of each node block in the quality management knowledge graph is obtained through... It is calculated using the PageRank algorithm. PageRank is a method for evaluating the importance of nodes in a graph based on their connectivity. Its calculation formula is formula (7): Formula (7) in, It is the damping factor, usually with a value of 0.85, representing the probability of a jump; It is a collection of nodes in a quality management knowledge graph; Indicates the linked node block The set of node blocks, which refers to all pointers to node blocks. Other node blocks; Indicates the outgoing node block A set of node blocks, this set of node blocks The number of other node blocks connected.

[0087] PageRank reflects the importance of nodes by considering their connectivity. Each node's PageRank score represents its influence in the knowledge graph; the higher the score, the greater its weight and importance in the graph.

[0088] Target score of the final node block By using the similarity scores of node blocks and importance assessment value The combined results show that the calculation formula is formula (8): Formula (8) Formula (8) embodies the idea of ​​semantic alignment of node blocks: the similarity of node blocks is combined with their importance in the knowledge graph to ensure that the retrieved node blocks not only match the query semantically, but also take into account their importance in the knowledge graph.

[0089] In practice, semantic alignment enables the retrieval of nodes that are both highly semantically matched to the query request and of high importance in the quality management knowledge graph, ensuring that the retrieval results are more accurate and relevant, meeting the needs of quality management queries.

[0090] S620: Based on the similarity score and edge weight of each edge block, perform semantic alignment on each edge to generate the target score for each edge block.

[0091] Edge blocks represent the relationships between nodes, and the strength and type of these relationships are crucial for generating query answers. Similar to node blocks, edge blocks also need to be evaluated based on their similarity scores and weights. This step ensures that the relationships between nodes are correctly understood during the retrieval process, thereby providing accurate query answers.

[0092] The target score for each edge block was obtained by calculating the similarity score and weight of the edge blocks. This target score reflects the semantic relevance of the edge block and the strength of its relationship in the knowledge graph, ensuring that the relationships between nodes can be effectively evaluated.

[0093] Among them, the weights of each edge block in the quality management knowledge graph This indicates the importance and strength of edge blocks in the knowledge graph. The calculation of edge block weights takes into account the frequency of edge blocks. Frequency-based weights are calculated using a simple addition mechanism, that is, the initial weight of each edge block is 1, and its weight increases by 1 each time it reappears. The weight value represents the importance assessment value of the edge in the prefabricated component quality management knowledge graph. The higher the value, the higher the importance of the relationship. Target score of the final edge piece By using edge similarity scores and edge weight Based on the calculations, the formula is (9): Formula (9) Formula (9) embodies the idea of ​​edge block semantic alignment: the similarity score of the edge block is combined with its edge weight in the knowledge graph to ensure that not only highly similar relationships are selected during retrieval, but also those relationships with strong connections in the graph are selected.

[0094] In practice, the target score calculation of edge blocks ensures that the most relevant and important relationships are selected, thereby improving the accuracy and relevance of the retrieval results. By combining weight and similarity, it can effectively handle complex relationships between node blocks in the retrieval process.

[0095] S630: Based on the target scores of each node block and each edge block, determine the target node block and target edge block with the highest score from the quality management knowledge graph.

[0096] After calculating the target scores for each node block and edge block, these scores need to be further filtered to determine which node blocks and edge blocks are most relevant to the quality management query request. The optimal solution, i.e., the nodes and relationships with the highest scores, is selected from the node blocks and edge blocks to provide accurate constituent elements for the final answer.

[0097] By sorting the target scores of each node block and each edge block, the target node block and target edge block with the highest scores are determined. These target node blocks and target edge blocks best reflect the intent of the query request and can provide the most relevant information and relationships to the query request.

[0098] In one possible implementation, based on the target scores of each node block and each edge block, multiple (Top N) high-scoring node blocks and edge blocks are selected from the quality management knowledge graph as target node blocks and target edge blocks, instead of just selecting the one with the highest score. By selecting the Top N node blocks and edge blocks, the accuracy and completeness of the answer can be improved, information loss can be avoided, and users can also gain more perspectives from the final answer, thus improving the comprehensiveness of the query results.

[0099] In practice, by selecting target nodes and target edge blocks, the focus of the search is concentrated on the most relevant and accurate parts, ensuring that the final search answer contains the most critical nodes and the most important relationships, which greatly improves the accuracy and effectiveness of the search results.

[0100] S640 generates the search answer based on the target node block and the target edge block.

[0101] By combining the target node block and the target edge block, a complete search answer is generated. This search answer not only conforms to the semantics of the query request, but also effectively reflects the relationships between node blocks in the quality management knowledge graph, helping users obtain accurate quality management-related knowledge.

[0102] Specifically, the extracted target node blocks and target edge blocks are input into the Large Language Model (LLM) along with the quality management query request. During this process, the LLM not only receives the user's query intent but also obtains the target node blocks and target edge blocks extracted from the knowledge graph, forming a powerful input set. The LLM utilizes this input, combined with its powerful natural language generation capabilities, to understand the context of the query request and, based on the semantic information provided by the target node blocks and target edge blocks, generates an answer that meets quality management requirements. This answer not only accurately reflects the intent of the query request but also integrates information from the node blocks and edge blocks in the quality management knowledge graph, providing a high-quality, structured response.

[0103] In one possible implementation, based on semantic association and logical structure, LLM will eventually generate a complete and accurate natural language answer that not only answers the user's query but also ensures the accuracy and readability of the answer in a professional context.

[0104] In this embodiment, the most relevant nodes and relationships—namely, target node blocks and target edge blocks—can be progressively filtered out based on the similarity of node blocks and edge blocks and their importance in the knowledge graph. This helps to find the most accurate answers in complex quality management knowledge graphs.

[0105] In the above Figure 4Based on the corresponding embodiments, in order to more clearly demonstrate the process of generating search answers, this application also provides a possible implementation method for generating search answers in hybrid knowledge retrieval enhancement generation. Figure 5 This is a schematic diagram illustrating the process of generating search answers in another hybrid knowledge retrieval enhancement method provided in this application. For example... Figure 5 As shown, in step S640 above, the process of generating the search answer based on the target node block and the target edge block includes: S6401, determine the target community report from the quality management knowledge graph based on the target node block and the target edge block.

[0106] In the quality management knowledge graph, knowledge is organized into several communities. Each community represents a set of closely related knowledge points within a topic. For example, nodes related to "quality control" and their associated edge nodes might form a "quality management" community. Each community contains multiple nodes (such as "quality control methods," "construction project management," etc.) and edge nodes (such as relationships like "impact" and "dependency") related to that topic. These nodes and edges collectively describe a complete knowledge domain.

[0107] In the previous steps, target node blocks and target edge blocks have been selected based on the target scores. In this step, we will use these target node blocks and target edge blocks to identify the communities related to them in the knowledge graph.

[0108] Each community will have a pre-generated community report that highlights the community's core characteristics and domain-related insights. Core characteristics refer to the most important nodes in the community and the relationships between them; domain-related insights refer to key insights into the field of quality management, such as how to implement quality control at different project phases and how to manage risks. Additionally, each community report has a unique community ID and is pre-indexed for quick access in subsequent searches.

[0109] Specifically, each community compares its contained node blocks and edge blocks with the target node block and target edge blocks to select the most relevant community. For example, if the target node block is "quality control method", then the community containing this node block and its related edge blocks will be searched as the target community, thus obtaining the target community report.

[0110] In practice, by combining target node blocks and target edge blocks, relevant community reports can be accurately identified, ensuring that the answers provided are more targeted. At the same time, by extracting relevant community reports from the knowledge graph, the most relevant topics can be quickly focused on, avoiding redundant information.

[0111] S6402, Generate search answers based on the target community report.

[0112] Specifically, at this stage, the target node blocks, target edge blocks, target community reports, and quality management query requests are input into the Large Language Model (LLM). This information collectively provides the LLM with sufficient context and expertise, enabling it to generate accurate answers. Upon receiving all input, the LLM identifies the specific semantics of the target node blocks and edge blocks, combining this information with the expertise from the target community reports. By integrating the relationships between node blocks and edge blocks with the insights from the community reports, the LLM organizes this information and generates a structured, coherent answer. This process ensures that the generated answer is not only logically clear but also highly relevant to the user's query request. Finally, the LLM organizes all the information into a natural language answer that meets quality management requirements.

[0113] For example, for a query like "How to improve quality control standards in construction projects," LLM will combine the relevance of target node blocks (such as "quality control methods") and edge blocks (such as "impacts"), leveraging insights from target community reports on "improving construction technology," to generate a well-structured and highly professional answer. In one possible implementation, LLM will also ensure the accuracy and actionability of the answer based on standards and best practices in the field of quality management.

[0114] In one possible implementation, based on the target scores of each node block and edge block, multiple (Top N) high-scoring node blocks and edge blocks are selected from the quality management knowledge graph. Based on these high-scoring node blocks and edge blocks, multiple target community reports are determined. Finally, these multiple target node blocks, multiple target edge blocks, multiple community reports, and the quality management query request are input into the LLM. This not only helps understand the relationships between the multiple target node blocks and edge blocks but also integrates information from the multiple community reports, ensuring that the generated answer considers perspectives and insights from multiple domains. Simultaneously, the LLM generates a structured and coherent answer based on this information, covering all aspects of the query while avoiding information duplication or redundancy. In this approach, based on multiple target node blocks, multiple target edge blocks, and multiple community reports, the LLM outputs a more multi-dimensional, professional, and in-depth answer that not only accurately answers the query request but also provides broader domain knowledge and insights, ensuring the answer is more practical.

[0115] In practice, since community reports gather core information from a specific domain within a knowledge graph, combining these core insights with the community reports can provide more professional answers that meet specific needs in the field of quality management. This makes the generated answers more relevant to the user's query requirements, thereby improving the user experience and satisfaction.

[0116] In this embodiment, community reports in the quality management knowledge graph can be effectively utilized to generate high-quality search answers from core features and insights across multiple domains, ensuring that the final search results can accurately answer practical questions in the field of quality management.

[0117] In the above Figure 1 Based on the corresponding embodiments, in order to more clearly demonstrate the process of obtaining a quality management knowledge graph, this application also provides a possible implementation method for constructing a quality management knowledge graph in a hybrid knowledge retrieval enhancement generation process. Figure 6 This is a schematic diagram illustrating the process of constructing a quality management knowledge graph in a hybrid knowledge retrieval enhancement generation method provided in this application. For example... Figure 6 As shown, based on S101-S106, it also includes: S710 converts the quality management knowledge document for prefabricated components into text knowledge blocks.

[0118] Knowledge documents on quality management of prefabricated components are generally in .pdf and .docx formats. First, a pre-defined multimodal document processing pipeline is used to process the documents, converting the text, tables, and graphics into actionable knowledge blocks. Specifically, for .docx documents: these documents can be directly parsed into an XML structure, allowing the extraction of text paragraphs, tables, and image content; for .pdf documents: due to their more complex format, Optical Character Recognition (OCR) technology must be used to extract the text content.

[0119] For table content extraction, not only is the table structure extracted from the document, but the data within the table is also converted into text knowledge blocks. For example, the data in each row and column of the table is extracted separately and represented as a text knowledge block, facilitating subsequent processing and analysis. Specifically, table content is converted into structured text using Table Structure Recognition (TSR) technology, preserving the row and column relationships of the table and converting it into knowledge blocks conforming to text formatting. For image content in the document, image recognition technology (such as OCR or image annotation) is used to convert the text information in the image into text content. If the image contains charts or flowcharts, key data or process steps are extracted and converted into relevant text knowledge blocks.

[0120] Furthermore, each information unit extracted from the document (text, tabular data, image descriptions, etc.) is transformed into an independent text chunk. Subsequently, all extracted text chunks are standardized into a uniform format to facilitate subsequent processing and analysis; these formatted text chunks will serve as input for subsequent steps.

[0121] In practice, it can not only process ordinary text content, but also effectively extract and transform table and image data, providing basic information for subsequent map construction.

[0122] S720: Based on the text knowledge block, the first prompt word is generated using the first preset structured indicator word template.

[0123] The first preset structured indicator template is used to extract relevant entities and relationships from text knowledge blocks. The first preset structured indicator template mainly includes three parts: task description; processing steps; and standardized output format instructions.

[0124] Specifically, the task description requires identifying entities and relationships from a given block of textual knowledge.

[0125] The processing steps include: 1. Entity Identification: Based on the information in the text knowledge block, identify all entities related to quality management control. For each entity, the following information needs to be extracted: Entity Name (the specific name of the entity in the text), Entity Type (selected from the following types: precast component, material, tolerance, method, quality issue, clause), and Entity Description (a general description of the entity's attributes and activities). The format for identified entities is (Entity|Entity Name|Entity Type|Entity Description); 2. Relationship Identification: Identify entity pairs (source entity, target entity) in the text knowledge block and extract the relationships between them. For each relationship, the following information needs to be extracted: Source Entity (the name of the entity from which the relationship originates), Target Entity (the name of the entity to which the relationship points), Relationship Description (explaining why the two entities are related), and Relationship Strength (a numerical score representing the strength of the relationship, a preset value). The format for identified relationships is (Relationship|Source Entity|Target Entity|Relationship Description|Relationship Strength).

[0126] The output format is: a list that combines all entities and relationships identified in steps 1 and 2 of the processing procedure.

[0127] Subsequently, based on each text knowledge block, the first preset structured indicator word template is used to fill the corresponding position of the text knowledge block to be processed into the template, thereby obtaining the first prompt word, which guides the model to extract entities and relationships from the text.

[0128] In practice, structured prompt word templates provide clear guidance, ensuring that the large model can accurately identify entities and relationships in the text.

[0129] S730, based on the first prompt word, uses a preset large model to generate semantic entities in the text knowledge block and the relationships between semantic entities.

[0130] The first cue word is input into a pre-defined large model, such as GPT (Generative Pre-trained Transformer) or BERT (Bidirectional Encoder Representations from Transformers). The first cue word contains structured information about relevant entities and relationships extracted from textual knowledge blocks, providing guiding context and helping the large model identify semantic entities in the text and the relationships between them.

[0131] The large model parses the semantic entities and relationships in the text knowledge blocks based on the structured information in the prompt words. Each entity and relationship is extracted and output in a list of tuples to ensure the clarity and consistency of the information structure.

[0132] Specifically, firstly, all semantic entities conforming to predefined types are identified and extracted; secondly, from the identified entities, all entity pairs with clear semantic relationships are found, and the relationships between them are determined, while relationship descriptions and strength information are extracted. The final output list of tuples is, for example, (quality defects, precast composite slabs, grout leakage at slab joints), (inspection methods, precast component installation, string lines and measurements), etc.

[0133] If a text knowledge block contains multiple entities and multiple relationships between them, multiple tuples will be generated, each tuple representing an independent entity-relation pair, thus obtaining the semantic entities in the text knowledge block and the relationships between them.

[0134] S740, construct the subgraph corresponding to the text knowledge block based on the semantic entities in the text knowledge block and the relationships between the semantic entities.

[0135] The semantic entities and relationships between them in a text knowledge block are transformed into graph-structured data. Specifically, each identified semantic entity is mapped to a node in the graph, with node attributes including entity name, type, and description. Each identified relationship is mapped to an edge connecting two corresponding entity nodes, with edge attributes including relationship description and relationship strength. Thus, the local knowledge contained in a single text knowledge block is constructed into a structured subgraph that reflects the entity network relationships within the original text knowledge block.

[0136] For example, within a sub-drawing, nodes represent precast composite slabs, grout leakage at slab joints, installation of precast components, string lines, and measurements, while edges represent precast composite slabs → (quality defects) grout leakage at slab joints, and installation of precast components → (inspection methods) string lines and measurements.

[0137] Furthermore, when the same relationship between the same entity pairs appears in multiple text blocks, these duplicate entity relationship pairs will be merged, that is, identical subgraphs will be merged, and the frequency of the relationship will be retained as the basis for the relationship strength. For example, if the relationship between "construction method" and "construction project management" appears multiple times, they will be merged into one relationship, and this relationship will be assigned a higher weight based on its frequency of occurrence.

[0138] In practice, by constructing a subgraph for each text knowledge block, the representation of entities and relationships in each text block can be refined, providing data support for subsequent merging and integration of the graphs. S750 merges the subgraphs corresponding to each text knowledge block to generate a quality management knowledge graph.

[0139] Specifically, the subgraphs generated from different text knowledge blocks will be merged to create a complete knowledge graph. During the merging process, entity semantics are first merged. Using methods such as semantic similarity calculation, it is determined whether nodes from different subgraphs point to the same entity (e.g., "PC wall panel" and "precast concrete exterior wall panel"), and these nodes are merged into a unified entity node, aggregating all its attributes and relationships. Subsequently, relationships are merged. For edges connecting merged entities, frequency-based weight calculations are performed to obtain the edge weight for each edge. Specifically, the initial weight of each edge is set to 1, and its weight increases by 1 whenever identical relations (relations of the same type connecting the same entity pairs) from different subgraphs are merged. Meanwhile, for each node, the PageRank algorithm can be used to calculate its importance evaluation value.

[0140] Ultimately, through this global merging and integration operation, all subgraphs are combined into a unified, de-redundant, weighted quality management knowledge graph. This graph fully covers the core knowledge in the quality management knowledge document. In this embodiment, by converting text, tables, and images in the document into structured text blocks, key information can be extracted efficiently and its structured nature can be ensured, facilitating subsequent analysis. Entity and relationship extraction through prompt word generation and large-scale modeling reduces manual intervention and improves the accuracy of entity recognition. Finally, a knowledge graph is generated by merging subgraphs, accurately reflecting the core concepts and relationships in the field of quality management. This not only improves data processing efficiency but also enhances the accuracy of the knowledge graph, ensuring reliable support for knowledge retrieval, intelligent analysis, and decision-making, and greatly improving the intelligence and operability of quality management.

[0141] In the above Figure 6 Based on the corresponding embodiments, in order to more clearly demonstrate the process of obtaining community reports, this application also provides a possible implementation of generating device reports in hybrid knowledge retrieval enhancement generation. Figure 7 This is a schematic diagram illustrating the process of generating a community report in a hybrid knowledge retrieval enhancement method provided in this application. For example... Figure 7 As shown, based on S710-S750, it also includes: S810 performs entity clustering on the quality management knowledge graph to obtain the knowledge communities corresponding to the quality management knowledge graph.

[0142] First, a pre-defined clustering algorithm is used to cluster entities in the knowledge graph, forming knowledge communities corresponding to the quality management knowledge graph with high internal connectivity. To evaluate the quality of the community structure, modularization is introduced. The function is used as an evaluation index, and its definition is as shown in formula (10): Formula (10) in, This is the sum of all edge weights in the graph, where each edge weight represents a frequency-based weight score. ; Represents a node and nodes Edge weights between them; and Represents nodes and nodes The sum of the weights of connected edges; and They are nodes and nodes The community This is the Kronecker function, which takes the value 1 when two nodes belong to the same community, and 0 otherwise.

[0143] Modular The higher the value, the more effective the community division, meaning that the nodes within the same community are closely connected, while the connections between different communities are weaker. This helps to identify highly cohesive entity clusters and provides a foundation for subsequent community analysis and report generation.

[0144] S820, based on the knowledge community, uses the second preset structured indicator word template to generate the second prompt word.

[0145] The second preset structured indicator template is used to extract relevant community information from the knowledge community. The second preset structured indicator template mainly includes three parts: task description; processing steps; and standardized output format instructions.

[0146] Specifically, the task description requires writing a comprehensive community analysis report from the knowledge community’s list of entities, relationships, and optional association claims. This report will assist quality inspectors and project managers in identifying quality issues, tracing responsibility, proposing solutions, and supporting decision-making.

[0147] It should be noted that the association declaration is generated by the large model based on the first cue word, indicating the specific relationships between various entities, and is used to help the model understand the interactions between different entities in each community.

[0148] The processing steps include: 1. Identifying the entity list: Extracting all entities from the knowledge community; 2. Identifying relationships: Identifying the relationships between entities within the knowledge community; 3. Generating report content: Generating report content based on entities and relationships, including a title (naming the community title according to representative entities within the community), a summary (outlining the core entities, relationships, and quality control topics involved in the community), an impact severity rating (assigning impact severity scores to entity relationships within the community), and detailed findings (listing 5 to 10 key findings within the community, each with a brief summary and a detailed explanatory paragraph), etc.

[0149] The output format is: structured text identified by "[Community Report]".

[0150] Subsequently, for each community, the identified knowledge community information is filled into the corresponding positions of the second preset structured prompt word template to obtain the second prompt word, which will provide detailed guidance for the subsequent generation of community reports by the large model.

[0151] S830, based on the second prompt word, uses a preset large model to generate a community report for the knowledge community.

[0152] The second prompt word is input into a pre-defined large model. The large model utilizes the structured information provided by the second prompt word to understand the task description, entity relationships, impact severity rating, and other content, highlighting its core features and domain insights to generate a community report for the knowledge community. Each report is identified by a unique community ID and stored as an independent data block, supporting subsequent semantic retrieval and ensemble analysis.

[0153] In this embodiment, different knowledge communities can be efficiently identified from the quality management knowledge graph, and detailed reports for each community can be generated. The generated reports can reflect the core knowledge and professional insights in the community, which not only improves the quality of the knowledge graph, but also provides strong support for intelligent decision-making and analysis.

[0154] To describe this application more clearly, a specific embodiment is described below.

[0155] First, a PDF specification document (e.g., "DBJT 15 171-201 Guangdong Province Prefabricated Concrete Building Construction Quality Acceptance Specification") is received and parsed to extract text knowledge blocks. Based on the extracted text knowledge blocks, using a first preset structured cue word template and a large model, all relevant entities and relationships are identified and constructed into a list of tuples. These tuple lists are mapped to nodes and edges in a graph, with nodes and edges corresponding to entities and relationships, forming the basic structure of the graph and creating several subgraphs. Simultaneously, nodes and edges in each graph are converted into corresponding node blocks and edge blocks, and through a preset text embedding model, these node blocks and edge blocks are converted into dense vectors. Then, the constructed subgraphs are merged to obtain a complete quality management knowledge graph for prefabricated components. For the generated quality management knowledge graph, relevant entities are clustered to obtain several communities, and using a second preset structured cue word template and a large model, several community reports are generated.

[0156] When a user submits a query (e.g., "What is the allowable deviation in the mesh size of the reinforcing mesh during rebar installation?"), a dense search is first performed based on the semantic entities in the query request, followed by a sparse search based on the terms in the query request. This results in a fused similarity score for each node block based on both dense and sparse similarity scores. Next, the process traverses the quality management knowledge graph. For example, when the query involves "allowable deviation in the mesh size of the reinforcing mesh", the process first finds nodes directly associated with "reinforcing mesh", such as "tying process requirements", and then continues to explore its neighboring nodes (e.g., "construction quality acceptance standards" or "rebar installation specifications"). By applying a distance decay function, the contribution of neighboring nodes is ensured to gradually decrease, resulting in the similarity score for each edge block.

[0157] Finally, the above scores are semantically aligned to obtain the node blocks and edge blocks most relevant to the query request, thereby matching relevant community reports. These reports, along with the query request, are then input into the large model to obtain the answer. For example, the answer might be: According to the "DBJT 15 171-201 Guangdong Province Prefabricated Concrete Building Construction Quality Acceptance Specification" and related materials, the allowable deviation of the mesh size of the reinforcing steel mesh during installation is... .

[0158] The following describes a hybrid knowledge retrieval enhancement generation apparatus and electronic device provided in this application, the specific implementation process and technical effects of which are described above and will not be repeated below.

[0159] Figure 8 A schematic diagram of a hybrid knowledge retrieval enhancement generation device provided in this application is shown below. Figure 8As shown, the hybrid knowledge retrieval enhancement generation device includes: Module 1000 is used to retrieve query requests for quality management knowledge related to prefabricated components. The calculation module 2000 is used to calculate the dense similarity score of each node block in the quality management knowledge graph of prefabricated components based on the semantic entities in the quality management knowledge query request and using a preset dense retrieval algorithm. The calculation module 2000 is also used to calculate the sparse similarity score of each node block in the quality management knowledge graph based on the terms in the quality management knowledge query request and using a preset sparse retrieval algorithm. The calculation module 2000 is also used to calculate the fusion similarity score of each node block based on the dense similarity score of each node and the sparse similarity score of each node block; The calculation module 2000 is also used to traverse the quality management knowledge graph based on the fusion similarity score of each node block to obtain the similarity score of each edge block in the quality management knowledge graph. The generation module 3000 is used to generate the retrieval answer corresponding to the quality management knowledge query request based on the fusion similarity score of each node block and the similarity score of each edge block.

[0160] Optionally, the calculation module 2000 is also used to encode semantic entities into dense vectors; calculate the cosine similarity between the dense vector of the semantic entity and the dense vector of each node block to obtain the dense similarity score of each node block.

[0161] Optionally, the calculation module 2000 is also used to calculate the inverse document frequency information of a term based on the total number of documents corresponding to the quality management knowledge graph and the number of documents in which the term appears; calculate the weight of a term based on the inverse document frequency information of the term, the named entity recognition information of the term, and the part-of-speech tagging information; and calculate the sparse similarity score of each node block based on the weight of the term.

[0162] Optionally, the calculation module 2000 is also used to traverse the quality management knowledge graph based on the fusion similarity score of each node block and the preset jump distance, using a preset distance decay function, to obtain the similarity score of each edge block.

[0163] Optionally, the calculation module 2000 is further configured to perform semantic alignment on each node block based on the fusion similarity score and the importance evaluation value of each node block to obtain the target score of each node block; perform semantic alignment on each edge block based on the similarity score and the edge weight of each edge block to generate the target score of each edge block; and determine the target node block and target edge block with the highest score from the quality management knowledge graph based on the target scores of each node block and the target scores of each edge block.

[0164] Optionally, the generation module 3000 is also used to determine the target community report from the quality management knowledge graph based on the target node block and the target edge block; and to generate the search answer based on the target community report.

[0165] Optionally, the generation module 3000 is also used to convert the quality management knowledge document of the prefabricated components into text knowledge blocks; generate first prompt words based on the text knowledge blocks using a first preset structured indicator word template; generate semantic entities and relationships between semantic entities in the text knowledge blocks using a preset large model based on the first prompt words; construct subgraphs corresponding to the text knowledge blocks based on the semantic entities and relationships between semantic entities in the text knowledge blocks; and merge the subgraphs corresponding to each text knowledge block to generate a quality management knowledge graph.

[0166] Optionally, the generation module 3000 is also used to perform entity clustering on the quality management knowledge graph to obtain the knowledge community corresponding to the quality management knowledge graph; based on the knowledge community, a second prompt word is generated using a second preset structured indicator word template; based on the second prompt word, a community report of the knowledge community is generated using a preset large model.

[0167] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more digital signal processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).

[0168] Figure 9 This is a schematic diagram of an electronic device provided in this application. The device may be a computing device or a server with computing processing capabilities.

[0169] This application does not limit the product form of the electronic device 10. When the electronic device 10 integrates a hybrid knowledge retrieval enhancement generation program, it can be used as a hybrid knowledge retrieval enhancement generation device to execute the above method embodiments.

[0170] The electronic device 10 includes a processor 11, a storage medium 12, and a bus 13. The storage medium 12 stores program instructions executable by the processor 11. When the electronic device 10 is executed, the processor 11 communicates with the storage medium 12 via the bus 13, and the processor 11 executes the program instructions to perform the above-described method embodiment. The specific implementation and technical effects are similar and will not be described in detail here.

[0171] Optionally, this application also provides a program product, such as a computer-readable storage medium, including a program that, when executed by a processor, performs the above-described method embodiments.

[0172] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0173] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0174] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0175] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0176] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A hybrid knowledge retrieval and enhanced generation method, characterized in that, The method includes: Request a query for quality management knowledge related to prefabricated components; Based on the semantic entities in the quality management knowledge query request, a preset dense retrieval algorithm is used to calculate the dense similarity score of each node block in the quality management knowledge graph of the prefabricated components. Based on the terms in the quality management knowledge query request, a preset sparse retrieval algorithm is used to calculate the sparse similarity score of each node block in the quality management knowledge graph. Based on the dense similarity scores of each node and the sparse similarity scores of each node block, the fusion similarity score of each node block is calculated. Based on the fusion similarity scores of each node block, the quality management knowledge graph is traversed to obtain the similarity scores of each edge block in the quality management knowledge graph; Based on the fusion similarity score of each node block and the similarity score of each edge block, the retrieval answer corresponding to the quality management knowledge query request is generated.

2. The method according to claim 1, characterized in that, The step of calculating the dense similarity score of each node block in the quality management knowledge graph of the prefabricated components based on the semantic entities in the quality management knowledge query request using a preset dense retrieval algorithm includes: The semantic entities are encoded into dense vectors; Calculate the cosine similarity between the dense vector of the semantic entity and the dense vector of each node block to obtain the dense similarity score of each node block.

3. The method according to claim 1, characterized in that, The step of calculating the sparse similarity score of each node block in the quality management knowledge graph based on the terms in the quality management knowledge query request using a preset sparse retrieval algorithm includes: Based on the total number of documents corresponding to the quality management knowledge graph and the number of documents in which the term appears, calculate the inverse document frequency information of the term; The weight of the term is calculated based on the inverse document frequency information of the term, the named entity recognition information of the term, and the part-of-speech tagging information; Based on the weights of the terms, the sparse similarity score of each node block is calculated.

4. The method according to claim 1, characterized in that, The step of traversing the quality management knowledge graph based on the fusion similarity scores of each node block to obtain the similarity scores of each edge block in the quality management knowledge graph includes: Based on the fusion similarity score of each node block and the preset jump distance, a preset distance decay function is used to traverse the quality management knowledge graph to obtain the similarity score of each edge block.

5. The method according to claim 1, characterized in that, The step of generating the retrieval answer corresponding to the quality management knowledge query request based on the fusion similarity score of each node block and the similarity score of each edge block includes: Based on the fusion similarity score and importance evaluation value of each node block, semantic alignment is performed on each node block to obtain the target score of each node block; Based on the similarity score and edge weight of each edge block, semantic alignment is performed on each edge block to generate a target score for each edge block; Based on the target scores of each node block and each edge block, the target node block and target edge block with the highest scores are determined from the quality management knowledge graph. The retrieval answer is generated based on the target node block and the target edge block.

6. The method according to claim 5, characterized in that, The step of generating the search answer based on the target node block and the target edge block includes: Based on the target node block and the target edge block, the target community report is determined from the quality management knowledge graph; The search answer is generated based on the target community report.

7. The method according to claim 1, characterized in that, The method further includes: The quality management knowledge document for the prefabricated components is converted into text knowledge blocks. Based on the text knowledge block, a first prompt word is generated using a first preset structured indicator word template; Based on the first prompt word, a preset large model is used to generate semantic entities in the text knowledge block and the relationships between semantic entities; Based on the semantic entities in the text knowledge block and the relationships between them, construct the subgraph corresponding to the text knowledge block; The subgraphs corresponding to each of the text knowledge blocks are merged to generate the quality management knowledge graph.

8. The method according to claim 7, characterized in that, The method further includes: Entity clustering is performed on the quality management knowledge graph to obtain the knowledge communities corresponding to the quality management knowledge graph; Based on the knowledge community, a second prompt word is generated using a second preset structured indicator word template; Based on the second prompt word, a community report of the knowledge community is generated using a preset large model.

9. A hybrid knowledge retrieval enhancement generation device, characterized in that, The device includes: The acquisition module is used to retrieve query requests for quality management knowledge related to prefabricated components. The calculation module is used to calculate the dense similarity score of each node block in the quality management knowledge graph of the prefabricated component based on the semantic entities in the quality management knowledge query request and using a preset dense retrieval algorithm. The calculation module is also used to calculate the sparse similarity score of each node block in the quality management knowledge graph based on the terms in the quality management knowledge query request and using a preset sparse retrieval algorithm. The calculation module is also used to calculate the fusion similarity score of each node block based on the dense similarity score of each node and the sparse similarity score of each node block; The calculation module is also used to traverse the quality management knowledge graph based on the fusion similarity score of each node block to obtain the similarity score of each edge block in the quality management knowledge graph. The generation module is used to generate the retrieval answer corresponding to the quality management knowledge query request based on the fusion similarity score of each node block and the similarity score of each edge block.

10. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores program instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the program instructions to perform the steps of the hybrid knowledge retrieval enhancement generation method as described in any one of claims 1 to 8.