Federal knowledge retrieval and big language model enhancement system and method
Through the coordinated design of the server-side cache module and hardware accelerator, a dynamic cache system is built, which solves the query and fusion problems of distributed multi-source knowledge graphs, and realizes efficient and secure knowledge graph retrieval and large language model enhancement, improving query response speed and result accuracy.
Patent Information
- Application Number
- CN202510263857.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-29
AI Technical Summary
The existing RAG framework is difficult to efficiently handle the query and fusion of distributed multi-source knowledge graphs, and there are shortcomings in privacy protection and efficient computing requirements, especially in the process of large-scale graph data processing and query.
The server-side cache module is used to work in collaboration with the hardware accelerator to build a two-level dynamic cache system, protect privacy through differential privacy technology, and use the parallel computing power of the hardware accelerator to segment and parallel search to optimize the search process of the knowledge graph.
It significantly improves the query efficiency of multi-source knowledge graphs and the inference performance of large language models, reduces the system network load, ensures data privacy and security, and improves query response speed and result accuracy.
Smart Images

Figure CN120386853A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent reasoning, and also relates to the technical field of information retrieval, and particularly relates to a federated knowledge retrieval and large language model enhancement system and method. Background Art
[0002] As a structured information storage and expression method, a knowledge graph organizes factual knowledge in the form of triples (head entity, relation, tail entity), and can effectively capture the complex relationships between entities. This technology is widely applied to knowledge-intensive tasks, such as intelligent question answering, recommendation systems, and semantic search and other fields. In contrast, traditional question answering systems often face problems such as low query efficiency, lagging knowledge updates, and incomplete expressions when dealing with knowledge-rich fields. The knowledge graph, through the associations between nodes, provides in-depth information retrieval and reasoning support, and becomes the key technology to solve these problems.
[0003] In recent years, large language models (LLMs) have made remarkable progress in the field of natural language processing, and have demonstrated powerful capabilities in semantic understanding and generation. However, the training of these models requires a large amount of computing resources, and the pre-training process is time-consuming and costly, which makes them have certain limitations when dealing with rapidly changing knowledge fields. In addition, the content generated by LLMs may have the "hallucination" problem, that is, although the grammar is correct, the content may contain errors or inaccurate information.
[0004] To solve the above problems, the Retrieval-Augmented Generation (RAG) framework has been proposed and widely applied. This framework improves the quality of input information by integrating accurate information from external knowledge bases (such as knowledge graphs) into the model reasoning process, thereby effectively reducing the occurrence of the "hallucination" problem. When applying RAG in a question answering system, by integrating knowledge graph information into the generation process, it provides a reliable reasoning basis for the LLM, enabling it to more accurately respond to user questions, especially suitable for task scenarios that need to process a large amount of dynamically updated knowledge.
[0005] However, the current Retrieval-Augmented Generation (RAG) framework generally adopts a centralized knowledge graph architecture and can only obtain knowledge from a single data source. However, in actual application scenarios, knowledge graphs usually exist in a distributed form, especially in the federated learning mode, where knowledge graph shards in the same domain are scattered and stored on multiple local computers. This conflict between the distribution characteristics and the centralized architecture leads to systematic defects in existing solutions:
[0006] First, the multi-source retrieval efficiency is low: When traversing the cross-device knowledge graph, the traditional hardware architecture lacks the collaborative design of the shard loading module and the task decomposition unit, resulting in an imbalance between data loading and computing resources. Specifically, when performing large-scale graph queries, insufficient on-chip memory capacity causes memory overflow, or unreasonable sub-task granularity division causes computing power to be idle, significantly reducing the retrieval speed.
[0007] Second, the irreconcilable contradiction between privacy and efficiency: Existing privacy protection solutions rely on full data encryption or desensitization processing. Although they can prevent the leakage of sensitive information, the encryption and decryption operations increase the retrieval latency by 200%-300%. The absence of a hardware-level differential privacy processing unit makes it impossible for the system to achieve the goal of parallel entity-level privacy protection and millisecond-level response through Laplace noise injection.
[0008] Third, the mismatch between the static cache mechanism and dynamic requirements: The difference in access popularity of multi-source knowledge items invalidates the traditional static cache strategy. Without the joint optimization of the hot and cold data hierarchical storage unit and the prefetch unit, high-frequency knowledge items cannot be preferentially resident in the SRAM medium. For example, low-frequency data occupies more than 30% of the SRAM storage resources, forcing the system to frequently access the low-speed DRAM medium, increasing the query latency by 1.5-2 times.
[0009] For example, CN118364916A discloses a news retrieval method and system based on a large language model and a knowledge graph. The method includes: obtaining multi-source news data and constructing a news database according to the multi-source news data; using an initial large language model to perform knowledge extraction on the multi-source news data in the news database to obtain news entities and news relationships, and constructing a news knowledge graph according to the news entities and news relationships; using a knowledge fusion module to integrate the news knowledge graph into the initial large language model to obtain a final large language model; performing news content retrieval based on the final large language model to obtain news retrieval results. This technical solution only performs retrieval based on a single knowledge graph, so it also has the above-mentioned defects.
[0010] The present invention hopes to provide a large language model enhancement system and method with a new mechanism, combined with a differential privacy protection mechanism to improve the inference effect of the large language model (LLM) and the multi-source knowledge retrieval efficiency.
[0011] In addition, on the one hand, there are differences in the understanding of those skilled in the art; on the other hand, although the applicant has studied a large number of documents and patents when making the present invention, due to space limitations, all details and contents are not listed in detail. However, this does not mean that the present invention does not possess the features of these prior arts. On the contrary, the present invention already possesses all the features of the prior arts, and the applicant reserves the right to add relevant prior arts in the background art. Summary of the Invention
[0012] However, most current RAG frameworks rely on a single knowledge graph, limiting their ability to acquire knowledge from a single data source. In reality, knowledge graphs are often distributed. In scenarios like federated learning, knowledge graphs for the same domain may be scattered across multiple clients or devices. Consequently, existing RAG frameworks struggle to efficiently handle the query and fusion of distributed multi-source knowledge graphs, and they also lack consideration for privacy protection and efficient computation.
[0013] In response to the shortcomings of the existing technology, the present invention provides a federated knowledge retrieval and large language model enhancement system from a first aspect, comprising a server and a hardware accelerator. The server preprocesses the input query and extracts the central word of the query; and retrieves the central word during the caching process through its own cache module. If the cache hits, the cache module provides the query result corresponding to the central word to the data processor for inference through the large language model and generation of an enhanced answer; if the cache of the cache module does not hit, the server sends a retrieval instruction to the hardware accelerator to search the local cache of the hardware accelerator; if the local cache still does not hit, the hardware accelerator calls the local knowledge graph and generates parallel subtasks associated with the central word by splitting the local knowledge graph, thereby performing a deep search, and adding noise to the knowledge items retrieved by the parallel subtasks to protect sensitive information. The hardware accelerator integrates the knowledge items with noise and sends them to the data processor through the cache module for inference through the large language model and generation of an enhanced answer.
[0014] Through the collaborative work of the server-side cache module and the hardware accelerator, the present invention constructs a two-level dynamic caching system. This architecture effectively reduces the frequency of cross-device communication, prioritizes local high-frequency cache responses to query requests, and reduces the overall network load of the system. When a cache miss occurs, the hardware accelerator implements task parallelization through a knowledge graph segmentation algorithm, fully leveraging the parallel processing capabilities of dedicated computing hardware to significantly improve task throughput in complex query scenarios.
[0015] Furthermore, within the federated learning framework, this paper employs a dynamic noise injection mechanism to maintain the privacy of search results. This approach, while maintaining knowledge availability, protects sensitive information through differential privacy techniques, ensuring that the knowledge graph processing process complies with international data security standards. The knowledge graph segmentation strategy further limits data processing to local nodes, preventing the risk of raw data being transmitted externally. The distributed management mechanism of the local knowledge graph effectively reduces the load on central nodes and, in conjunction with the task scheduling algorithm, enables dynamic resource allocation, adapting to the needs of application scenarios of varying scales.
[0016] According to a preferred embodiment, the hardware accelerator includes a data loader, a computing module, a differential privacy processing unit, and a result aggregation unit. The data loader is used to batch obtain data from the local knowledge graph and divide the knowledge graph into multiple data blocks suitable for the size of the on-chip memory on the hardware accelerator; the computing module is used for parallel computing, and the differential privacy processing unit is used to add noise to sensitive data during the processing to ensure that data privacy is not leaked; the result aggregation unit is used to summarize the output information of each computing unit and perform comprehensive analysis and processing to ensure the integrity and consistency of the final result.
[0017] The data loader adopts a block loading strategy based on on-chip memory, realizes efficient reuse of memory resources by dynamically dividing the data scale, and can significantly reduce the I / O bottleneck and memory access latency during the processing of large-scale knowledge graphs. The parallel architecture of the computing module optimizes the thread scheduling strategy for the irregular access characteristics of graph-structured data, and can effectively improve the throughput in typical tasks such as entity relationship reasoning and graph embedding calculation, and is more adaptable to the high-dimensional feature calculation requirements of knowledge graphs than the traditional serial processing method.
[0018] Moreover, the differential privacy processing unit realizes the time and space optimization of the noise injection mechanism at the hardware level, adopts a dynamically adjustable privacy budget allocation strategy, and ensures minimizing the distortion impact of noise on the analysis result under the premise of satisfying the (ε,δ)-differential privacy constraint. The pipeline design of the present invention deeply integrates the privacy protection process with the computing process, avoiding the privacy leakage risk and additional performance loss in the traditional software-level implementation. The result aggregation unit synchronously completes data consistency verification when integrating the parallel computing results through a distributed redundant calculation verification and adaptive weight allocation algorithm, ensuring that the statistical characteristics of the final output are consistent with the theoretical expectation.
[0019] According to a preferred embodiment, the data loader further includes a task decomposition unit. During the data loading process, the task decomposition unit searches for the central word of the query in all the divided sub-regions of the knowledge graph, which is beneficial to the rapid progress of the task.
[0020] Specifically, the task decomposition unit of the present invention adopts a distributed retrieval mechanism to map the central word of the query task to the pre-divided sub-regions of the knowledge graph for parallel search. This architecture design has three core advantages: (1) converting the global search into multiple local searches through a space segmentation strategy, effectively reducing the computational complexity of a single search; (2) using the local association characteristics of the knowledge graph and adopting an optimized graph traversal algorithm within the sub-region to reduce the number of accesses to invalid nodes; (3) supporting a modular architecture with dynamic expansion. When the scale of the knowledge graph expands, the retrieval efficiency can be maintained by increasing the number of sub-regions, avoiding the performance degradation problem faced by traditional full-graph searches.
[0021] According to a preferred embodiment, the hardware accelerator further includes an on-chip controller, and the on-chip controller includes: a task distribution unit that receives and distributes subtasks from the task queue to ensure workload balance among the various computing units in the computing module; an instruction decoder that converts high-level task instructions into specific hardware operation instructions to guide the computing units to execute corresponding tasks; and a status management unit that monitors the running status of each module in real time and coordinates resource allocation to ensure the correctness of task execution and the stability of the system.
[0022] The task distribution unit adopts a dynamic load balancing algorithm and performs real-time scheduling according to the working status of the computing units and task characteristics, so as to improve the utilization rate of computing resources; the instruction decoder converts abstract operation instructions into specific micro-operation sequences by establishing a configurable instruction mapping table, which not only maintains the usability of the upper-layer programming interface but also realizes the efficient drive of the underlying hardware; the status management unit constructs a multi-dimensional monitoring system, collects and analyzes more than 20 parameters such as the power consumption, temperature, and instruction throughput of the computing cores in real time, and realizes the adaptive adjustment of abnormal states in combination with the preset resource configuration strategy. The coordinated work of these three-layer control mechanisms enables the system to maintain a peak computing power of more than 95% while controlling the task execution error rate below the order of magnitude of 10^-6.
[0023] According to a preferred embodiment, the hardware accelerator further includes a prefetch unit. The prefetch unit predicts potential data requirements using the historical access pattern and performs data prefetch on the knowledge graph nodes based on the Markov chain model, so as to load data in advance to reduce the query latency. Through the Markov chain model, the prefetch unit can load the path data that may be accessed from the data loader into the global memory in the computing module of the hardware accelerator before the query task officially starts. This preloading mechanism reduces the performance bottleneck caused by data access latency during the query process, thereby improving the query efficiency.
[0024] According to a preferred embodiment, the prefetch unit of the hardware accelerator predicts the accessible nodes during the query process based on the Markov chain model; among them, before the query starts, a Markov chain model is established based on the historical access pattern; and the node access path is predicted based on the Markov chain model. During the query process, this prediction enables the system to make data preparations in advance, so that the query task can obtain data immediately when needed, thus significantly improving the system response speed.
[0025] According to a preferred embodiment, the server includes a data processor and a cache module. The data processor extracts the central word and runs a primary search program, and sends the extracted central word to the cache module; the cache module searches for the existence of a corresponding knowledge item; in the case where the knowledge item of the central word exists, the cache module directly returns the query result to enter the enhanced answering stage, and the query result is composed of the neighbors of the central word. The data processor texturizes the query result and inputs it into a large language model for reasoning to generate an enhanced answer.
[0026] The primary search program of the present invention completes the extraction and preliminary retrieval of the central word in the data processor, avoiding the high computational overhead caused by directly invoking the large language model. The cache module adopts a knowledge item pre-storage mechanism, and when a historical query pattern is detected, it can directly reuse the structured knowledge graph data, reducing the redundant calculation by about 30-50% compared with the traditional full-process processing method. At the same time, the enhanced answering stage only starts the large language model for valid queries with cache hits, and through the dynamic resource allocation mechanism, the utilization rate of the computing resources of the high-precision model is increased by more than 20%. This hierarchical processing strategy significantly reduces the computing power consumption per query while ensuring the accuracy of the results.
[0027] Moreover, the query result returned by the cache module is constructed based on the neighbor node network of the knowledge graph. This design brings two technical advantages: on the one hand, using the topological connection characteristics of the graph data to automatically generate an associated entity set can improve the result relevance by more than 15% compared with the traditional keyword matching method; on the other hand, the data processor converts the graph structure data into a text sequence with semantic relationships, and by retaining the edge attribute information between nodes (such as relationship type, connection weight), the semantic density of the input to the large language model is increased by about 40%, significantly improving the logical coherence and factual accuracy of the generated answer.
[0028] According to a preferred embodiment, the processing steps of the data processor in the server further include: in the case where no relevant knowledge item is found in the primary search, sending a deep search request to each local computer through the Internet; the local computer receives the query request of the server through the network adapter; the local computer runs a deep search program based on the query request, and the deep search program accesses the hardware accelerator chip through the high-speed bus, which helps to improve the retrieval efficiency of knowledge.
[0029] When the cache module misses, the present invention triggers a deep search request and its related protocol: the data processor encapsulates the query request into a standard data packet and distributes it to local nodes through a load balancing algorithm. This process reuses the scheduling strategy of the task distribution unit, which can optimize the network transmission efficiency; the local computer activates the task queue management function of the on-chip controller of the hardware accelerator, and compiles the deep search program into hardware micro-instructions through an instruction decoder; the high-speed bus is directly connected to the hardware accelerator architecture, combined with the execution of parallel subtasks of the task distribution unit, to improve the execution efficiency of the graph traversal algorithm.
[0030] The present invention provides a federated knowledge retrieval and large language model enhancement method from a second aspect. The method includes: the server preprocesses the input query to extract the central word of the query; and retrieves the central word during the caching process through its own cache module. If the cache hits, the cache module uses the large language model to reason about the query result corresponding to the central word and generates an enhanced answer; if the cache of the cache module misses, the server sends a retrieval instruction to the hardware accelerator to retrieve in the local cache of the hardware accelerator; if the local cache still misses, the hardware accelerator calls the local knowledge graph and generates parallel subtasks associated with the central word in a way of splitting the local knowledge graph, so as to perform a deep search, and adds noise to the knowledge items retrieved by the parallel subtasks to protect sensitive information. The hardware accelerator integrates the knowledge items with added noise and uses the large language model to reason and generate an enhanced answer.
[0031] The advantages brought by the method of the present invention include: when the preprocessing accurately identifies the core semantic unit, the memory cache hit rate can reach the benchmark level of 85%, enabling 68.5% of the query requests to complete the closed-loop processing at the server level, avoiding triggering the resource consumption of the downstream hardware accelerator. This feedforward optimization mechanism reduces the overall system resource overhead by about 32% while ensuring sub-second response for high-frequency requests.
[0032] When the request is sent to the hardware accelerator, the knowledge graph dynamic segmentation algorithm realizes the optimization of the computation-communication ratio through task granularity control: fine-grained segmentation adapts to the parallel computing units of the hardware accelerator, while the task aggregation strategy based on semantic communities effectively suppresses the communication overhead caused by data sharding. During this process, the differential privacy module adopts a streaming noise injection mechanism to complete noise superposition at the parallel subtask generation stage, avoiding the 30-40% additional delay caused by traditional post-processing methods. This proactive privacy protection design aligns the data processing flow and the security mechanism in space and time, while maintaining a semantic fidelity of 86.7%, the system throughput is increased by 22.5%.
[0033] According to a preferred embodiment, the local cache is provided with a dual storage mechanism for hot and cold data to improve the retrieval efficiency; wherein, the cold data and the hot data are stored on different types of storage media; the access times or timestamps of the data items are monitored, and when it is detected that the access frequency of a certain data exceeds a preset threshold, the migration process is started; the low-frequency data is removed from the hot data cache area and stored in the slower cold data area; at the same time, the frequently accessed data is migrated from the cold data area to the hot data cache to ensure its fast access.
[0034] The method of the present invention adopts a dual storage mechanism for hot and cold data, and the data with high access frequency stays in the hot data cache area preferentially to ensure the optimal access speed. The cold data storage is responsible for storing the data with low access frequency. When the access frequency of the data changes, the system will regularly adjust the position of the data so that the high-frequency data can be quickly migrated to the hot data cache, thereby reducing the access latency and improving the query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a hardware architecture diagram of the federated knowledge retrieval and large language model enhancement system provided by the present invention;
[0036] Figure 2 It is a schematic diagram of the hardware structure of the server provided by the present invention;
[0037] Figure 3 It is a hardware structure diagram of each local computer provided by the present invention;
[0038] Figure 4 It is a logical schematic diagram of the hardware accelerator module provided by the present invention;
[0039] Figure 5 It is an overall flowchart of the federated knowledge retrieval and large language model enhancement system provided by the present invention;
[0040] Figure 6 It is a schematic diagram of the process flow of the in-depth search process of the present system provided by the present invention;
[0041] Figure 7 It is a flowchart of the differential privacy processing process in the in-depth search stage provided by the present invention.
[0042] LIST OF REFERENCE NUMERALS
[0043] 100: Server; 110: Data Processor; 120: Cache Module; 200: Hardware Accelerator; 210: I / O Adapter; 220: Network Adapter; 230: RAM; 240: Client; 250: External Memory; 260: Accelerator Chip; 270: Data Loader; 271: DMA Controller; 272: Task Decomposition Unit; 273: On-chip Memory; 280: Prefetch Unit; 281: Node Transition Matrix; 282: Historical Access Record; 290: Computing Module; 300: Differential Privacy Processing Unit; 310: High-speed Bus; 311: On-chip Controller; 312: Task Distribution Unit; 313: Instruction Decoder; 314: Status Management Unit; 320: Hot and Cold Data Hierarchical Storage Module; 321: Hot Data Storage Medium; 322: Cold Data Storage Medium; 330: Result Aggregation Unit; 340: On-chip Bus; 400: Local Computer; 410: Local Knowledge Graph. Detailed Implementation Manner
[0044] The following is a detailed description in conjunction with the accompanying drawings.
[0045] The present invention explains some noun terms.
[0046] Federated Knowledge Retrieval: An application branch of federated learning that allows multiple distributed data sources to collaborate on knowledge discovery and model training without sharing actual data, improving model performance while ensuring data privacy and security.
[0047] Query: A form of request used to retrieve specific information from a database or other information storage system. A query usually contains keywords, phrases, or more complex conditional expressions entered by the user.
[0048] Prompt: The input text provided to an artificial intelligence model, intended to guide the model to generate or process content according to the user's intention.
[0049] Large Language Model (LLM): A large-scale neural network based on deep learning that can understand and generate natural language text. The LLM is trained with a large amount of text data and can exhibit language capabilities close to those of humans in various tasks.
[0050] Cache (Server-side Cache) Module 120: Belonging to the first-level cache, it is a temporary data storage area stored on the server 100, used to save copies of frequently accessed data to reduce repeated calculations and database queries, thereby accelerating the response speed.
[0051] Local Cache: It belongs to the secondary cache and is the cache on the client side. It is used to store recent or frequently used data to reduce the dependence on server 100 and improve the response speed and user experience of the application.
[0052] Graph Processing Module (GPM): A hardware or processing module designed specifically for efficient processing of graphic data, which optimizes the storage and retrieval processes of graphic structures.
[0053] Knowledge Graph: A structured form of knowledge representation that uses nodes to represent entities and edges to represent relationships between entities, forming a complex information network that supports advanced information retrieval and reasoning functions.
[0054] Retrieval-Augmented Generation (RAG): A method that combines traditional retrieval techniques with modern generation models. It enriches the generated content by retrieving relevant documents to improve the relevance and accuracy of the generated text.
[0055] Differential Privacy: A technology for protecting individual data privacy that adds noise during the data analysis process to ensure that the presence or absence of a single record does not significantly affect the analysis results.
[0056] Cosine Similarity: A measure for evaluating the consistency of the directions of two vectors, commonly used in fields such as text analysis. It assesses similarity based on the cosine value of the angle between vectors.
[0057] Hardware Accelerator 200: Hardware components such as GPUs or TPUs that are used to accelerate specific types of computing tasks, such as matrix operations in machine learning or convolution operations in image processing.
[0058] Data Loader 270: A processing module used to read data from a storage medium and transfer it to the memory of a computing device, which is part of the data preprocessing process.
[0059] Task Dispatch Unit 312: A mechanism for managing and allocating tasks to different computing units, aiming to achieve parallel processing and load balancing to improve overall efficiency.
[0060] Task Queue: An ordered list of tasks used to schedule the execution order to ensure that tasks with higher priorities are executed first.
[0061] Computation Module 290: A hardware or processing module that performs specific computational tasks, such as the arithmetic logic unit (ALU) executed by the CPU or the computational core on the GPU.
[0062] Result Aggregation: The process of summarizing the results from multiple computing nodes to generate a final output.
[0063] BERT (Bidirectional Encoder Representations from Transformers): A pre-trained language model based on the Transformer architecture that captures context information through bidirectional encoders and performs well in natural language processing tasks.
[0064] Cold and Hot Data Dual Storage Mechanism: A data management strategy that stores frequently used "hot data" in fast storage media (such as caches), while storing infrequently accessed "cold data" in slower but more economical media to optimize storage resources and access speed.
[0065] SRAM (Static Random-Access Memory): A high-speed semiconductor memory that uses static storage cells to save data without periodic refreshing, and is commonly used for caching and fast storage requirements.
[0066] DRAM (Dynamic Random-Access Memory): A commonly used type of main memory where data is stored in capacitors and needs to be refreshed periodically to maintain the data. Due to its high capacity and low cost, it is widely used in computer systems.
[0067] Markov Chain Prediction Algorithm: A prediction method based on a probability transition matrix that uses the current state of the system to predict the next state, and is commonly used in the modeling and analysis of time series data and stochastic processes.
[0068] The present invention provides a federated knowledge retrieval and large language model enhancement system and method, and can also provide a federated knowledge retrieval and large language model enhancement device, and can also provide a processor for the federated knowledge retrieval and large language model enhancement system.
[0069] Existing RAG frameworks mainly rely on a single knowledge graph, usually limited to a single data source for knowledge retrieval. However, in reality, knowledge graphs are often distributed, especially in scenarios such as federated learning, where knowledge graphs from the same domain are often distributed across multiple clients or devices. Therefore, existing RAG frameworks are difficult to efficiently handle the query and fusion problems of distributed multi-source knowledge graphs, and do not fully consider privacy protection and efficient computing requirements during knowledge graph retrieval, especially facing performance bottlenecks during the processing and query of large-scale graph data.
[0070] Therefore, how to achieve efficient retrieval and fusion of multi-source knowledge graphs while ensuring privacy security is an important challenge in current knowledge-enhanced generation technologies.
[0071] The present invention proposes a software-hardware collaborative federated knowledge retrieval and large language model enhancement system, and its main idea is as follows:
[0072] Through software-hardware collaborative design, combining the data processor 110, cache module 120, and hardware accelerator 200 in the server 100, the efficiency of federated knowledge graph retrieval and the inference performance of the large language model (LLM) can be significantly improved, and differential privacy technology is used to ensure data privacy. Specifically, the system of the present invention adopts a multi-level cache mechanism (including server-side cache and local cache) to cooperate with the hardware accelerator 200 to optimize the retrieval process of the knowledge graph, accelerate the query response speed, and implement privacy protection measures during the entire retrieval process.
[0073] The basic operation process of the system of the present invention is as follows: First, the user's query request is sent to the cache module 120 on the server side for preliminary inspection to determine whether the required information can be directly obtained from the cache module 120, so as to avoid unnecessary repeated calculations. If further search is required, the request will be passed to the hardware accelerator 200. The hardware accelerator 200 can quickly execute complex graph structure data processing tasks. After the hardware accelerator 200 completes the in-depth search, the results are returned to the cache module 120. The cache module 120 aggregates and organizes the knowledge fragments from different sources. Finally, these summarized knowledge is finally processed by the data processor 110 on the server side and provided as enhanced input to the LLM to generate more accurate answers.
[0074] This design scheme not only improves the retrieval efficiency and answer quality, but also strengthens the security of user data through technical means such as differential privacy, forming a solution that is both efficient and secure, and is particularly suitable for application in knowledge graph scenarios involving multiple distributed data sources. In this way, whether it is the need for real-time information or the analysis of large-scale data sets, the system can provide strong support.
[0075] Embodiment 1
[0076] The federated knowledge retrieval and large language model enhancement system of the present invention includes a server 100. At least a data processor 110, a cache module 120, and a hardware accelerator 200 are provided inside the server 100. The client 240 is connected to the server 100 through an I / O adapter 210. The I / O adapter 210 is connected to the server 100 through a high-speed bus 310 and is used to process external input and output operations. The network adapter 220 is connected to the server 100 through the high-speed bus 310 and is responsible for communicating with the Internet or other computers. In the present invention, it specifically refers to communicating with multiple local computers 400 through the Internet.
[0077] Preferably, the central processing unit (CPU) of the server 100 is also connected to the main memory RAM 230 through the high-speed bus 310. Key modules such as a data processing program and a primary search program are stored in the RAM 230, that is, the data processing program on which the data processor 110 depends.
[0078] In addition, the cache module 120 is connected to the high-speed bus 310 and is used to store recently used or frequently accessed data to reduce the access frequency to the RAM 230, thereby improving the data retrieval speed.
[0079] Preferably, as Figure 3 shown, the hardware accelerator 200 includes an accelerator chip 260, a CPU, a RAM, a network adapter 220, and an external memory 250, etc. The accelerator chip 260 includes multiple functional modules for accelerating data processing and computing tasks. The hardware accelerator 200 is deployed at each local end. The local end also refers to the client 240 and the client side.
[0080] As Figure 3 shown, the network adapter 220 is responsible for obtaining data from the external network and the local knowledge graph 410 and transmitting the data to the CPU and the RAM 230 through the high-speed bus 310. Modules such as a deep search program are stored in the RAM 230. The CPU is connected to the RAM 230 and is used to execute program instructions and process the data stored in the RAM 230.
[0081] A chip-internal controller 311, a data loader 270, a task distribution unit 312, a computing module 290, a differential privacy processing unit 300, a result aggregation unit 330, etc. are integrated in the accelerator chip 260. As Figure 3 shown, the physical composition of the hardware accelerator 200 consists of multiple key components, which are responsible for storing local personalized knowledge and performing efficient graph retrieval while ensuring privacy.
[0082] As Figure 3As shown in the figure, the accelerator chip 260 of the hardware accelerator in the present invention adopts a multi-layer on-chip bus architecture to realize the interconnection communication and data collaboration among various functional modules. Figure 3 The on-chip bus 340 in it is a highly optimized on-chip system interconnection structure.
[0083] The on-chip bus 340 adopts a hierarchical topology structure and is divided into two major physical channels: a high-performance computing bus and a general control bus. Among them, the high-performance computing bus is configured as a 256-bit wide bidirectional transmission channel, and its operating frequency is synchronized with the main frequency of the computing module. It is dedicated to data-intensive transmission between the computing module, the data loader, and the global memory; the general control bus is configured as a 64-bit wide time-division multiplexing bus, which carries the control signal and status information interaction among various modules.
[0084] The on-chip bus architecture adopts a master-slave access mechanism. Among them, the on-chip controller 311, as the bus master device, has the highest arbitration priority. The task decomposition unit 272 sends task configuration parameters and instruction streams to the task distribution unit 312 through the general control bus to implement dynamic bandwidth allocation. The data loader 270, the computing module 290, the differential privacy processing unit 300, and the result aggregation unit 330 are mounted on the on-chip bus 340 as slave devices through the on-chip bus interface unit.
[0085] The connection relationship of the on-chip bus 340 is specifically as follows:
[0086] The task decomposition unit 272 sends task configuration parameters and instruction streams to the task distribution unit 312 through the general control bus.
[0087] The DMA controller 271 of the data loader 270 is directly connected to the high-performance computing bus to realize the burst transfer mode data path with the global memory of the computing module 290.
[0088] The computing module 290 includes a computing unit, a local memory, and also includes a global memory. Each graphics processing computing unit (GPM) in the computing module 290 accesses through the interface unit of the on-chip bus 340 in a dual-channel manner. Among them, the computing instruction channel is connected to the general control bus, and the data channel is directly connected to the high-performance computing bus.
[0089] The prefetch unit 280 captures the access feature data of the node transition matrix 281 through the on-chip bus 340 listening mechanism and injects the prefetch data during the idle period of the on-chip bus 340.
[0090] The differential privacy processing unit 300 is embedded in a bridging manner on the on-chip bus 340 to implement real-time noise injection on the bus transmission path between the computing module 290 and the result aggregation unit 330.
[0091] The result aggregation unit 330 is configured as a master-slave dual-mode device of the on-chip bus 340, which not only receives the output data stream of the computing module 290, but also feeds back the system status to the on-chip controller 311 through the on-chip bus 340.
[0092] Through the combination of physical layer timing optimization and protocol layer intelligent scheduling, this on-chip bus architecture effectively solves the bus contention problem during multi-module concurrent access, providing a reliable interconnection infrastructure for the coordinated operation of each functional module of the hardware accelerator.
[0093] The data loader 270 is connected to the external memory 250 through the on-chip bus 340 and the high-speed bus 310, and is used to obtain data in batches from the local knowledge graph 410. The data loader 270 is connected to the DMA controller 271, which is used to efficiently transfer data into the hardware accelerator 200, reducing the waiting time. The DMA controller 271 is connected to the knowledge graph in the off-chip memory and the on-chip memory 273 of the hardware accelerator 200, and is used to directly and efficiently transfer data without passing through the central processing unit CPU. The on-chip controller 311 is respectively connected to the task distribution unit 312, the instruction decoder 313, and the status management unit 314, and is used to coordinate and manage the allocation and execution of computing tasks.
[0094] Preferably, in order to support efficient data storage and access, the hardware accelerator 200 integrates a cache structure - local cache, which is built using static random access memory (SRAM) to ensure fast access to data at different computing stages.
[0095] Preferably, the DMA controller 271 in the data loader 270 is used to obtain data in batches from the local knowledge graph 410 and quickly load the required information into the hardware accelerator 200 through an optimized data transfer channel, thereby reducing the waiting time. The DMA controller 271 can directly and efficiently transfer data between the knowledge graph in the external memory 250 and the on-chip memory 273 of the hardware accelerator 200 without passing through the central processing unit at the local end.
[0096] Specifically, the data loader 270 first divides the knowledge graph in the external memory 250 into multiple data blocks suitable for the size of the on-chip memory 273 of the hardware accelerator 200. Each data block contains a part of the graph structure, such as nodes and edges. In addition, the data loader 270 includes a task decomposition unit 272. During the data loading process, the task decomposition unit 272 processes the central word of each query to ensure that search operations are performed in all segmented sub-regions of the knowledge graph.
[0097] Preferably, the hardware components of the on-chip controller 311 include a task distribution unit 312, an instruction decoder 313, and a status management unit 314. The task distribution unit 312 is responsible for receiving and distributing subtasks from the task queue to ensure workload balance among the computing units in the on-chip computing module 290. The instruction decoder 313 converts high-level task instructions into specific hardware operation instructions to guide each module to execute corresponding tasks. The status management unit 314 monitors the running status of each module in real time and coordinates resource allocation to ensure the correctness of task execution and the stability of the system.
[0098] Figure 3 The task distribution unit 312 in [description] is responsible for receiving and parsing requests from the task queue, dividing tasks into subtasks that can be processed in parallel, and then distributing them to each computing unit in the computing module 290, thereby improving the overall processing efficiency. In this way, the task distribution unit 312 can dynamically allocate computing tasks to different computing units to achieve parallel computing and load balance.
[0099] The computing module 290 consists of multiple customized graph processing computing units (Graph Processing Module, GPM), simply referred to as computing units. These computing units provide highly parallel computing capabilities and are particularly good at processing complex graph data. Each computing unit is not only equipped with local memory internally but also can interact with the global memory in the computing module 290. This architecture provides the basic conditions for the prefetch unit 280 to be introduced later.
[0100] The differential privacy processing unit 300 ensures that data privacy is not leaked by adding noise to sensitive data during the processing. The differential privacy processing unit 300 is embedded in the computing module 290 in the form of hardware logic and executes the privacy protection algorithm in real time when each query subtask is completed. The differential privacy processing unit 300 is connected to the result aggregation unit 330.
[0101] Finally, the result aggregation unit 330 aggregates the outputs of each computing unit and performs comprehensive analysis and processing to ensure the integrity and consistency of the final result. The result aggregation unit 330 is implemented using a high-performance processor to ensure fast response under high load conditions.
[0102] As Figure 1 shown, the system structure and operation process of the hardware accelerator 200 according to the present invention show an exemplary embodiment. In Figure 1 and Figure 2Among them, the system includes a server 100 and multiple local computers 400. The server 100 can be multiple computers, and can also be called a server-side computer. These components work together through network connections to support efficient knowledge retrieval and reasoning functions.
[0103] The server 100 is configured with a data processor 110 and a cache module 120. The user inputs n query requests on the server-side computer. These query requests are used by the data processor 110 to extract central words. Then, the server-side computer runs a primary search program, and the central words extracted by the data processor 110 are transmitted to the cache module 120 on the server-side to search for corresponding knowledge items. If the knowledge item of the central word exists, the cache module 120 directly returns the query result to enter the enhanced answer stage. The query result consists of the neighbors of the central word. The data processor 110 texturizes the query result and inputs it into a large language model (LLM) for reasoning to generate an enhanced answer.
[0104] If the primary search does not find relevant knowledge items, the server 100 sends a deep search request to each local computer 400 through the Internet network. Each local computer 400 is equipped with a network adapter 220 for receiving the query request from the server 100. The query request is used by the local computer 400 to run a deep search program. The deep search program accesses the accelerator chip 260 of the hardware accelerator 200 through a high-speed bus 310.
[0105] Figure 3 Shows the internal structure of the accelerator chip 260 of the hardware accelerator 200. The accelerator chip 260 is built-in with an on-chip controller 311, which is responsible for decoding instructions and controlling the data flow.
[0106] First, the local computer 400 runs a deep search program to search for the knowledge item of the central word cached locally. If found, the result is immediately returned to the server-side computer; otherwise, the data loader 270 in the accelerator chip 260 first divides the knowledge graph in the external memory 250 into multiple data blocks suitable for the size of the on-chip memory 273 of the hardware accelerator 200. Each data block contains a part of the graph structure, such as nodes and edges. This block division method ensures that each data block can be efficiently stored and processed in the on-chip memory 273. During the data loading process, the data loader 270 also decomposes the query task into multiple subtasks according to the central word. The central word of each query task is searched in all the divided knowledge graph sub-regions. Each search corresponds to a subtask. Each subtask retains its own central word, and the central word is distributed to the task queue and queued according to the priority for processing.
[0107] Tasks in the task queue are transmitted to the computing module 290 through the task distribution unit 312. The computing module 290 consists of multiple computing units and performs graph query operations on subtasks in parallel. The calculation results of each subtask are processed by the differential privacy processing unit 300 on the computing module 290. The differential privacy processing unit 300 adds noise to protect sensitive information and ensure that user privacy is not leaked. Finally, the processing results are transmitted by the differential privacy processing unit 300 to the result aggregation unit 330. The result aggregation unit 330 is responsible for integrating the results of subtasks, removing duplicates, and ensuring the integrity of the results.
[0108] After the hardware accelerator 200 of each local computer 400 completes the depth search, the query results are summarized to the cache module 120. The cache module 120 merges and organizes the retrieval results from different sources to ensure data consistency and integrity. Finally, the cache module 120 sends the merged knowledge data to the data processor 110 to enhance the answering ability of the LLM.
[0109] Embodiment 2
[0110] This embodiment is a further elaboration of Embodiment 1, and repeated content will not be described again.
[0111] Preferably, for the federated knowledge retrieval and large language model enhancement method of the present invention, the flow of its steps is as Figure 1 , Figure 4 , Figure 5 , Figure 6 and Figure 7 shown, including five stages: query preprocessing, primary search, depth search, knowledge aggregation, and enhanced answering. Figure 6 and Figure 7 belong to Figure 5 the detailed steps in the
[0112] As Figure 5 shown, the overall steps of the federated knowledge retrieval and large language model enhancement method of the present invention include:
[0113] S001: Batch process user queries, extract central words, and perform a primary search.
[0114] S002: Check the cache on the server side.
[0115] S003: Determine whether there is relevant knowledge. If not, call the hardware accelerator 200 to perform a depth search.
[0116] S004: If so, obtain relevant knowledge items and get a prompt word prompt from the data processor 110.
[0117] S005: The LLM infers to get an answer.
[0118] S006: Determine whether there are any remaining questions. If so, restart step S001; if not, end.
[0119] S007: Store the knowledge aggregation results in the cache on the server side.
[0120] The detailed steps of the five stages of the present invention are described below.
[0121] S100: Query preprocessing stage: preprocess the input query and extract the core words of the query.
[0122] like Figure 1 As shown, the data processor 110 of the server 100 processes n queries accumulated by the user within a period of time. Specifically for each query (ie query request), the data processor 110 parses the query q currently input by the user. i , i∈[1,n], identify the core words with the most information, namely the central word c q This process includes steps such as word segmentation of the query text, removal of stop words, and identification of keywords.
[0123] For example, the query q i The question is “Which areas of China do giant pandas mainly live in?”
[0124] First, the data processor 110 performs word segmentation on the query based on the open source tool NLTK. NLTK is a natural language processing toolkit that can effectively decompose queries into words or phrases. The query is decomposed into "giant pandas | mainly | live | in | China | which | regions". Then, the data processor 110 uses the stop word list provided by NLTK to remove non-informative words such as "mainly", "in", "of", and "region". Finally, BERT is used to perform keyword recognition on the entire query and extract the most informative words "giant panda" and "China" as the central word c q .
[0125] For another example, when a user enters the query "What other animals live in the hummingbird's habitat?", first, the data processor 110 performs word segmentation on the query based on the open source tool NLTK. NLTK is a natural language processing toolkit that can effectively decompose queries into words or phrases, decomposing the query into "hummingbird|of|habitat|also|lives|which|animals". Next, the data processor 110 uses NLTK's built-in stop word list to remove non-informative words such as "of", "also", "which", and "animals". Finally, BERT is used to perform keyword recognition on the entire sentence, extracting the most informative words "hummingbird" and "habitat" as the central words c q。BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model that understands the semantics of each word through bidirectional context. In keyword recognition, BERT can capture the importance score of a word in a specific context by generating context-sensitive word vectors and output it, and select several (configurable) words with the highest importance scores as the central words c q 。
[0126] In keyword recognition, BERT calculates the importance score of each word by generating context-sensitive word vectors. Word vectors are vector representations that map words or phrases to a high-dimensional space to capture their semantic and contextual relationships.
[0127] BERT uses the attention mechanism to calculate the attention weights of each word with other words.
[0128] The calculation formula is:
[0129]
[0130] Among them, Q i is the query vector of word i, K j is the key vector of word j, and α ij represents the attention weight of word i to word j .
[0131] BERT obtains the final importance score of each word by calculating the attention score of each word and the weighted average of the relevant word vectors. Finally, BERT outputs several words with the highest importance scores as the central words. Words with higher scores indicate greater semantic contributions in the query.
[0132] The word vectors and attention mechanism of BERT can accurately extract keywords according to the context of the query, avoiding the limitations of traditional methods. By dynamically adjusting the representation of words, BERT can capture semantic differences and accurately identify the most informative words in the query, thereby improving the accuracy of the query.
[0133] By extracting the central words, the key points of the query can be highlighted, providing a more explicit direction for subsequent knowledge retrieval and reasoning. The central word c q will be sent by the data processor 110 to the cache module 120 for primary search to support subsequent knowledge retrieval and reasoning.
[0134] S200: Primary search stage: Retrieve the central words during the caching process. If the cache is hit, directly provide them to the data processor 110; if the cache is not hit, retrieve them from the local cache.
[0135] Such asFigure 1 As described above, in the primary search phase, in response to the query requirement of the data processor 110, the cache module 120 searches for the central word c q to check if it has been stored. The data in the cache module 120 is organized in the JSON format as shown in the cache module 120 in Figure 1 , including entities, neighbors, the queries processed by the system when storing in the cache, and the time of the last access to this knowledge item. The cache module 120 matches the entity name of the central word c q against all entries cached in itself to query the corresponding knowledge item. If a knowledge item K L1 related to the queried central word is found in the cache module 120, then this knowledge item is directly returned as the query result and enters the enhanced answer phase. If no valid result is queried in the cache module 120, the data processor 110 turns to the local cache (L2Cache) in each local hardware accelerator 200 for in-depth search.
[0136] After receiving the in-depth search request from the cache module 120, the hardware accelerator 200 enters the in-depth search phase.
[0137] S300: In-depth search phase: If still not hit, the hardware accelerator 200 calls the local knowledge graph 410 for in-depth search, and adds noise to each retrieved candidate knowledge item through the differential privacy processing unit 300 to protect sensitive information, and stores the result of the in-depth search after adding noise in the cache module 120.
[0138] The steps of the in-depth search phase are as shown in Figure 6 .
[0139] S310: Check the local cache to reduce the need to access the local knowledge graph 410.
[0140] This is determined by the difference in access speed between accessing the local cache and accessing the local knowledge graph 410. The query speed of accessing the local cache is much faster than that of accessing the local knowledge graph 410. The local cache in the hardware accelerator 200 is used to store local personalized knowledge items. The personalized knowledge items are the knowledge items with personalized characteristics stored at each local end due to the distributed storage of knowledge sources and the existence of privacy protection measures.
[0141] The hardware accelerator 200 receives the central word c q from the cache module 120 and queries the local cache data. Specifically, the hardware accelerator 200 uses the hot and cold data hierarchical storage module 320 to query the corresponding knowledge item in the local cache according to the entity name of the central word c q .
[0142] S320: Determine whether there is relevant knowledge.
[0143] S321: If relevant knowledge K is found in the local cache L2 , the hardware accelerator 200 obtains the knowledge items related to the central word of the query, and returns the query result to enter the knowledge aggregation stage.
[0144] S322: Otherwise, call the hardware accelerator 200 to access the local knowledge graph 410.
[0145] When the query result cannot be hit in the local cache, the hardware accelerator 200 accesses the local knowledge graph 410. The local knowledge graph 410 contains the neighbor relationships and semantic information between entities, and is the initial source of the knowledge required to enhance the answers of the LLM. The hardware accelerator 200 reduces the calculation time by parallelly calculating multiple graph query tasks.
[0146] S330: Obtain all valid neighbors of the central word with differential privacy. In parallel, store the data in the local cache.
[0147] S340: Aggregate the knowledge sampling results of each local end.
[0148] Preferably, as Figure 3 shown, the hot and cold data dual storage mechanism of the hot and cold data hierarchical storage module 320 is the data management strategy of the local cache. It divides the data into "hot data" and "cold data" according to the access frequency and importance, and stores them on different types of storage media respectively. Hot data is the data that is frequently accessed and updated, and is stored in the hot data storage medium 321 - SRAM. SRAM has a faster access speed and is usually used for the high-speed part of the cache. Cold data is the data that is less frequently accessed, and is stored in the cold data storage medium 322 - DRAM. DRAM has a higher capacity and a relatively slower speed, and is usually used for the low-speed part of the cache system. The advantage of this hot and cold data dual storage mechanism is that it can significantly improve the access efficiency of the cached data and the system performance. By storing the frequently accessed hot data in the faster SRAM, i.e., the hot data storage medium 321, the query latency can be greatly reduced, ensuring that critical data can be quickly obtained to ensure the optimal access speed. Storing the infrequently accessed cold data in the larger-capacity DRAM, i.e., the cold data storage medium 322, can optimize the overall resource allocation of the system without sacrificing the storage capacity.
[0149] Traditional single - layer cache architectures often lead to excessive server - side loads when dealing with frequent knowledge retrieval requests. At the same time, due to the inability to fully utilize local storage resources, data transfer latency and overall query inefficiency become prominent problems. In contrast, the dual - storage mechanism for hot and cold data can significantly improve these issues. Specifically, the server - side is responsible for caching frequently accessed and general knowledge, while the local side focuses on caching personalized knowledge. This design enables the system to preferentially query the local cache, thereby reducing dependence on the server - side and effectively reducing data transfer and query times. In addition, the dual - storage mechanism for hot and cold data ensures that frequently used data always resides in the local cache, not only increasing the data reuse rate but also maximizing the utilization efficiency of knowledge, and thus significantly enhancing the overall system performance and response speed.
[0150] Preferably, when the access frequency of data changes, the hardware accelerator 200 can trigger data migration according to the historical access record 282 and the importance of the data through the pre - set pre - fetch unit 280, thereby dynamically adjusting the division of hot and cold data, enabling the system to maintain efficient operation when dealing with different loads.
[0151] The pre - fetch unit 280 is used to predict and load data nodes that may be accessed during the query process, in order to reduce data access latency and improve the overall system performance. As Figure 3 shown, the pre - fetch unit 280 includes two types of modules: a node transition matrix 281 and a historical access record 282.
[0152] The node transition matrix 281 is used to analyze the access relationships between nodes and predict the nodes that may be accessed during the query process based on hardware logic. The historical access record 282 is used to save the node access trajectories in previous queries.
[0153] The pre - fetch unit 280 monitors the access count or timestamp of each data item according to the historical access record 282. Once it detects that the access frequency of a certain data exceeds the preset threshold, the pre - fetch unit 280 will initiate the migration process. Then, the low - frequency data will be removed from the hot data storage medium 321 and stored in the slower cold data storage medium 322. At the same time, the frequently accessed data will be migrated from the cold data storage medium 322 to the hot data storage medium 321 to ensure its fast access. The migration process is controlled by the pre - fetch unit 280 to ensure the efficient and low - latency completion of data movement.
[0154] As described above, the hardware accelerator 200 reduces the number of cache misses by improving the locality of cache access and optimizes the data loading efficiency.
[0155] This optimization not only improves the response speed of the query but also reduces the waste of memory resources, ultimately achieving higher query throughput and lower power consumption.
[0156] Preferably, in order to reduce data loading latency, the node transition matrix 281 predicts the nodes that may be accessed during the query process based on the Markov chain prediction algorithm.
[0157] Specifically, before the query starts, the node transition matrix 281 establishes a Markov chain model based on the historical access pattern to predict the next node access path. The Markov chain prediction algorithm is a probability model based on historical state transitions, used to predict the future state of the system according to the current state, assuming that the future state is only related to the current state and has nothing to do with the past state.
[0158] As Figure 3 shown, before the start of each subtask, the node transition matrix 281 analyzes the historical access record data and constructs a Markov chain model based on the sub-regions R1, R2, …, R k sub-regions divided by the knowledge graph.
[0159] S351: Collect historical access patterns.
[0160] Record the access relationship between nodes: P(i→j), representing the transition probability from node i to node j. For example, the access frequency from node A to node B is expressed as P(A→B).
[0161] S352: Construct a Markov chain model.
[0162] The node transition matrix 281 calculates the probability transition matrix P of each node transition based on the historical access record data. The entire transition process can be represented by the probability transition matrix P:
[0163]
[0164] S353: Predict node access.
[0165] According to the currently accessed node X, predict the next node access path. The predicted next node Y can be determined by maximizing the transition probability P(X→Y), that is:
[0166]
[0167] where Y represents the most likely node to transfer from the current node X.
[0168] In the actual operation process, such as Figure 3 and Figure 4As shown in the figure, the prefetch unit 280 in the hardware accelerator 200 constructs a Markov chain model by analyzing the historical access record 282 and the node transition matrix 281 to predict the graph node paths that may be accessed during the query process. Through the Markov chain model, the prefetch unit 280 can load the data of these possible access paths from the data loader 270 into the global memory in the computing module 290 of the hardware accelerator 200 in advance before the query task officially starts. This preloading mechanism reduces the performance bottleneck caused by data access latency during the query process, thereby improving the query efficiency.
[0169] The reason why the Markov chain model can achieve this is that it describes the access patterns between nodes based on the probability transition matrix, so as to load the nodes that may be accessed next into the internal memory of the computing module 290 in advance. During the query process, this prediction enables the system to prepare the data in advance, so that the query task can obtain the data immediately when needed, thus significantly improving the system response speed.
[0170] The query process of the local knowledge graph 410 is as follows:
[0171] S361: The hardware accelerator 200 accesses the local knowledge graph 410G through the data loader 270 local . The data loader 270 divides the knowledge graph into multiple data blocks suitable for the size of the on-chip memory 273 of the hardware accelerator 200. Each data block contains a part of the graph structure, such as nodes and edges. Suppose the knowledge graph is divided into R1, R2, …, R k regions. The data loader 270 decomposes the task into multiple subtasks job i according to the corresponding relationship between the central word of the query and the sub-regions of the knowledge graph (see the following example), and distributes them to the task queue. Each subtask still retains its own central word c q .
[0172] Specifically, for each query task q i , i ∈ [1, n], where n is the total number of query tasks. The data loader 270 decomposes the query task q i into m subtasks job i , i ∈ [1, m], where m is the total number of subtasks. Each subtask corresponds to a graph region for query.
[0173] Specifically, the central word of each query will be queried in all graph regions. Therefore, each central word will generate multiple subtasks, and these subtasks are executed in parallel in each region. Suppose the query q iThe central words are "hummingbird" and "habitat", and the knowledge graph is divided into two regions, R1 and R2. Then, the query task will be decomposed into six subtasks, job1, job2, job3…, job6. Each subtask corresponds to a combination of a central word and a region. For example, job1 represents querying for the central word "hummingbird" in region R1, and returning all the neighbors of the central word "hummingbird" in region R1 and their relationships with the central word "hummingbird", such as "habitat: tropical rainforest" and "food: nectar" of the "hummingbird".
[0174] The advantage of this task decomposition and distribution strategy of the data loader 270 is that it can significantly improve the parallel processing ability of the query task and make full use of the parallel computing advantage of the hardware accelerator 200. By dividing the knowledge graph into multiple regions and decomposing the query task into multiple subtasks, the hardware accelerator 200 can process multiple tasks simultaneously, reducing the overall processing time.
[0175] The effectiveness of this optimization mainly stems from the rationality of region division and task decomposition. Dividing the knowledge graph by region can not only effectively narrow the search scope of each query, but also reduce memory access conflicts and improve the locality of data access. Decomposing the task into subtasks and maintaining their respective central words can ensure the independence and integrity of the tasks, enabling the hardware accelerator 200 to efficiently process complex queries and quickly return results. This strategy ultimately improves the throughput and response speed of the system, especially when dealing with large-scale graph data, and the performance advantage is particularly significant.
[0176] S362: The decomposed subtasks are transferred by the data loader 270 to the task queue and queued according to the priority for processing. The task queue ensures the orderly execution of tasks and can efficiently schedule computing resources.
[0177] Specifically, the determination of the priority can be based on the following factors:
[0178] Regional data load: Give priority to processing regions with lighter loads to avoid delays caused by task backlogs in some regions.
[0179] Task arrival time (in accordance with the FIFO first-in, first-out principle): When the loads are the same, give priority to processing the tasks that arrive first to ensure that the tasks are processed in order.
[0180] In the present invention, the priority formula can be defined as:
[0181]
[0182] Among them, P job represents the priority, Load represents the load of the current region, T arrivalLet \(\lambda\) represent the task arrival time, and \(\delta\) be the time weight parameter. This priority formula combines task load and arrival time to ensure efficient allocation of system resources and reduce task latency. Through this mechanism, the system can effectively balance the load and optimize task scheduling.
[0183] S363: The tasks in the task queue are sent to the internal computing module 290 by the task distribution unit 312 in the on-chip controller 311 according to the priority results.
[0184] Specifically, assume that the central words of query q i are "giant panda" and "China", and the knowledge graph is divided into three regions R1, R2, and R3. The query task is decomposed into six subtasks job1, job2, job3…, job6, and each subtask corresponds to a combination of a central word and a region. For example, job1 represents querying for the central word "giant panda" in region R1. In the case of parallel processing of all subtasks job i The formula for calculating the total time of parallel computing is:
[0185]
[0186] where is the execution time of the \(m\)th subtask, and finally the query task is completed through parallel processing.
[0187] This parallel execution method can greatly improve the computing efficiency and reduce the overall latency of query operations. By simultaneously processing multiple subtasks with multiple computing units, the hardware resources can be fully utilized, and the concurrent processing ability of the system can be improved, so as to complete a large number of query tasks in a short time. The advantage of this architecture is that it can significantly reduce the load of a single computing unit and avoid resource bottlenecks.
[0188] The benefits of parallel execution are also reflected in the flexibility of task scheduling. Since the computing units can simultaneously process graph query tasks in different regions, the system can dynamically adjust the task distribution strategy according to the real-time load situation, further optimizing the resource utilization rate. Such a design can ensure that the system still maintains a high response speed and processing efficiency under high load, and is suitable for large-scale and high-frequency data query scenarios.
[0189] S364: The computing module 290 transmits the calculation results of each subtask to the differential privacy processing unit 300. The differential privacy processing unit 300 adds noise to protect sensitive information and ensure that user privacy is not leaked in the scenario of multi-source knowledge extraction.
[0190] The differential privacy processing unit 300 performs differential privacy noise addition processing and further cosine similarity screening on the candidate neighbors retrieved by the graph query. Cosine similarity screening is a method based on the vector space model for measuring the similarity between two vectors, and it determines the degree of similarity by calculating the cosine value of the angle between them. Cosine similarity screening is commonly used in scenarios such as text, images, and knowledge graphs, and it can effectively screen out the items most relevant to the query term, reduce noisy data, and improve the relevance and accuracy of the results. The purpose of the further cosine similarity screening is to prevent the situation where there is too much neighbor information of the central word in the current query, and ensure that the knowledge items most likely to be reused are stored in the multi-level cache.
[0191] The main steps of the differential privacy processing unit 300 are as Figure 7 shown.
[0192] S371: Receive the graph query subtask.
[0193] Receive the query q i 、the central word c q 、the knowledge graph G local 、the privacy budget ∈ and the number of entities n to be returned.
[0194] S372: Calculate the connection strength vector of the neighbors.
[0195] According to the neighbor nodes related to the central word c local in the local knowledge graph 410G q ,calculate the connection strength vector X. The connection strength vector X contains all entities in the current graph and initializes each dimension to 0, sets the vector dimension corresponding to the central word c q and all its neighbor entities to 1, and sets the dimension of unrelated entities to 0.
[0196] Specifically, the principle of calculating the connection strength vector is to measure the degree of association between the central word of the query and the neighbor nodes in the local knowledge graph 410. For each central word in the central word c q (each central word of the query is a set and can include multiple words, see the previous example of the giant panda), each dimension of the connection strength vector X corresponds to an entity, and the initial value is 0. When the central word c q has a direct connection with a certain neighbor entity e i , set the corresponding dimension of the connection strength vector X to 1.
[0197] The calculation formula for the connection strength vector is:
[0198]
[0199] Here, X i represents the connection strength vector X at the entity ei Values on the dimension. In this way, entities related to the query can be quickly determined, thus optimizing the efficiency of graph queries.
[0200] S373: Add Laplace noise to the connection strength vector to obtain candidate neighbors.
[0201] To protect the privacy of neighbor information, Laplace noise Lap(1 / ∈) is added to the strength vector X to obtain the noisy connection strength vector X ′ . X ′ = X + Lap(1 / ∈).
[0202] S374: Calculate the cosine similarity between the candidate neighbors and the query.
[0203] For the noisy connection strength vector X ′ Sort according to the numerical values of each dimension of the vector, and select the top n entities as the candidate neighbor set R for preliminary retrieval nbr .
[0204] For each entity v nbr in the candidate neighbor R i , calculate its cosine similarity with the query q as the relevance score. The calculation formula is:
[0205]
[0206] where I(q, v i ) represents the cosine similarity, and q and v i represent the query q and the entity v i embeddings respectively.
[0207] S375: Screen to obtain the final neighbor set for the current subtask. Select the entity with the highest score as the final knowledge set R.
[0208] The purpose of using cosine similarity screening is to ensure that only the most relevant part of the neighbor information of the central word of the current query is retained, preventing waste of cache resources due to excessive neighbor information. If the candidate neighbor set only processed by differential privacy is not further screened, when the scale of the knowledge graph is large, the number of neighbors will be a quite large order of magnitude, which will lead to too long processing time for each subtask and a large amount of redundant data that is not helpful for answering questions. By calculating the cosine similarity scores between each neighbor entity and the query q, their semantic similarity can be effectively measured, and the neighbors with the closest semantics to the central word can be screened out.
[0209] The beneficial effect of this screening process is that it can significantly improve the utilization efficiency of the multi-level cache. Since the cache stores only the knowledge items most relevant to the query, this not only reduces the probability of cache miss failures but also enhances the overall performance and response speed of the query. In addition, preferentially storing neighbor information with high similarity can also more accurately provide valuable data in subsequent queries, avoiding unnecessary data transmission and processing, thereby improving the overall resource utilization rate and query accuracy of the system.
[0210] S376: Determine whether there are remaining subtasks. If so, restart and then receive the graph query subtask. If not, proceed to step S350.
[0211] That is to say, the calculation result is processed by the differential privacy processing unit 300 to form the knowledge set Re i , i ∈ [1, m], where m is the total number of subtasks. Transfer the knowledge set Re i to the result aggregation unit 330, and end.
[0212] S350: The knowledge set Re i is sent from the differential privacy processing unit 300 to the result aggregation unit 330 to integrate the calculation results of multiple subtasks.
[0213] During the integration process, the result aggregation unit 330 mainly performs a deduplication operation on the results of each subtask, removing duplicate items to ensure that each result appears only once. After all the deduplicated results are aggregated, a complete query result set R kw is formed. The advantage of this deduplication operation is that it significantly improves the accuracy of the query results, ensures that each knowledge item appears only once, and avoids information redundancy and user confusion. In addition, the streamlined result set reduces the storage and transmission overheads, improving the overall efficiency of the system. Deduplication also optimizes the speed and resource utilization of subsequent processing, enabling the system to respond more quickly to query requests.
[0214] Specifically, assume that the central words of the query are "hummingbird" and "habitat". The results returned by each subtask include "Habitat: Tropical rainforest" and "Food: Nectar" for "hummingbird", and "Environment; Tropical rainforest" and "Habitat animals: Hummingbird, American crocodile" for "habitat". The result aggregation unit 330 integrates these results together, removes duplicate items through the deduplication operation, and finally forms a complete query result set, including the habitat and food information of "hummingbird", as well as the environmental and habitat animal-related information of "habitat". The generated query result R kw is written back by the result aggregation unit 330 to the local cache of the hardware accelerator 200 so that subsequent similar queries can be quickly hit, reducing duplicate calculations.
[0215] S400: Knowledge Aggregation Stage: The caching module 120 combines and organizes the retrieval results from different sources to ensure data consistency and integrity.
[0216] The combination and organization process is similar to the result aggregation unit 330 processing process in the above deep search. During the combination and organization process, the caching module 120 combines and deduplicates the results to ensure that each result appears only once. The combined and organized data is stored in the server cache, providing a faster access path for subsequent queries and processing, reducing the overhead of repeated retrievals, and improving the overall system performance.
[0217] S500: Enhanced Answer Stage: The caching module 120 on the server side sends the combined and organized knowledge data to the data processor 110, which is used to enhance the answer ability of the large language model.
[0218] Specifically, the data processor 110 textifies the query result K in the initial search stage from the caching module 120 L1 or the query result R returned in the knowledge aggregation stage kw and inputs it as a prompt into the large language model. The LLM performs an inference process to generate an enhanced answer with text information.
[0219] In this way, the answer of the LLM can incorporate information from multiple local knowledge graphs, enhancing the accuracy, depth, and relevance of the answer. This answer method based on aggregated knowledge enables the system to provide more in-depth and comprehensive answers when processing complex queries, meeting the diverse needs of users.
[0220] The technical advantages of the present invention include:
[0221] Traditional knowledge retrieval systems usually only rely on software-level optimizations and are difficult to fully utilize the advantages of hardware resources, resulting in slow processing speeds and inability to meet the needs of large-scale concurrent queries. To solve this problem, the hardware accelerator 200 of the present invention introduces various optimization means, such as parallel computing, cache access optimization, and prefetching technology, etc., significantly improving the knowledge retrieval efficiency. Among them, parallel computing accelerates the processing process of query tasks; cache locality optimization reduces data access latency; the prefetching technology preloads data blocks that may be needed in advance, further reducing the waiting time. These measures work together to effectively alleviate the bottleneck problem in the retrieval process, enabling the system to extract the required information from the local knowledge graph in a shorter time, significantly improving the retrieval efficiency, and meeting the real-time requirements.
[0222] Current knowledge retrieval systems are prone to exposing sensitive information when integrating multi-source data and lack effective privacy protection mechanisms. Especially in federated learning and distributed data environments, the risk of user privacy leakage is particularly prominent. To address this shortcoming, the differential privacy mechanism proposed in this invention provides protection by adding noise to sensitive data to ensure that user privacy is not leaked. Specifically, during the federated knowledge retrieval process, the differential privacy mechanism applies noise processing to the connection strength vectors of neighbor nodes to mask specific access patterns and sensitive relationships. Even when cross-source data integration is performed in a distributed environment, the differential privacy mechanism can effectively prevent the leakage of sensitive information, ensure the security and privacy of user data, and thus enhance the reliability of the system when processing sensitive data.
[0223] It should be noted that the above-mentioned specific embodiments are exemplary, and those skilled in the art can come up with various solutions inspired by the disclosure of the present invention, and these solutions also fall within the scope of the disclosure of the present invention and fall within the scope of protection of the present invention. Those skilled in the art should understand that the present invention specification and its drawings are illustrative and do not constitute a limitation on the claims. The scope of protection of the present invention is defined by the claims and their equivalents. The present invention specification contains multiple inventive concepts, such as "preferably" and "according to a preferred embodiment", which means that the corresponding paragraph discloses an independent concept, and the applicant reserves the right to file a divisional application based on each inventive concept.
Claims
1. A federated knowledge retrieval and large language model enhancement system, characterized in that, It includes a server (100) and a hardware accelerator (200). The server (100) preprocesses the input query to extract the central word of the query; and retrieves the central word during the caching process through its own cache module (120). If the cache hits, the cache module (120) provides the query result corresponding to the central word to the data processor (110) to perform inference through a large language model and generate an enhanced answer. If the cache of the cache module (120) misses, the server (100) sends a retrieval instruction to the hardware accelerator (200) to retrieve in the local cache of the hardware accelerator (200). If the local cache still misses, the hardware accelerator (200) calls the local knowledge graph (410) and generates parallel subtasks associated with the central word in a way of splitting the local knowledge graph (410), so as to perform the in-depth search, and adds noise to the knowledge items retrieved by the parallel subtasks to protect sensitive information. The hardware accelerator (200) integrates the knowledge items with added noise and sends them to the data processor (110) through the cache module (120) to perform inference through a large language model and generate an enhanced answer.
2. The system according to claim 1, wherein The hardware accelerator (200) includes: A data loader (270) that batches data from the local knowledge graph (410) and splits the knowledge graph into multiple data blocks suitable for the size of the on-chip memory (273) on the hardware accelerator (200). A computing module (290) that performs parallel computing. A differential privacy processing unit (300) that adds noise to sensitive data during the processing to ensure that data privacy is not leaked. A result aggregation unit (330) that summarizes the output information of each computing unit and performs comprehensive analysis and processing to ensure the integrity and consistency of the final result.
3. The system according to claim 1 or 2, characterized in that, The data loader (270) further includes a task decomposition unit (272). During the data loading process, the task decomposition unit (272) searches for the central word of the query in all the split sub-regions of the knowledge graph.
4. The system according to any one of claims 1 to 3, characterized in that, The hardware accelerator (200) further includes an on-chip controller (311), and the on-chip controller (311) includes: A task distribution unit (312) that receives and distributes subtasks from the task queue to ensure workload balance of each computing unit in the computing module (290). An instruction decoder (313) that converts high-level task instructions into specific hardware operation instructions to guide the computing unit to execute corresponding tasks. A status management unit (314) that monitors the running status of each module in real time and coordinates resource allocation to ensure the correctness of task execution and the stability of the system.
5. The system according to any one of claims 1 to 4, characterized in that, The hardware accelerator (200) further includes a prefetch unit (280). The prefetch unit (280) predicts potential data requirements using the historical access pattern, pre-fetches data for knowledge graph nodes based on the Markov chain model, and thus pre-loads data in advance to reduce query latency.
6. The system according to any one of claims 1 to 5, characterized in that The prefetch unit (280) of the hardware accelerator (200) predicts the nodes accessible during the query process based on a Markov chain model; wherein, Before the query starts, a Markov chain model is established based on the historical access patterns; Based on the Markov chain model, the node access path is predicted.
7. The system according to any one of claims 1 to 6, characterized in that, The server (100) includes a data processor (110) and a cache module (120), The data processor (110) extracts the central word and runs a primary search program, and sends the extracted central word to the cache module (120); The cache module (120) searches for whether there is a corresponding knowledge item; When the knowledge item of the central word exists, the cache module (120) directly returns the query result to enter the enhanced answering stage, and the query result is composed of the neighbors of the central word, The data processor (110) texturizes the query result and inputs it into a large language model for reasoning to generate an enhanced answer.
8. The system according to any one of claims 1 to 7, characterized in that, The processing steps of the data processor (110) in the server (100) further include: When no relevant knowledge item is found in the primary search, a deep search request is sent to each local computer (400) via the Internet; The local computer (400) receives the query request from the server (100) through the network adapter (220); The local computer (400) runs a deep search program based on the query request, The deep search program accesses the accelerator chip (260) of the hardware accelerator (200) through the high-speed bus (310).
9. A method for federated knowledge retrieval and large language model enhancement, characterized in that The method includes: The server (100) preprocesses the input query to extract the central word of the query; and retrieves the central word through its own cache module (120) during the caching process. If the cache hits, the cache module (120) performs reasoning on the query result corresponding to the central word through a large language model and generates an enhanced answer; If the cache of the cache module (120) misses, the server (100) sends a retrieval instruction to the hardware accelerator (200) to retrieve in the local cache of the hardware accelerator (200); If the local cache still misses, the hardware accelerator (200) calls the local knowledge graph (410) and generates parallel subtasks associated with the central word in a way of splitting the local knowledge graph (410), so as to perform the deep search, and adds noise to the knowledge items retrieved by the parallel subtasks to protect sensitive information, The hardware accelerator (200) integrates the knowledge items with noise and performs reasoning through a large language model and generates an enhanced answer.
10. The method according to claim 9, wherein The local cache is set with a dual storage mechanism for hot and cold data to improve the retrieval efficiency; Among them, cold data and hot data are stored on different types of storage media; Monitor the access times or timestamps of data items, When it is detected that the access frequency of a certain data exceeds a preset threshold, start the migration process; Remove the low-frequency data from the hot data cache area and store it in the slower cold data area; At the same time, migrate the frequently accessed data from the cold data area to the hot data cache to ensure its fast access.
Citation Information
Cited By
Retrieval enhancement generation large language model system and method
CN120562570A
Retrieval-augmented generation large language model system and method
CN120562570B
Chip, intelligent agent system based on LLM and intelligent agent equipment
CN120973417A
Deep learning task-oriented operating system kernel parameter tuning system and method
CN121092223A
Operating system kernel parameter tuning system and method for deep learning tasks
CN121092223B