A memory-aware dynamic permission scheduling and knowledge retrieval method for multi-agent medical question answering

CN122838581APending Publication Date: 2026-09-29SHENZHEN GUOJIAN INTELLIGENT COMPUTING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611237658.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-14
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0009]本发明旨在克服现有多智能体医疗问答系统显存竞争激烈、权限检索效率低、知识串话严重、资源与权限无法协同调度的缺陷,提供一种面向多智能体医疗问答的显存感知动态权限调度与知识检索方法,实现低配离线终端下多专科智能体稳定并发、细粒度知识权限管控、低延迟精准检索与动态显存资源优化

Benefits of technology

[0028]与现有技术相比,本发明具有如下技术优势:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838581A_ABST
    Figure CN122838581A_ABST
Patent Text Reader

Abstract

This invention discloses a memory-aware dynamic permission scheduling and knowledge retrieval method for multi-agent medical question answering, belonging to the field of medical artificial intelligence technology. This invention constructs a multi-level knowledge base including a global public library, a specialty private library, and a cross-specialty shared library, configuring permission tag vectors containing identity, level, and domain information for each knowledge element. This invention monitors the GPU memory load of the terminal in real time, dynamically executes hierarchical loading scheduling of the knowledge base based on the agent's call popularity index, embeds three-dimensional permission verification into the retrieval process, achieves fine-grained knowledge access control, performs domain fusion and conflict resolution on multi-agent retrieval results, and implements hierarchical memory caching management based on popularity. This invention solves the technical problems of intense concurrent memory contention, low permission retrieval efficiency, and cross-referencing of specialty knowledge on low-configuration offline terminals in hospitals. It can significantly improve the system's concurrency and question-answering accuracy while ensuring the localization and compliance of medical data, and is suitable for hospital-private medical agent question answering scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical artificial intelligence large model knowledge base retrieval and edge resource scheduling technology. Specifically, it relates to a memory-aware dynamic permission scheduling and knowledge retrieval method for multi-agent medical question answering, which is applicable to the private deployment of offline industrial control terminals in hospitals and concurrent medical question answering scenarios involving multiple specialty agents. Background Technology

[0002] With the popularization of large-scale medical models and multi-agent technologies, in-hospital medical question-answering systems based on multi-specialty agent collaboration have been gradually implemented. However, in actual deployment scenarios involving offline, private, and low-configuration GPU industrial control terminals in hospitals, existing technologies suffer from several key technical defects, severely limiting system concurrency, retrieval accuracy, and data security. These defects include:

[0003] 1. Severe competition for video memory resources among multiple agents. Existing medical multi-agent systems adopt a static video memory allocation mechanism, in which each specialty agent independently loads vector knowledge bases and model resources. During concurrent calls, video memory usage increases linearly, which can easily trigger video memory overflow and crashes. This makes the system unsuitable for the long-term stable operation of low-configuration fixed computing power terminals within hospitals.

[0004] 2. Disjointed and high-latency access control between permission verification and knowledge retrieval processes. Traditional knowledge base access control uses a token verification mechanism independent of the retrieval process. Each retrieval requires complete permission verification, resulting in a large amount of redundant computational overhead. Furthermore, the permission policy is static and fixed, making it impossible to dynamically adapt access priorities based on the real-time load of the device.

[0005] 3. The granularity of specialist knowledge isolation is coarse, leading to knowledge cross-referencing and redundant storage. The existing two-layer isolation scheme of public and private repositories only achieves repository-level isolation and does not impose permission and domain constraints on the smallest knowledge unit. General medical knowledge is stored repeatedly across specialties, and cross-specialty mismatches and question-answer cross-referencing problems are prone to occur, affecting the accuracy of medical Q&A.

[0006] 4. Resource scheduling, access control, and knowledge retrieval are optimized independently without a closed-loop collaborative mechanism. Existing technologies separate memory resource management, access control, and knowledge retrieval into independent modules, failing to form an end-to-end dynamic linkage strategy. This results in low overall system concurrency throughput, poor resource utilization, and an inability to adapt to scenarios with limited offline computing power within the institution.

[0007] In summary, existing technologies lack an integrated solution that is suitable for offline deployment within hospitals, can dynamically allocate resources based on memory load and agent popularity, and can simultaneously achieve fine-grained knowledge access control and knowledge conflict resolution. Summary of the Invention

[0008] Purpose of the invention

[0009] This invention aims to overcome the shortcomings of existing multi-agent medical question-answering systems, such as intense memory contention, low efficiency of permission retrieval, serious knowledge crosstalk, and inability to coordinate resource and permission scheduling. It provides a memory-aware dynamic permission scheduling and knowledge retrieval method for multi-agent medical question-answering, which enables stable concurrency of multiple specialty agents, fine-grained knowledge permission control, low-latency accurate retrieval, and dynamic memory resource optimization on low-configuration offline terminals.

[0010] Technical solution

[0011] To achieve the above objectives, the present invention adopts the following technical solution.

[0012] A memory-aware dynamic permission scheduling and knowledge retrieval method for multi-agent medical question answering is proposed, which runs on an offline industrial control computing terminal within a hospital. The terminal deploys a quantitative medical large language model, and includes the following steps:

[0013] S1. Construct a multi-level medical knowledge base system and a knowledge element permission tag system;

[0014] A three-layered, isolated, and cross-domain-shareable medical knowledge base system is constructed, including a global public medical knowledge base, independent private knowledge bases for each specialty agent, and a cross-specialty shared knowledge base. Knowledge elements are used as the smallest knowledge unit, and permission tag vectors are configured for each knowledge element. The permission tag vectors include a list of accessible agent identities, the minimum access level, and the specialty domain code.

[0015] S2. Establish an intelligent agent identity mapping and call popularity statistics mechanism;

[0016] Each specialized medical intelligent agent is assigned a unique identity, and a binding mapping relationship is established between the intelligent agent and the accessible knowledge base. The call frequency of each intelligent agent is counted in real time based on a sliding time window, and an intelligent agent call heat index is generated for resource scheduling.

[0017] S3. Receive and parse the user's medical consultation request, locate the target and invoke the intelligent agent;

[0018] It receives medical Q&A requests initiated by in-hospital terminals, parses the request business domain and Q&A intent, identifies the target specialty intelligent agent to be invoked, and obtains the real-time invocation popularity index of the target intelligent agent.

[0019] S4. Perform dynamic memory-aware resource scheduling based on memory load and agent popularity.

[0020] The system collects the GPU memory usage status of the terminal in real time and dynamically switches the knowledge base loading strategy based on the memory load range and the agent popularity index, so as to realize hierarchical resource loading and rate limiting scheduling under different loads.

[0021] S5. Perform multi-dimensional linked permission verification and intelligent retrieval based on knowledge element permission tag vectors;

[0022] Permission verification is embedded into the knowledge retrieval process. The target knowledge element is jointly verified from three dimensions: agent identity, access permission level, and professional domain matching. After the verification is passed, a precise retrieval is performed. For low-confidence retrieval results, cross-database supplementary retrieval is automatically triggered.

[0023] S6. Domain fusion and knowledge conflict resolution of multi-agent parallel retrieval results;

[0024] The knowledge fragments output concurrently by multiple agents are categorized by domain encoding, and knowledge within the same domain is fused using confidence weighting. When there are content conflicts between cross-domain knowledge, priority is determined based on the knowledge element permission level, resulting in a unified and consistent set of medical knowledge results.

[0025] S7. Perform hierarchical cache cleanup and dynamic release of video memory based on the agent's popularity;

[0026] After a single question-and-answer task is completed, the vector cache is retained or cleared based on the intelligence agent's call popularity. The cache of high-popularity intelligence agents is retained with a delay, while the cache of low-popularity intelligence agents is released immediately. The background automatically resumes the preloading mechanism based on the decline of the video memory load.

[0027] Beneficial effects

[0028] Compared with the prior art, the present invention has the following technical advantages:

[0029] 1. This invention enables dynamic adaptive scheduling of video memory resources, significantly improving the concurrency capabilities of low-spec terminals. It implements a tiered loading strategy based on video memory load ranges and agent popularity, avoiding resource waste and overflow issues caused by static allocation, and significantly improving the number of concurrent agents and system stability.

[0030] 2. Deep integration of permission verification and knowledge retrieval reduces retrieval latency. The traditionally separate permission verification process is embedded into the knowledge element retrieval process, eliminating redundant verification overhead and improving the response speed of medical question-and-answer retrieval.

[0031] 3. Fine-grained knowledge element access control effectively suppresses cross-disciplinary knowledge mismatch. Through three-dimensional tag constraints based on identity, level, and domain, precise access control at the individual knowledge level is achieved, resolving issues of mismatch and cross-disciplinary knowledge and improving the professionalism of question-and-answer sessions.

[0032] 4. Resource, permission, and retrieval closed-loop collaborative optimization, adapted to in-hospital offline private scenarios. The entire process is executed entirely locally, without external network data interaction, meeting medical data privacy compliance requirements, and is compatible with common in-hospital industrial control terminal hardware conditions such as 8GB and 16GB. Attached Figure Description

[0033] Figure 1This is a flowchart of the overall method of the present invention;

[0034] Figure 2 This is a schematic diagram of the knowledge element permission tag vector structure;

[0035] Figure 3 This is a schematic diagram of the dynamic scheduling state machine for video memory.

[0036] Figure 4 This is a schematic diagram of the three-dimensional permission verification and retrieval process;

[0037] Figure 5 A schematic diagram of the multi-agent knowledge fusion and conflict resolution process;

[0038] Figure 6 This is a schematic diagram of the overall offline deployment architecture of the system. Detailed Implementation

[0039] The present invention will be further described in detail below with reference to specific embodiments, but the present invention is not limited to the following embodiments.

[0040] The method of this invention is deployed on an offline industrial control terminal within a hospital, equipped with a quantitative large language model, constructing a three-layer medical knowledge base, and configuring three-dimensional permission tag vectors for all knowledge elements; the system monitors memory usage in real time and counts the call popularity of intelligent agents, realizing a closed-loop process of dynamic scheduling, permission retrieval, knowledge fusion and hierarchical memory reclamation.

[0041] Example 1: Complete loading mode in low-load video memory idle scenarios

[0042] When the terminal's GPU memory usage is in a low-load range, the system is considered idle and performs a full private knowledge base vector index loading on the target specialty agent. After receiving a user's specialty medical consultation request, it completes three-dimensional permission verification and precise retrieval. If the confidence level of the retrieval results meets the requirements, the knowledge content is directly output. After the question-and-answer session, frequently accessed, high-population agents are cached and retained with a delay to ensure smooth continuous question-and-answer flow. This mode can maximize the completeness of retrieval and the accuracy of question-and-answer when the terminal's computing power is sufficient.

[0043] Example 2: Heat-based Tiered Scheduling Mode under Medium Memory Load Scenarios

[0044] When GPU memory usage is in a moderate load range, the system initiates a tiered scheduling strategy. For frequently accessed, high-frequency specialist agents, full access to the private library is retained; for low-frequency agents, only a subset of core high-frequency knowledge elements from the private library is loaded to reduce GPU memory usage. After concurrent retrieval by multiple agents, domain classification and confidence fusion are performed, and the question-and-answer results are output after no conflicts are found. The cache of low-frequency agents is immediately cleared after the question-and-answer session ends to quickly release GPU memory resources and achieve load balancing.

[0045] Example 3: Lightweight Search and Delayed Loading Mode in Urgent Scenarios with High Memory Load

[0046] When GPU memory usage reaches a high-load critical range, the system enters an emergency scheduling state. All new agent requests only load the lightweight index of the cross-specialty shared knowledge base, no longer loading the complete private library vector resources. Priority is given to retrieving data from the shared knowledge base. If the result confidence is insufficient, private knowledge elements are loaded late as needed and released immediately after each call to avoid GPU memory overflow. This mode ensures that the terminal does not crash or interrupt under ultra-high concurrency scenarios, significantly improving the fault tolerance of offline deployment.

Claims

1. A memory-aware dynamic permission scheduling and knowledge retrieval method for multi-agent medical question answering, applied to an offline industrial control computing terminal within a hospital, wherein the terminal is equipped with a GPU and a quantized large language model, characterized in that... Includes the following steps: S1. Construct a multi-level medical knowledge base system: Construct a three-layer knowledge system in the local storage of the terminal, including a global public knowledge base, a specialist private knowledge base, and a cross-specialty shared knowledge base; Using vectorized knowledge elements as the smallest retrieval unit, each knowledge element is configured with a permission tag vector, which includes an accessible agent identity bitmap, a minimum access level field, and a specialty domain code. S2. Statistical analysis of agent call frequency: By monitoring the call logs of each specialized agent through the processor, the call frequency is counted based on a sliding time window, and a call frequency index representing the activity level of the agent is generated. S3. Analyze the target intelligent agent: Receive the medical question and answer request input by the user, analyze the request to determine the target specialty intelligent agent, and read the real-time call popularity index of the target intelligent agent; S4. Perform memory-aware dynamic scheduling: The GPU memory usage is obtained in real time through the memory monitoring module. Based on the load range of the memory usage and the call popularity index, the loading granularity of the three-layer knowledge system in the GPU memory is dynamically adjusted to control the memory usage to not exceed the preset threshold. S5. Execute embedded permission retrieval: In response to the retrieval request of the target intelligent agent, extract the permission tag vector of the candidate knowledge element, and perform three-dimensional verification in parallel in the GPU memory: compare the intelligent agent's identity with the identity bitmap for identity authentication, compare the intelligent agent's permission level with the level field for authorization, and compare the intelligent agent's domain with the domain code for domain filtering; after the verification is passed, return the retrieval results based on the vector similarity algorithm. S6. Fusion and Conflict Resolution: The retrieval result set returned by multiple agents is parsed, knowledge elements are clustered according to domain encoding, knowledge elements within the same cluster are fused with confidence weighting, and knowledge elements with semantic conflicts between different clusters are prioritized according to their permission level field to select target knowledge elements. S7. Tiered memory reclamation: In response to the signal that the question-answering task has ended, a differentiated memory release strategy is executed according to the call popularity index: the knowledge base vector index corresponding to the high-popularity agent is retained, and the knowledge base vector index corresponding to the low-popularity agent is released.

2. The method according to claim 1, characterized in that, The dynamic adjustment of loading granularity mentioned in step S4 specifically includes: When the video memory usage is lower than the first preset threshold, the full amount of the specialized private knowledge base corresponding to the target intelligent agent is loaded into the GPU video memory. When the video memory usage is between the first preset threshold and the second preset threshold, if the target agent's call popularity index is higher than the preset popularity threshold, then its specialized private knowledge base is fully loaded; otherwise, only the core high-frequency subset in the specialized private knowledge base is loaded. When the video memory usage exceeds the second preset threshold, only the lightweight index of the cross-specialty shared knowledge base is loaded into the GPU video memory, and specific knowledge elements in the private knowledge base are loaded into the video memory as needed when a retrieval instruction is received.

3. The method according to claim 1, characterized in that, In step S5, if the confidence level of the search results based on the private knowledge base of the specialty is lower than the preset confidence threshold, a cross-database search mechanism is triggered to call the global public knowledge base or the cross-specialty shared knowledge base for supplementary search, and the supplementary search process reuses the three-dimensional verification logic.

4. The method according to claim 1, characterized in that, The quantized large language model is a 4-bit or 5-bit quantized version of the DeepSeek-14B model, deployed locally on the terminal without an external network communication link.

5. The method according to claim 1, characterized in that, The formula for calculating the call popularity index is: H = T × N, where N is the number of times the agent calls within the sliding time window T.

6. The method according to claim 1, characterized in that, Step S7 is followed by: The background daemon continuously monitors the GPU memory load. When the memory load drops below the first preset threshold, it preloads the knowledge base vector index that is likely to be called into the GPU memory based on the historical call popularity index of each agent.

7. A memory-aware dynamic permission scheduling and knowledge retrieval system for multi-agent medical question answering, characterized in that, include: The knowledge base module is configured in the terminal's local storage and is used to store the global public knowledge base, the specialized private knowledge base, and the cross-specialty shared knowledge base; The knowledge base module uses vectorized knowledge elements as the smallest unit, and associates each knowledge element with a permission tag vector containing an identity bitmap, a level field, and a domain code. The popularity statistics module, coupled to the knowledge base module, is configured to count the call frequency of each intelligent agent based on a sliding time window and output the call popularity index. The memory-aware scheduling module, coupled to the GPU and the heat statistics module, is configured to monitor the GPU memory usage in real time and dynamically control the loading granularity of the knowledge base module to the GPU memory based on the memory load range and the call heat index. The permission retrieval module, coupled to the memory-aware scheduling module and the knowledge base module, is configured to respond to a retrieval request by performing a three-dimensional parallel verification of the permission tag vector of the candidate knowledge element in the GPU memory, including identity, level, and domain, and returning the retrieval result that passes the verification. The knowledge fusion and resolution module, coupled to the permission retrieval module, is configured to cluster and fuse retrieval results by domain encoding, and to make a decision based on the permission level field when semantic conflicts are detected. The cache reclamation module, coupled to the memory-aware scheduling module, is configured to perform hierarchical retention or release of vector indices in GPU memory based on the call popularity index.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.