Method and device for security management of privately deployed large language model

Through API Key authentication, access frequency limitation, priority scheduling and interface standardization, the security and management efficiency issues of privately deployed large language models are solved, and secure user authentication, resource scheduling and compliance auditing are achieved, thereby improving the security and management efficiency of the system.

CN120705844APending Publication Date: 2025-09-26INSPUR SOFTWARE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510788808.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing private deployments of large language models suffer from weak access control, uneven resource allocation, lack of auditing capabilities, and insufficient interface compatibility, resulting in low security and management efficiency.

Method used

Through API Key authentication, access frequency limitation, priority scheduling, interface standardization and audit log recording, a security management system is built to ensure user identity authentication, reasonable resource allocation and compliance audit.

Benefits of technology

It improves system security and management efficiency, prevents unauthorized access, optimizes resource allocation, meets corporate compliance requirements, and reduces integration difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705844A_ABST
    Figure CN120705844A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and particularly provides a method and device for security management of a privatized deployment large language model, and the method comprises the following steps: S1, user registration and API Key distribution; s2, authentication of the API Key is carried out; s3, limiting the access frequency; s4, performing priority scheduling; s5, interface standardization; and S6, recording an audit log. Compared with the prior art, access control, resource scheduling, audit log and interface standardization functions can be integrated, and the security and management efficiency of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and specifically provides a method and device for secure management of privately deployed large language models. Background Art

[0002] In recent years, demand for large language models (LLMs) has surged in enterprise applications. To meet data privacy, compliance, and low latency requirements, many enterprises have adopted frameworks such as Ollama and vLLM for private deployment. However, existing private deployment solutions have significant shortcomings in access control and resource management, mainly reflected in the following aspects:

[0003] (1) Weak access control: Existing frameworks usually do not provide APIKey authentication mechanisms, which makes it difficult to effectively verify user identities, easily leading to unauthorized access and increasing the risk of data leakage.

[0004] (2) Uneven resource allocation: Due to the lack of access frequency restrictions and priority scheduling mechanisms, some users may over-occupy the model computing power, affecting the service quality of other users, especially in high-concurrency scenarios.

[0005] (3) Lack of auditing function: The existing system cannot systematically record user calling behavior, making it difficult to meet the company's needs for compliance review and security audit.

[0006] (4) Insufficient interface compatibility: The interface formats of private deployment models are diverse and lack standardized output (such as the OpenAI format), which increases the complexity of integration with existing tools.

[0007] The above problems limit the security and practicality of privately deployed large models in enterprise scenarios. There is an urgent need for a comprehensive solution that integrates access control, resource scheduling, audit logs, and interface standardization functions to improve system security and management efficiency. Summary of the Invention

[0008] The present invention addresses the deficiencies of the above-mentioned prior art and provides a highly practical method for secure management of privately deployed large language models.

[0009] A further technical task of the present invention is to provide a device for secure management of large language models in a privately deployed manner that is rationally designed, safe and applicable.

[0010] The technical solution adopted by the present invention to solve its technical problem is:

[0011] A method for securely managing a privately deployed large language model includes the following steps:

[0012] S1. User registration and API Key allocation;

[0013] S2, API Key authentication;

[0014] S3, access frequency limit;

[0015] S4, priority scheduling;

[0016] S5, interface standardization;

[0017] S6. Audit log records.

[0018] Furthermore, in step S1, the user management module assigns a unique APIKey to the user, sets user priority and access frequency limit, and stores the user information and APIKey in the database to ensure that the user identity is bound to the authority.

[0019] Furthermore, in step S2, the user initiates a model call request through the APIKey, and the authentication module queries the database to verify the validity of the Key. If the APIKey is invalid or expired, the request is rejected and the error code 401Unauthorized is returned.

[0020] Furthermore, in step S3, the frequency limiting module uses a bucket algorithm of a time sliding window to dynamically record the number of calls made by the user within a specified time window. If the call frequency exceeds the limit, the error code 429 Too Many Requests is returned; otherwise, the request is allowed to continue processing.

[0021] Furthermore, in step S4, the priority scheduling module allocates requests to a high-priority queue or a normal queue according to user priority. In a high-load scenario, requests in the high-priority queue are processed first to ensure the service quality of key users.

[0022] Furthermore, in step S5, the interface encapsulation module converts the original output of the large model into the standard OpenAI format and streams back requests and responses in JSON format to ensure that the output is compatible with existing tools and frameworks.

[0023] Furthermore, in step S6, the audit log module records detailed information of each call, including user ID, API key, request time, input content and model output. The log is stored in the database, queried and exported to meet compliance audit requirements.

[0024] A device for secure management of a privately deployed large language model, comprising: at least one memory and at least one processor;

[0025] The at least one memory is configured to store a machine-readable program;

[0026] The at least one processor is configured to call the machine-readable program to execute a method for secure management of a privately deployed large language model.

[0027] Compared with the prior art, the method and device for secure management of a privately deployed large language model of the present invention have the following outstanding beneficial effects:

[0028] The present invention can integrate access control, resource scheduling, audit log and interface standardization functions to improve system security and management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0030] Figure 1 This is an architectural diagram of a method for securely managing large language models deployed in a private manner;

[0031] Figure 2 The figure is a flowchart of a method for securely managing a privately deployed large language model. DETAILED DESCRIPTION

[0032] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention will be further described in detail below in conjunction with specific embodiments. Obviously, the embodiments described are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0033] A best embodiment is given below:

[0034] like Figure 1 、 2 As shown, a method for secure management of a privately deployed large language model in this embodiment includes the following steps:

[0035] S1. User registration and API Key allocation;

[0036] Assign unique API keys to users through the user management module, set user priority (high, medium, low) and access frequency limits (limit the maximum number of calls per minute). Store user information and API keys in the database to ensure that user identity is bound to permissions.

[0037] S2, API Key authentication;

[0038] When a user initiates a model call request using an API Key, the authentication module queries the database to verify the validity of the Key. If the API Key is invalid or expired, the request is rejected and the error code 401 (Unauthorized) is returned.

[0039] S3, access frequency limit;

[0040] The frequency limit module uses a bucket algorithm within a sliding window to dynamically record the number of user calls within a specified time window. If the call frequency exceeds the limit, the error code 429 (Too Many Requests) is returned; otherwise, the request is allowed to continue processing.

[0041] S4, priority scheduling;

[0042] The priority scheduling module assigns requests to either a high-priority queue or a normal queue based on user priority. In high-load scenarios, requests in the high-priority queue are prioritized to ensure service quality for critical users.

[0043] S5, interface standardization;

[0044] The interface encapsulation module converts the raw output of large models into the standard OpenAI format ( / v1 / completions), supporting streaming requests and responses in JSON format. This ensures compatibility with existing tools and frameworks, improving system usability.

[0045] S6, audit log records;

[0046] The audit log module records detailed information about each call, including user ID, API key, request time, input content, and model output. Logs are stored in a database and can be queried and exported to meet compliance audit requirements.

[0047] Based on the above method, a device for privately deploying large language model security management in this embodiment includes: at least one memory and at least one processor;

[0048] The at least one memory is configured to store a machine-readable program;

[0049] The at least one processor is configured to call the machine-readable program to execute a method for secure management of a privately deployed large language model.

[0050] This invention uses refined user authentication based on APIKey to identify the calling user, prevent unauthorized access, and improve model security. Through access frequency limiting and priority scheduling, it optimizes model resource allocation and ensures service quality for high-priority users in high-load scenarios.

[0051] Provides comprehensive audit logging capabilities to record user call behavior and meet corporate compliance requirements. Standardized model interface output is in OpenAI format, converting various large model interface protocols to reduce integration difficulty and improve system compatibility.

[0052] Build an efficient and scalable management system to meet the diverse needs of enterprise-level private deployment of large models.

[0053] The above-mentioned specific implementation methods are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above-mentioned specific implementation methods. Any technical solutions that conform to the above-mentioned specific implementation methods of the present invention and any appropriate changes or substitutions made thereto by ordinary technicians in the relevant technical field shall fall within the patent protection scope of the present invention.

[0054] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for secure management of privately deployed large language models, characterized in that: The steps are as follows: S1. User registration and API Key allocation; S2, API Key authentication; S3, access frequency limit; S4, priority scheduling; S5, interface standardization; S6. Audit log records.

2. A method for secure management of a privately deployed large language model according to claim 1, characterized in that: In step S1, the user management module assigns a unique API Key to the user, sets user priority and access frequency limit, and stores user information and API Key in the database to ensure that user identity is bound to permissions.

3. The method for secure management of a privately deployed large language model according to claim 2, characterized in that: In step S2, the user initiates a model call request using the API Key. The authentication module queries the database to verify the validity of the Key. If the API Key is invalid or expired, the request is rejected and the error code 401 Unauthorized is returned.

4. A method for secure management of a privately deployed large language model according to claim 3, characterized in that: In step S3, the frequency limit module uses a bucket algorithm based on a sliding window to dynamically record the number of user calls within a specified time window. If the call frequency exceeds the limit, an error code 429 Too Many Requests is returned. Otherwise, the request is allowed to continue processing.

5. A method for secure management of a privately deployed large language model according to claim 4, characterized in that: In step S4, the priority scheduling module allocates requests to a high-priority queue or a normal queue according to user priority. In a high-load scenario, requests in the high-priority queue are processed first to ensure the service quality of key users.

6. A method for secure management of a privately deployed large language model according to claim 5, characterized in that: In step S5, the interface encapsulation module converts the raw output of the large model into the standard OpenAI format and streams back requests and responses in JSON format to ensure that the output is compatible with existing tools and frameworks.

7. A method for secure management of a privately deployed large language model according to claim 6, characterized in that: In step S6, the audit log module records detailed information of each call, including user ID, API key, request time, input content and model output. The log is stored in the database and can be queried and exported to meet compliance audit requirements.

8. A device for privately deploying large language model security management, characterized in that: include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • A privatized large-scale model secure access control system and method

    CN122419916A