Large language model access method, device and medium
By determining the demand function information and cluster information of the target caller, and selecting matching target service clusters with performance and concurrent support, the instability and high cost problems in the use of large language models are solved, and stability and resource efficiency are improved.
Patent Information
- Application Number
- CN202511026140.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-09-02
AI Technical Summary
When using a large language model, direct connection with independent deployment clusters and multiple third-party cloud vendor clusters may lead to problems such as instability, complex access, high cost, complex management, and functional limitations.
By providing an access method of a large language model, the demand function information of the target caller is determined, the performance and concurrent support information of each service cluster are obtained, the matching target service cluster is selected to process the access request, and forward it to the caller when the target service cluster returns the result. The weight priority of the independent deployment cluster is higher than that of the third-party cloud vendor cluster, and the configuration center and distributed cache are combined to optimize request scheduling.
It achieves the improvement of the stability of the large language model while meeting functional needs, reduces comprehensive costs, simplifies management complexity, and supports multiple functional expansion, ensuring the stability of services and efficient utilization of resources.
Smart Images

Figure CN120583149A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, and medium for accessing a large language model. Background Art
[0002] Currently, when using large language models, business services typically directly connect to independently deployed clusters and various third-party cloud vendor clusters to meet their own requirements for large language model inference performance, as well as other functional and cost requirements. However, direct connection with independently deployed clusters and various third-party cloud vendor clusters may lead to instability. In summary, in the process of implementing the present invention, the inventors have at least discovered that the existing technology has the problem of instability when using large language models. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide a method, device, and medium for accessing a large language model that can meet functional requirements while ensuring the stability of using the large language model. The specific solution is as follows:
[0004] In a first aspect, the present application provides a method for accessing a large language model, comprising:
[0005] When receiving an access request from a target caller, determining required function information corresponding to the target caller;
[0006] Obtaining cluster information of each service cluster for providing a large language model service, wherein the cluster information includes performance information and concurrency support information;
[0007] Selecting, from among the service clusters, a target service cluster whose performance information matches the required function information and whose concurrent support information matches the number of access requests;
[0008] The access request is forwarded to the target service cluster, and when the target service cluster returns an access result for the access request, the access result is forwarded to the target caller.
[0009] Optionally, the major language model clusters include independently deployed clusters and third-party cloud vendor clusters;
[0010] Selecting, from each of the service clusters, a target service cluster whose performance information matches the required function information and whose concurrent support information matches the number of access requests, includes:
[0011] Selecting, from among the service clusters, a candidate service cluster whose performance information matches the required function information and whose concurrent support information matches the number of access requests;
[0012] When the number of candidate service clusters is more than two, selecting a target service cluster from the candidate service clusters based on the weights of the candidate service clusters;
[0013] Among them, the weight of the independently deployed cluster is greater than the weight of the third-party cloud vendor cluster.
[0014] Optionally, determining the required function information corresponding to the target caller includes:
[0015] Configuration information corresponding to the caller information is acquired based on the caller information carried in the access request, where the configuration information includes required function information.
[0016] Optionally, obtaining configuration information corresponding to the caller information based on the caller information carried in the access request includes:
[0017] Obtaining configuration information corresponding to the caller information from a memory based on the caller information carried in the access request;
[0018] If the configuration information corresponding to the caller information does not exist in the memory, the configuration information corresponding to the caller information is obtained from the configuration center based on the caller information carried in the access request, and the configuration information corresponding to the caller information is stored in the memory.
[0019] Optionally, the configuration information also includes a request quota limit;
[0020] Before obtaining the cluster information of each service cluster used to provide the large language model service, the following steps are also included:
[0021] Determine whether the current number of requests corresponding to the caller information exceeds the request quota limit;
[0022] When the current number of requests corresponding to the caller information does not exceed the request quota limit, the step of obtaining cluster information of each service cluster for providing large language model services is triggered.
[0023] Optionally, before determining whether the current number of requests corresponding to the caller information exceeds the request quota limit, the method further includes:
[0024] Obtaining request information corresponding to the caller information from the target cache, wherein the request information is information corresponding to the access request added to the target cache after determining the target service cluster corresponding to any access request;
[0025] The current number of requests corresponding to the caller information is counted based on the request information.
[0026] Optionally, the configuration information also includes authentication information;
[0027] After obtaining configuration information corresponding to the caller information based on the caller information carried in the access request, the method further includes:
[0028] The access request is authenticated based on the authentication information.
[0029] Optionally, also include:
[0030] The call information corresponding to the access request is written into a message queue so that a consumer can consume the call information from the message queue and obtain resource usage information.
[0031] In a second aspect, the present application provides a device for accessing a large language model, comprising:
[0032] A demand function information determination module is used to determine the demand function information corresponding to the target caller when receiving an access request from the target caller;
[0033] A cluster information acquisition module, configured to acquire cluster information of each service cluster providing a large language model service, wherein the cluster information includes performance information and concurrency support information;
[0034] a target service cluster selection module, configured to select, from among the service clusters, a target service cluster whose performance information matches the required function information and whose concurrent support information matches the number of access requests;
[0035] The access request processing module is configured to forward the access request to the target service cluster, and when the target service cluster returns an access result for the access request, forward the access result to the target caller.
[0036] In a third aspect, the present application provides an electronic device, including a memory and a processor, wherein:
[0037] The memory is used to store computer programs;
[0038] The processor is configured to execute the computer program to implement the aforementioned method for accessing the large language model.
[0039] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the aforementioned method for accessing a large language model when executed by a processor.
[0040] In a fifth aspect, the present application provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the aforementioned method for accessing a large language model.
[0041] From the above scheme, it can be seen that the present application provides a method for accessing a large language model, including: when receiving an access request from a target caller, determining the required function information corresponding to the target caller; obtaining cluster information of each service cluster used to provide large language model services, wherein the cluster information includes performance information and concurrency support information; among each of the service clusters, selecting a target service cluster whose performance information matches the required function information and whose concurrency support information matches the number of access requests; forwarding the access request to the target service cluster, and when the target service cluster returns an access result for the access request, forwarding the access result to the target caller.
[0042] It can be seen that the beneficial effects of the present application are: when receiving an access request from a target caller, the required functional information corresponding to the target caller is determined, and the cluster information of each service cluster providing large language model services is obtained, and then, among each service cluster, a target service cluster is selected whose performance information matches the required functional information and whose concurrent support information matches the number of access requests, and the access request is forwarded to the target service cluster for processing and the corresponding access result is forwarded to the target caller. In this way, the target service cluster that processes the access request is a cluster whose performance information matches the required functional information and whose concurrent support information matches the number of access requests. The service is relatively stable, and while meeting the functional requirements, it can ensure the stability of using a large language model.
[0043] Correspondingly, the electronic device and readable storage medium provided by this application also have the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0045] Figure 1 A flow chart of a method for accessing a large language model provided in an embodiment of the present application;
[0046] Figure 2 A schematic diagram of a large language model proxy solution provided in an embodiment of the present application;
[0047] Figure 3 A cluster determination flow chart provided in an embodiment of the present application;
[0048] Figure 4A schematic diagram of the structure of a large language model access device provided in an embodiment of the present application;
[0049] Figure 5 A structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0051] Currently, business services directly connect to self-built clusters or APIs (Application Programming Interfaces) from various third-party cloud vendors to meet their inference needs. However, these services present the following challenges: Lack of stability: Third-party API services suffer from certain stability issues; Complex integration: Parameter specifications vary across platforms (e.g., max_token (maximum token) limits, differences between non-streaming and streaming output, and model names), resulting in varying usage by callers and the need to maintain multiple sets of call code; High cost of independently deploying clusters: Deploying the required GPUs (Graphics Processing Units) is expensive, and fixed node deployments lead to wasted computing power during idle times; Complex management: Different businesses may need to call multiple API services, requiring multiple accounts. As more and more businesses connect, account management and billing become complex issues; Functionality limitations: Support for online search and the maximum token context length are fixed and difficult to adjust dynamically, leaving businesses to decide which API service to call. Therefore, this application proposes a large language model proxy solution to address these issues.
[0052] See also Figure 1 As shown, the embodiment of the present application discloses a method for accessing a large language model, including:
[0053] Step S11: When an access request from a target caller is received, required function information corresponding to the target caller is determined.
[0054] When a concurrent or non-concurrent access request from a target caller is received, the required function information corresponding to the target caller can be determined. Concurrency can be understood as the number of access requests sent simultaneously being greater than or equal to 2.
[0055] The large language model access method disclosed in this application is applied to a proxy service. The proxy service provides a unified interface format, enabling callers to access a multi-party large language model cluster using a single format, i.e., a service cluster providing large language model services. Access requests can be HTTP (Hypertext Transfer Protocol) requests initiated by the business service (i.e., the caller) to the proxy service. Business services can locate the proxy service and initiate requests using either HTTP domain names or microservice command words. Both office network and service intranet domain names are provided.
[0056] In an optional implementation, determining the required functional information corresponding to the target caller includes: obtaining configuration information corresponding to the caller information based on the caller information carried in the access request, wherein the configuration information includes the required functional information.
[0057] The caller information can be information representing the caller's identity, such as an account number. The required functional information includes functional requirements for large language models, such as whether online search is supported, whether context preset data size is supported, such as 64KB context (i.e., 64KB of questions and answers combined), whether only private, independently deployed clusters are requested, and whether the completion interface is supported.
[0058] In an optional embodiment, obtaining configuration information corresponding to the caller information based on the caller information carried in the access request includes: obtaining the configuration information corresponding to the caller information from the memory based on the caller information carried in the access request; if the configuration information corresponding to the caller information does not exist in the memory, obtaining the configuration information corresponding to the caller information from the configuration center based on the caller information carried in the access request, and storing the configuration information corresponding to the caller information in the memory.
[0059] In addition, you can set an expiration time for configuration information in memory. When the expiration time is reached, the corresponding configuration information is deleted from memory. The configuration center can be implemented based on MySQL (a relational database management system) and Redis (a remote dictionary server). After the proxy service obtains the configuration from the configuration center, it caches it in memory and sets an expiration time. The next time the configuration information is read, it will be read from memory first to improve service performance. If the cache does not match, it will be pulled from the configuration center again.
[0060] Furthermore, in an optional embodiment, the configuration information also includes authentication information; after obtaining the configuration information corresponding to the caller information based on the caller information carried in the access request, the process further includes: authenticating the access request based on the authentication information. If authentication succeeds, subsequent steps are triggered; otherwise, no subsequent processing is performed. In this embodiment of the present application, upon obtaining the configuration information, authentication is first performed. Authentication can be used to determine whether the requester, i.e., the caller, is from a legitimate source.
[0061] In an optional implementation, the configuration information also includes a request quota limit; before obtaining the cluster information of each service cluster used to provide large language model services, it also includes: determining whether the current number of requests corresponding to the caller information exceeds the request quota limit; if the current number of requests corresponding to the caller information does not exceed the request quota limit, triggering the step of obtaining the cluster information of each service cluster used to provide large language model services.
[0062] Among them, the request quota limit is the upper limit of the number of concurrent requests. In the embodiment of the present application, the number of concurrent requests of the caller can be limited. If the current request corresponding to the caller information exceeds the request quota limit, the access request can be temporarily stored in a preset queue. If the current number of requests corresponding to the caller information does not exceed the request quota limit, the step of obtaining cluster information of each service cluster used to provide large language model services is triggered. Alternatively, after obtaining the cluster information of each service cluster, the access request is temporarily stored in the preset queue for waiting.
[0063] In an optional embodiment, before determining whether the current number of requests corresponding to the caller information exceeds the request quota limit, it also includes: obtaining the request information corresponding to the caller information from the target cache, wherein the request information is the information corresponding to the access request added to the target cache after determining the target service cluster corresponding to any access request; and counting the current number of requests corresponding to the caller information based on the request information.
[0064] The request information includes caller information, such as an account number, and the target cache may be a distributed cache. Embodiments of the present application can count the currently processed request information of the same caller in the target cache to obtain the current number of requests corresponding to the caller. Furthermore, in embodiments of the present application, when the inference result of any access request is returned to the caller, the request information corresponding to the access request is deleted from the target cache.
[0065] Step S12: obtaining cluster information of each service cluster for providing large language model services, wherein the cluster information includes performance information and concurrency support information.
[0066] The concurrent support information may include the upper limit of the number of requests supported by the service cluster and the current concurrent number.
[0067] The service clusters used to provide large language model services, known as large language model clusters, can provide large language model inference services, including but not limited to large AI models. A cluster is a set of services used to directly provide model inference. The underlying layer may be one or more GPU machine nodes, and can be either an API service provided by a cloud provider or a self-deployed API service.
[0068] In an optional implementation, cluster information for each service cluster can be obtained from a distributed cache. The current concurrency count can be calculated based on request information. That is, request information can be stored by cluster dimension.
[0069] Step S13: From each of the service clusters, a target service cluster is selected whose performance information matches the required function information and whose concurrency support information matches the number of access requests. That is, the target service cluster is the service cluster that meets the preset remaining concurrency sufficient condition and the required function information.
[0070] In an optional implementation, the required function information is the first function tag corresponding to the required function; the performance information in the cluster information can be the second function tag; the target service cluster is a service cluster whose second function tag is consistent with the first function tag and meets the preset remaining concurrency sufficient condition.
[0071] In an optional implementation, the major language model clusters include an independently deployed cluster and a third-party cloud vendor cluster; in each of the service clusters, a target service cluster is selected whose performance information matches the required functional information and whose concurrent support information matches the number of access requests, including: in each of the service clusters, a candidate service cluster is selected whose performance information matches the required functional information and whose concurrent support information matches the number of access requests; when the number of candidate service clusters is two or more, a target service cluster is filtered out from the candidate service clusters based on the weights of the candidate service clusters; wherein the weight of the independently deployed cluster is greater than the weight of the third-party cloud vendor cluster.
[0072] In an optional implementation, the target service cluster may be determined from the major language model clusters based on the parameter information carried in the request, the configuration information, and the cluster information. For example, the required function information may be specified by the parameter information.
[0073] In this way, the weight of the independent deployment cluster is greater than the weight of the third-party cloud vendor cluster, and requests can be dispatched to the more stable independent deployment cluster with a higher probability, thereby further ensuring the stability of the service.
[0074] This embodiment can determine weights based on the upper limit of requests supported by the cluster. The weights can be obtained by multiplying the upper limit by a coefficient. Furthermore, based on historical request data, busy and idle periods can be determined, and different upper limits on requests can be configured to schedule requests, ensuring service stability and rational resource utilization.
[0075] System administrators can adjust weights based on actual needs. A higher weight can be assigned to a private, independently deployed cluster, increasing the probability that requests will be dispatched to it, improving its utilization and preventing idleness. For privately deployed clusters, the weight is calculated by multiplying the concurrency limit by a coefficient. The coefficient can be adjusted based on the situation. For example, if the limit is 100 and the coefficient is 100, the weight is 100 * 100 = 10,000. A private cluster's coefficient might be 1, for example, if the concurrency is 10, the weight is 10 * 1 = 10. This means that weights are not fixed and are adjusted to ensure that the weight of a private cluster is several orders of magnitude higher than that of a third-party cluster. If the concurrency of a private cluster is low, the coefficient is increased to ensure that the final weight (concurrency * coefficient) is higher than that of the third-party cluster.
[0076] Step S14: forwarding the access request to the target service cluster, and when the target service cluster returns an access result for the access request, forwarding the access result to the target caller.
[0077] That is, the embodiment of the present application can dispatch the access request to the target service cluster, obtain the model inference result based on the target service cluster's response to the access request, and forward the access result to the target caller based on the request method of the access request.
[0078] For non-streaming requests, the complete access result is returned to the caller. For streaming requests, after receiving the chunked data streamed back, the chunked data is pushed sequentially to the caller. The request mode is determined based on the stream parameter in the request parameters. A true value indicates a streaming call; otherwise, it is a non-streaming call.
[0079] Furthermore, in an optional embodiment, after forwarding the access result to the target caller, the method further includes: writing the call information corresponding to the access request into a message queue so that a consumer can consume the call information from the message queue to obtain resource usage information. The call information is information that can represent resource usage information. In the embodiment of the present application, the resource usage information of the caller can be obtained by counting the call information.
[0080] It can be seen that when the embodiment of the present application receives an access request from the target caller, it determines the required functional information corresponding to the target caller, obtains the cluster information of each service cluster that provides large language model services, and then selects a target service cluster from each service cluster whose performance information matches the required functional information and whose concurrent support information matches the number of access requests, forwards the access request to the target service cluster for processing and forwards the corresponding access result to the target caller. In this way, the target service cluster that processes the access request is a cluster whose performance information matches the required functional information and whose concurrent support information matches the number of access requests. The service is relatively stable and can ensure the stability of using the large language model while meeting the functional requirements.
[0081] For further information, see Figure 2 As shown, Figure 2 A schematic diagram of a large language model proxy solution provided for an embodiment of the present application. The embodiment of the present application provides a proxy service, and the large language model proxy solution is applied to the proxy service. The business request first reaches the proxy service. The proxy service selects an optimal cluster based on the model, function and other requirements of the request and the concurrent use of the underlying large language model cluster (that is, the service cluster used to provide large language model services), schedules the request to this cluster, and then forwards the final result to the business service. Large language model clusters include but are not limited to artificial intelligence large model clusters, such as DeepSeek clusters, Tencent Cloud clusters, etc. The underlying artificial intelligence large model cluster includes independently deployed self-built clusters and API clusters of third-party cloud vendors. An example of a self-built cluster deployment is: Eight H20 GPU machines are used for hardware, and a single-machine model is deployed, meaning a large language model is deployed on one GPU machine, forming a self-built cluster. The deployment engine is SGLang, version v0.4.4, and the KV Cache (key-value cache) algorithm is RadixAttention (a radix tree-based attention algorithm). For large language models, depending on business needs, such as DeepSeek, the parameter precision is FP8 community open source version or INT8 quantized version. The processing flow for a single request may include:
[0082] 1. The business service initiates an HTTP request to the proxy service. The business service can find the proxy service and initiate the request using either the HTTP domain name or the microservice command. Because business services may be located in different network environments, two domain names are provided: the office network and the service intranet.
[0083] 2. The proxy service obtains configuration information. Configuration information such as business account authentication, quota limits, and function tags is stored in the configuration center, which is implemented based on MySQL and Redis. After the proxy service obtains the configuration information from the configuration center, it caches it in memory and sets an expiration time. The next time, the configuration information will be read from the cache first to improve service performance. If the cache is not hit, it will be pulled from the configuration center again. The configuration center can also be implemented based on other distributed components, such as etcd, zookeepr, memcached, etc. The selection can be based on high availability or high performance, and the combination of operation and maintenance capabilities must also be considered.
[0084] 3. The proxy service queries the distributed cache for concurrency information across all clusters. Because each cluster has a maximum limit on the number of requests it can support simultaneously (concurrency), it's necessary to check whether any cluster has remaining concurrency quotas. Each account also has a concurrency limit to prevent certain businesses from arbitrarily occupying system resources and impacting other businesses. Request information is stored in the distributed cache, Redis, to improve query efficiency. Information such as each request's ID (identifier), account, request time, and model are stored by cluster. Furthermore, the distributed cache's data structure allows cluster and account request records to be stored separately. Updates require atomic updates to both the corresponding cluster and account data simultaneously. Furthermore, the storage of cached request records can be replaced with other distributed cache components, such as Redis, Memcached, and etcd. Request details for each cluster and account are stored, and the request ID is guaranteed to be globally unique. This ensures idempotent writes and prevents inaccurate concurrency counts.
[0085] 4. The proxy service dispatches the request. Based on the above configuration information and concurrency information, it counts whether the concurrency of the business account has reached the upper limit, and whether there is a cluster that can meet this request in terms of functions and concurrency resources. After the proxy service determines the cluster to process this request, it records the request information in the distributed cache, and then initiates an inference request to this cluster to obtain the inference result. Figure 3 As shown, Figure 3A cluster determination flow chart provided for an embodiment of the present application. According to the data in the distributed cache, all request records on each cluster are counted to obtain the current concurrency of each cluster. Since each request contains account information, the number of concurrent requests of the business account can be counted at the same time; the account's concurrency upper limit configuration is compared with the current concurrency of the account obtained by statistics to determine whether the business has reached its available concurrency limit and whether the request can continue. Compare the concurrency upper limit configuration of each cluster with the current concurrency of the cluster obtained by statistics to determine whether there is an idle cluster. Further filter the available clusters according to the functional label of the account to obtain a list of available clusters, and randomly select a cluster according to the weight configuration of the cluster. In other embodiments, the request scheduling algorithm can be implemented according to other strategies, such as cluster priority ranking, load status calculation, etc.
[0086] 5. The proxy service returns the inference results. For non-streaming requests, the complete inference results are returned to the requester (i.e., the caller) at this point. For streaming requests, after receiving the cluster block data streaming response in step 4 above, the streaming block data is pushed sequentially to the requester.
[0087] 6. The proxy service releases the request. After the request is completed, the corresponding record must be deleted from the distributed cache, freeing up the resources occupied by the request so that subsequent requests can allocate resources normally. Prometheus monitoring metrics in the proxy service are also updated.
[0088] 7. The proxy service writes the call information to the message queue. This is used to calculate resource usage for the business service and subsequently generate billing statements. The message queue uses Kafka components. The call information may include consumed prompt tokens (prompt tokens), completion tokens (completion tokens), account number, requested model, globally unique request ID, whether the request is streaming, and whether search is enabled.
[0089] 8. Consumers consume flow data from the message queue.
[0090] 9. The consumer transfers the transaction data to the database. Based on the request transaction data, the resource usage of each business can be calculated and counted.
[0091] 10. Grafana regularly queries the monitoring metrics of the proxy service through the configured Prometheus data source. The query results of the metrics are displayed on the monitoring dashboard, allowing administrators to observe the operation of the system. Prometheus metrics can include the total number of requests received after the service is started; the first word of the streaming request: the interval from initiating the request to the underlying cluster to receiving the first streaming response packet; request information: account number, time, success, and reason for failure; the current concurrency of each cluster, that is, the number of requests currently being processed. Building a quasi-real-time monitoring system with Prometheus and Grafana can help administrators observe the operation of the entire system.
[0092] This solution primarily serves services that require API access to large language model inference capabilities. Users must first apply for an API account for permission verification, request scheduling, and cost statistics. Accounts are managed on the configuration platform. To be compatible with the OpenAI SDK (Software Development Kit), authentication uses a standardized HTTP header (Bearer) token authentication method. The interface protocol uses the open-source OpenAI interface format. Different addressing methods are provided for different network environments: domain names and microservice command words. Each account can be configured in multiple ways to meet the cost and functionality requirements of the business. Automatic switching is enabled by default to ensure request stability. Exclusive resource allocation is allowed to avoid scheduling to other resource-constrained or performance-constrained clusters. Functional tags include whether large context tokens, such as 64KB, are required, and whether network search is required. The concurrency limit can be adjusted during off-peak hours, dynamically adjusting the business's concurrency limit based on busy and busy times. During busy periods, when there are many users, concurrency needs to be reduced to prevent increased resource contention and ensure the quality of business requests.
[0093] This embodiment of the application features a proxy-layer-based system architecture: It builds a cross-platform abstraction layer, using standardized OpenAI protocol-compatible interfaces to shield access differences between self-built GPU clusters (such as NVIDIA H20) and third-party APIs (Tencent Cloud / Volcano Engine), supporting zero-code migration for callers. A heterogeneous resource fusion scheduling algorithm supports calls from both self-built clusters and third-party APIs, automatically selecting the appropriate cluster to handle requests based on cluster load, user functionality, and cost requirements. Multi-threaded automatic switching provides stable inference services. A low-cost deployment solution for self-built clusters: DeepSeek is deployed using the SGlang inference engine, and the RadixAttention algorithm is used to effectively reduce KV Cache memory usage. A single-machine deployment utilizes eight NVIDIA H20 GPUs. An intermediate proxy layer decouples callers from the cluster, preventing them from directly requesting the model cluster. This leaves room for resource scheduling, feature upgrades, account management, and billing models. Self-built clusters optimize performance and cost at the deployment level. The inference engine uses SGLang, and the model is available in versions with FP8 parameter precision (full-capability) and INT8 quantization. This meets the varying requirements of callers for model inference performance and effectiveness, reducing the cost of building self-built clusters. By integrating heterogeneous computing resources and innovative scheduling mechanisms, we achieve coordinated optimization of service stability, cost-effectiveness, and functional scalability.
[0094] In this way, the embodiment of the present application can improve service stability, and through multi-path redundancy and intelligent switching routes, it can ensure the caller's demand for stable large language model services. Since each cluster has a concurrency limit, some clusters have reached the limit and cannot process new requests. The embodiment of the present application will select healthy clusters with remaining concurrency to process, rather than offline, unavailable clusters, or clusters with no remaining concurrency; it can reduce overall costs and optimize resource efficiency, and through intelligent scheduling algorithms, reasonably combine clusters of different costs: for online businesses with high stability requirements, they are scheduled to independent clusters with high costs but good stability; for offline businesses with low stability requirements, they are scheduled to independent clusters during idle time, and to cloud vendor clusters that are relatively unstable but low in cost and high in concurrency during busy time, so as to make full use of resources. At the same time, considering the stability demands of other businesses, for self-built clusters, the SGLang inference engine and RadixAttention's KV are combined in model deployment. Cache caching technology further optimizes GPU costs, and by adjusting the weight of the cluster, prioritizes scheduling to independently deployed clusters to improve cluster resource utilization; it breaks through the limitations of third-party functional fixation and has great scalability to meet the needs of different business scenarios. The proxy layer can increase support for multimodality, RAG, contexts of different scales (such as 64K Tokens), other models (such as DeepSeek-V3) and other functions; it can meet both general business scenarios and customized needs in vertical fields such as music, novels, and podcasts. For example, if a 64K context label is attached to the account, then when scheduling and selecting a cluster, a cluster that supports 64K will be selected. The cluster is also labeled and matched by the label. In addition to configuring the account, you can also directly specify which functions are required in the request parameters, such as enabling search capabilities. The request parameters will also be combined during scheduling to match the cluster; the complexity of business access is greatly reduced: through a standardized interface layer, it is compatible with OpenAI The SDK's API protocol allows businesses to migrate existing applications with zero code, shielding the diversity and functional differences of the underlying clusters. This solves the code and interface fragmentation problems caused by switching between different cloud vendors. Use OpenAI's API protocol to make requests, replacing the base_url with the proxy service address. To enable different functions, simply modify the configuration items or request parameters without reconnecting to the new vendor. For account tags, the administrator can modify the tags in the configuration platform. For request parameters, simply modify the request parameter values, such as whether to enable search. For subsequent new features, businesses can quickly experience them without changing the access method. There is no need to directly connect to the underlying cluster, ensuring compatibility with heterogeneous hardware. This supports deployment of various GPU clusters and reduces the risk of dependence on individual GPU hardware.
[0095] That is, to address the lack of stability: the embodiments of the present application support access to multiple API services, and through multi-path redundancy and automatic line switching, ensure that stable services are provided to the caller; to address the complexity of access: the embodiments of the present application provide a unified OpenAI interface format, and the caller can access the reasoning services of multiple parties using one format; to address the high cost of independently deployed clusters: the embodiments of the present application use the SGLang inference engine for deployment, and models with different parameter accuracies, and by adjusting the weights of the clusters, give priority to scheduling to independently deployed clusters, thereby improving the resource utilization of the cluster and reducing the overall cluster cost; to address the complexity of management: the embodiments of the present application only need to assign an access account to each business party by building a configuration center, and there is no need to assign accounts to the underlying clusters; to address functional limitations: the embodiments of the present application access multiple clusters, and select the appropriate cluster to process requests based on the caller's configuration information, model and functional requirements; by adjusting the underlying clusters, the capabilities of the entire system are expanded without the business party's awareness.
[0096] For further information, see Figure 4 As shown, an embodiment of the present application provides a device for accessing a large language model, including:
[0097] The function requirement information determination module 11 is configured to determine the function requirement information corresponding to the target caller upon receiving an access request from the target caller;
[0098] A cluster information acquisition module 12 is used to obtain cluster information of each service cluster used to provide large language model services, wherein the cluster information includes performance information and concurrency support information;
[0099] A target service cluster selection module 13 is configured to select, from each of the service clusters, a target service cluster whose performance information matches the required function information and whose concurrent support information matches the number of access requests;
[0100] The access request processing module 14 is configured to forward the access request to the target service cluster, and when the target service cluster returns an access result for the access request, forward the access result to the target caller.
[0101] In an optional implementation, the major language model clusters include independently deployed clusters and third-party cloud vendor clusters; the target service cluster selection module 13 is specifically used to: select a candidate service cluster from each of the service clusters whose performance information matches the required function information and whose concurrent support information matches the number of access requests; when the number of candidate service clusters is two or more, filter out the target service cluster from the candidate service clusters based on the weights of the candidate service clusters; wherein the weight of the independently deployed cluster is greater than the weight of the third-party cloud vendor cluster.
[0102] In an optional implementation manner, the required function information determining module 11 is specifically configured to: obtain configuration information corresponding to the caller information based on the caller information carried in the access request, where the configuration information includes the required function information.
[0103] The demand function information determination module 11 is specifically used to: obtain the configuration information corresponding to the caller information from the memory based on the caller information carried in the access request; if the configuration information corresponding to the caller information does not exist in the memory, obtain the configuration information corresponding to the caller information from the configuration center based on the caller information carried in the access request, and store the configuration information corresponding to the caller information into the memory.
[0104] In an optional embodiment, the configuration information also includes a request quota limit, and the device is further used for a request quota limit judgment module, which is used to judge whether the current number of requests corresponding to the caller information exceeds the request quota limit before obtaining the cluster information of each service cluster used to provide large language model services; if the current number of requests corresponding to the caller information does not exceed the request quota limit, the step of obtaining the cluster information of each service cluster used to provide large language model services is triggered.
[0105] In an optional embodiment, the request quota limit judgment module is also used to obtain the request information corresponding to the caller information from the target cache before determining whether the current number of requests corresponding to the caller information exceeds the request quota limit, wherein the request information is the information corresponding to the access request added to the target cache after determining the target service cluster corresponding to any access request; and the current number of requests corresponding to the caller information is counted based on the request information.
[0106] In an optional implementation, the configuration information further includes authentication information; and the apparatus is further configured to authenticate the access request based on the authentication information.
[0107] In an optional embodiment, the device is further configured to write the call information corresponding to the access request into a message queue after forwarding the access result to the target caller, so that a consumer can consume the call information from the message queue to obtain resource usage.
[0108] See also Figure 5 As shown, an embodiment of the present application discloses an electronic device 20, including a processor 21 and a memory 22; wherein the memory 22 is used to store a computer program; the processor 21 is used to execute the computer program, and the access method of the large language model disclosed in the above embodiment.
[0109] For the specific process of the method for accessing the large language model, please refer to the corresponding content disclosed in the aforementioned embodiment, which will not be repeated here.
[0110] Furthermore, the memory 22 as a carrier for resource storage may be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the storage method may be temporary storage or permanent storage.
[0111] In addition, the electronic device 20 also includes a power supply 23, a communication interface 24, an input / output interface 25 and a communication bus 26; wherein the power supply 23 is used to provide an operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and an external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input / output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0112] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program, wherein when the computer program is executed by a processor, the method for accessing the large language model disclosed in the aforementioned embodiment is implemented.
[0113] For the specific process of the method for accessing the large language model, please refer to the corresponding content disclosed in the aforementioned embodiment, which will not be repeated here.
[0114] Furthermore, an embodiment of the present application also discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the method for accessing a large language model disclosed in the aforementioned embodiment.
[0115] For the specific process of the method for accessing the large language model, please refer to the corresponding content disclosed in the aforementioned embodiment, which will not be repeated here.
[0116] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0117] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0118] The above is a detailed introduction to the access method, device and medium for a large language model provided by this application. Specific examples are used in this article to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core idea of this application. At the same time, for general technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.
Claims
1. A method for accessing a large language model, characterized in that: include: When receiving an access request from a target caller, determining required function information corresponding to the target caller; Obtaining cluster information of each service cluster for providing a large language model service, wherein the cluster information includes performance information and concurrency support information; Selecting, from among the service clusters, a target service cluster whose performance information matches the required function information and whose concurrent support information matches the number of access requests; The access request is forwarded to the target service cluster, and when the target service cluster returns an access result for the access request, the access result is forwarded to the target caller.
2. The method for accessing a large language model according to claim 1, characterized in that: The major language model clusters include independently deployed clusters and third-party cloud vendor clusters; Selecting, from each of the service clusters, a target service cluster whose performance information matches the required function information and whose concurrent support information matches the number of access requests, includes: Selecting, from among the service clusters, a candidate service cluster whose performance information matches the required function information and whose concurrent support information matches the number of access requests; When the number of candidate service clusters is more than two, selecting a target service cluster from the candidate service clusters based on the weights of the candidate service clusters; Among them, the weight of the independently deployed cluster is greater than the weight of the third-party cloud vendor cluster.
3. The method for accessing a large language model according to claim 1, wherein: Determine the required functional information corresponding to the target caller, including: Configuration information corresponding to the caller information is acquired based on the caller information carried in the access request, where the configuration information includes required function information.
4. The method for accessing a large language model according to claim 3, wherein: Acquiring configuration information corresponding to the caller information based on the caller information carried in the access request, including: Obtaining configuration information corresponding to the caller information from a memory based on the caller information carried in the access request; If the configuration information corresponding to the caller information does not exist in the memory, the configuration information corresponding to the caller information is obtained from the configuration center based on the caller information carried in the access request, and the configuration information corresponding to the caller information is stored in the memory.
5. The method for accessing a large language model according to claim 3, wherein: The configuration information also includes a request quota limit; Before obtaining the cluster information of each service cluster used to provide the large language model service, the following steps are also included: Determine whether the current number of requests corresponding to the caller information exceeds the request quota limit; When the current number of requests corresponding to the caller information does not exceed the request quota limit, the step of obtaining cluster information of each service cluster for providing large language model services is triggered.
6. The method for accessing a large language model according to claim 5, characterized in that: Before determining whether the current number of requests corresponding to the caller information exceeds the request quota limit, the method further includes: Obtaining request information corresponding to the caller information from the target cache, wherein the request information is information corresponding to the access request added to the target cache after determining the target service cluster corresponding to any access request; The current number of requests corresponding to the caller information is counted based on the request information.
7. The method for accessing a large language model according to claim 3, wherein: The configuration information also includes authentication information; After obtaining configuration information corresponding to the caller information based on the caller information carried in the access request, the method further includes: The access request is authenticated based on the authentication information.
8. The method for accessing a large language model according to claim 7, wherein: After forwarding the access result to the target caller, the method further includes: The call information corresponding to the access request is written into a message queue so that a consumer can consume the call information from the message queue and obtain resource usage information.
9. An electronic device, characterized in that: comprising a memory and a processor, wherein: The memory is used to store computer programs; The processor is configured to execute the computer program to implement the method for accessing a large language model according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that Used to store a computer program, wherein when the computer program is executed by a processor, the method for accessing a large language model according to any one of claims 1 to 8 is implemented.