AI large model calling load balancing method, device, equipment and medium

By acquiring information about model vendors and load balancing patterns, the target large AI model is identified, solving the problem of load balancing in existing large AI model calling schemes and improving efficiency and stability.

CN120803720APending Publication Date: 2025-10-17SUZHOU DAJIAYING INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510922717.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The existing AI large model calling solution cannot effectively utilize the AI ​​large models provided by multiple different model suppliers for load balancing, resulting in difficulty in ensuring the efficiency and stability of the calling process.

Method used

By acquiring model application information from various model vendors' large AI models, including model IDs and traffic limit information, and based on key load balancing parameters and available load balancing modes, the current load balancing mode is determined. Upon detecting a problematic request, the target large AI model is selected from among the large AI models provided by each vendor for processing, according to the current load balancing mode. Available load balancing modes include round-robin balancing, traffic balancing, traffic weight balancing, and traffic ratio balancing.

Benefits of technology

It achieves load balancing of the AI ​​large model calling process, improves the efficiency and stability of the calling process, and ensures that problem requests can be diverted to the AI ​​large models provided by various model suppliers for processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803720A_ABST
    Figure CN120803720A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric digital data processing, in particular to an AI large model calling load balancing method, device and equipment and a medium. The method comprises the steps of obtaining model application information of an AI large model provided by each model supplier; determining a current load balancing mode according to the key load balancing parameters and selectable load balancing modes; the selectable load balancing modes comprise a polling balancing mode, a flow balancing mode, a flow weight comprehensive balancing mode and a flow proportion comprehensive balancing mode; when a problem request is detected, determining a target AI large model corresponding to the problem request in the AI large models provided by the model suppliers according to the current load balancing mode and the model application information; and calling the target AI large model to process the question request to obtain an answer text. According to the embodiment of the invention, load balancing can be carried out on the calling process of the AI large models, and the problem requests are shunted to the AI large models provided by all model suppliers to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electric digital data processing, and particularly relates to an AI large model calling load balancing method and device, equipment and a medium. BACKGROUND

[0002] With the development of artificial intelligence (AI) technology, more and more enterprises begin to call AI large models of business servers provided by model suppliers to process obtained question requests of users and obtain answer texts corresponding to the question requests.

[0003] In the related art, a commonly used AI large model calling scheme is as follows: an AI large model is selected from AI large models provided by multiple different model suppliers as an AI large model used for processing question requests obtained by a business server. After obtaining a question request each time, the AI large model used for processing question requests obtained by the business server is uniformly called to process the question request and obtain an answer text corresponding to the question request. The AI large model calling scheme in the related art calls an AI large model provided by a fixed model supplier to process a question request, and cannot use AI large models provided by multiple different model suppliers to balance the AI large model calling process and distribute question requests to AI large models provided by each model supplier for processing, so that the efficiency and stability of the AI large model calling process are difficult to guarantee. SUMMARY

[0004] The present application provides an AI large model calling load balancing method, device, equipment and medium to solve the problem that the AI large model calling scheme in the related art cannot use AI large models provided by multiple different model suppliers to balance the AI large model calling process and distribute question requests to AI large models provided by each model supplier for processing, so that the efficiency and stability of the AI large model calling process are difficult to guarantee.

[0005] According to an aspect of the present application, an AI large model calling load balancing method is provided, comprising:

[0006] obtaining model application information of AI large models provided by each model supplier; wherein the model application information comprises a model number and flow limitation information;

[0007] determining a current load balancing mode according to a key load balancing parameter and a selectable load balancing mode; wherein the selectable load balancing mode comprises a round robin balancing mode, a flow balancing mode, a flow weight comprehensive balancing mode and a flow proportion comprehensive balancing mode;

[0008] determine, when detecting a question request, a target AI large model corresponding to the question request from AI large models provided by each model provider according to the current load balancing mode and model application information of the AI large models provided by each model provider;

[0009] invoke the target AI large model to process the question request to obtain an answer text corresponding to the question request.

[0010] According to another aspect of the present application, an AI large model invocation load balancing device is provided, comprising:

[0011] an information acquisition module configured to acquire model application information of AI large models provided by each model provider; wherein the model application information comprises a model number and traffic limit information;

[0012] a mode determination module configured to determine a current load balancing mode according to a key load balancing parameter and a selectable load balancing mode; wherein the selectable load balancing mode comprises a round robin balancing mode, a traffic balancing mode, a traffic weight comprehensive balancing mode and a traffic proportion comprehensive balancing mode;

[0013] a model determination module configured to determine, when detecting a question request, a target AI large model corresponding to the question request from AI large models provided by each model provider according to the current load balancing mode and model application information of the AI large models provided by each model provider;

[0014] a model invocation module configured to invoke the target AI large model to process the question request to obtain an answer text corresponding to the question request.

[0015] According to another aspect of the present application, an electronic device is provided, comprising:

[0016] at least one processor;

[0017] and a memory in communication connection with the at least one processor;

[0018] wherein the memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the AI large model invocation load balancing method according to any one of the embodiments of the present application.

[0019] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to execute the AI large model invocation load balancing method according to any one of the embodiments of the present application.

[0020] According to another aspect of the present application, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the AI large model calling load balancing method according to any of the embodiments of the present application.

[0021] The technical solution of the embodiments of the present application acquires the model application information of the AI large models provided by each model provider, the model application information comprising a model number and traffic limit information; then determines the current load balancing mode according to the key load balancing parameters and the optional load balancing modes, the optional load balancing modes comprising: a round robin balancing mode, a traffic balancing mode, a traffic weight comprehensive balancing mode and a traffic proportion comprehensive balancing mode; when a problem request is detected, determines the target AI large model corresponding to the problem request from the AI large models provided by each model provider according to the current load balancing mode and the model application information of the AI large models provided by each model provider, calls the target AI large model to process the problem request, and obtains the answer text corresponding to the problem request, solving the problem that the AI large model calling solution in the related art cannot utilize the AI large models provided by multiple different model providers to balance the AI large model calling process, and divert the problem request to the AI large models provided by each model provider for processing, and the efficiency and stability of the AI large model calling process are difficult to guarantee, and the AI large model calling process can be balanced based on the model application information of the AI large models provided by each model provider, the problem request is diverted to the AI large models provided by each model provider for processing, and the efficiency and stability of the AI large model calling process are improved.

[0022] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0024] Figure 1 A flowchart of an AI large model calling load balancing method provided for the first embodiment of the present application.

[0025] Figure 2 A flowchart of an AI large model calling load balancing method provided for the second embodiment of the present application.

[0026] Figure 3A flow chart of an AI large model calling load balancing method provided for the third embodiment of the present application.

[0027] Figure 4 A flow chart of an AI large model calling load balancing method provided for the fourth embodiment of the present application.

[0028] Figure 5 A flow chart of an AI large model calling load balancing method provided for the fifth embodiment of the present application.

[0029] Figure 6 A structural schematic diagram of an AI large model calling load balancing device provided for the sixth embodiment of the present application.

[0030] Figure 7 A structural schematic diagram of an electronic device for implementing the AI large model calling load balancing method of the embodiments of the present application. DETAILED DESCRIPTION

[0031] In order to make the personnel in the technical field better understand the present application scheme, the technical scheme in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application. It should be noted that the terms "target", "first", "second" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include", "contain" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device containing a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] Embodiment one

[0033] Figure 1A flowchart of an AI large model calling load balancing method provided for Embodiment One of the present application. This embodiment can be applicable to the case of load balancing of the AI large model calling process in the scenario of calling the AI large model to process the problem request. The method can be executed by an AI large model calling load balancing device, which can be realized in the form of hardware and / or software, and can be configured in an electronic device, such as a business server set in an enterprise, which is used in cooperation with the client sending the problem request. The business server can be a server for processing the business of the enterprise. As shown in Figure 1 The method comprises the following steps:

[0034] Step 101, obtaining the model application information of the AI large model provided by each model provider.

[0035] Among them, the model application information includes model number and traffic limit information.

[0036] Optionally, the AI large model can refer to a large language model (LLM) trained using artificial intelligence (AI) technology and capable of performing inference processing such as problem answering, machine translation, and text generation. Multiple different model providers provide the trained AI large model to the business server of the enterprise for calling. Each model provider can be an enterprise that provides the trained AI large model to the business server of other enterprises for calling. The AI large model provided by each model provider is usually deployed on the server used by the model provider.

[0037] Optionally, for each AI large model provided by a model provider, the model application information of the AI large model can be information related to the AI large model and the calling process of the AI large model. The model application information of the AI large model includes a model number. The model number can be a pre-set digital number for identifying the AI large model. The model application information of the AI large model also includes traffic limit information. The traffic limit information can be the maximum number of calls per hour of the AI large model. The maximum number of calls per hour of the AI large model can refer to the maximum number of calls allowed by the enterprise server within 1 hour under the condition of ensuring the normal operation of the AI large model.

[0038] Optionally, the model application information of the AI large model provided by each model provider is obtained, including: obtaining the model application information of the AI large model provided by each model provider from a preset database. The preset database can be a local database set in the business server, and can also be a database set on another server connected with the business server. The model application information of the AI large model provided by each model provider is stored in the preset database. The model application information of the AI large model provided by each model provider can be obtained from the preset database. The target user can be a technical personnel in an enterprise. The model application information of the AI large model provided by each model provider stored in the preset database can be stored in the preset database by the target user.

[0039] In step 102, a current load balancing mode is determined according to the key load balancing parameter and the optional load balancing mode.

[0040] The optional load balancing mode includes a round robin balancing mode, a traffic balancing mode, a traffic weight comprehensive balancing mode, and a traffic proportion comprehensive balancing mode.

[0041] Optionally, the key load balancing parameter can be a key parameter that needs to be referred to when load balancing is performed on the AI large model calling process. A parameter storage area is set in the business server. The parameter storage area can be a memory or a register for storing the key load balancing parameter. The target user can evaluate and confirm the key parameter that needs to be referred to when load balancing is performed on the AI large model calling process, that is, evaluate and confirm the key load balancing parameter, and specify the parameter that needs to be referred to when load balancing is performed on the AI large model calling process by writing the key load balancing parameter into the parameter storage area. The key load balancing parameter can be a model number, can also be a traffic quota, can also be traffic quota and weight information, and can also be traffic quota and current request processing proportion.

[0042] Optionally, the traffic quota can refer to the current available calling number of the AI large model. The current available calling number of the AI large model can refer to the total number of AI large models that can be called by the business server in the hour at the current time. The current available calling number of the AI large model is equal to the difference between the maximum calling number of the AI large model per hour and the total number of AI large models that have been called by the business server in the hour at the current time. A counter corresponding to each AI large model provided by each model provider is set in the business server. For each hour of 24 hours in a day, the count value in the counter corresponding to the AI large model is cleared at the beginning of the hour, and then the count value in the counter is incremented by 1 after each call of the AI large model. Reading the count value in the counter corresponding to the AI large model can obtain the total number of AI large models that have been called by the business server in the hour at the current time.

[0043] Optionally, the weight information can be a numerical value representing the high or low degree of the AI large model's ability to process the problem request. The greater the weight information of the AI large model, the higher the AI large model's ability to process the problem request. The smaller the weight information of the AI large model, the lower the AI large model's ability to process the problem request. The current request processing ratio can be the ratio of the total number of problem requests that the business server has invoked the AI large model to process to the total number of problem requests that the business server has invoked all AI large models provided by the model provider to process.

[0044] Optionally, the optional load balancing mode can refer to a plurality of modes that can be used to balance the calling process of the AI large model provided by each model provider. The optional load balancing mode includes: a round robin balancing mode, a traffic balancing mode, a traffic weight balancing mode, and a traffic proportion balancing mode. The round robin balancing mode can refer to a mode that balances the calling process of the AI large model by using a round robin method based on the model number of the AI large model provided by each model provider, and distributes the problem request to each AI large model provided by each model provider for processing. The traffic balancing mode can refer to a mode that balances the calling process of the AI large model based on the traffic quota of the AI large model provided by each model provider, and distributes the problem request to each available AI large model in the AI large model provided by each model provider for processing. The traffic weight balancing mode can refer to a mode that balances the calling process of the AI large model based on the traffic quota and the weight information of the AI large model provided by each model provider, and distributes the problem request to each available AI large model in the AI large model provided by each model provider for processing. The traffic proportion balancing mode can refer to a mode that balances the calling process of the AI large model based on the traffic quota and the current request processing ratio of the AI large model provided by each model provider, and distributes the problem request to each available AI large model in the AI large model provided by each model provider for processing.

[0045] Optionally, the current load balancing mode is the mode that is currently used to balance the calling process of the AI large model provided by each model provider. According to the key load balancing parameter and the optional load balancing mode, the current load balancing mode is determined, including: when the key load balancing parameter is the model number, the current load balancing mode is determined to be the round robin balancing mode; when the key load balancing parameter is the traffic quota, the current load balancing mode is determined to be the traffic balancing mode; when the key load balancing parameter is the traffic quota and the weight information, the current load balancing mode is determined to be the traffic weight balancing mode; when the key load balancing parameter is the traffic quota and the current request processing ratio, the current load balancing mode is determined to be the traffic proportion balancing mode.

[0046] Optionally, it can be monitored whether the key load balancing parameters in the parameter storage area are changed, and when it is monitored that the key load balancing parameters in the parameter storage area are changed, the current load balancing mode is updated in a timely manner according to the changed key load balancing parameters.

[0047] In step 103, when detecting a problem request, a target AI large model corresponding to the problem request is determined from the AI large models provided by the model vendors according to the current load balancing mode and the model application information of the AI large models provided by the model vendors.

[0048] Optionally, the problem request can be a problem text input by a user. The problem text can be a text used to describe a problem related to an enterprise that the user is concerned about. The answer text corresponding to the problem request can be a text used to describe an answer to the problem described by the problem request. The target AI large model corresponding to the problem request can be an AI large model selected from the AI large models provided by the model vendors and capable of being used to process the problem request to obtain the answer text corresponding to the problem request. It can be detected whether the business server receives the problem request, and the target AI large model corresponding to the problem request can be determined from the AI large models provided by the model vendors according to the current load balancing mode and the model application information of the AI large models provided by the model vendors each time it is detected that the business server receives the problem request.

[0049] Optionally, determining the target AI large model corresponding to the problem request from the AI large models provided by the model vendors according to the current load balancing mode and the model application information of the AI large models provided by the model vendors includes: when the current load balancing mode is a round robin balancing mode, determining a model number of an AI large model used to process the problem request this time according to a number polling sequence corresponding to the AI large models provided by the model vendors and a historical model number; wherein the historical model number is a model number of an AI large model used to process the problem request last time; and determining an AI large model having the same model number as the model number of the AI large model used to process the problem request this time from the AI large models provided by the model vendors as the target AI large model corresponding to the problem request.

[0050] Optionally, the business server stores a number polling sequence corresponding to the AI large models provided by the model vendors. The number polling sequence corresponding to the AI large models provided by the model vendors can be a sequence formed by arranging the model numbers of the AI large models provided by the model vendors.

[0051] Optionally, the historical model number corresponding to the AI large model provided by each model provider can be the model number of the AI large model used last time by the business server to process a problem request. A number storage area is provided in the business server. The number storage area can be a memory or a register for storing the historical model number corresponding to the AI large model provided by each model provider. In the initial state, no AI large model is called to process a problem request, and the number storage area is empty. After the AI large model is called for the first time to process a problem request, the model number of the AI large model called for the first time is the model number of the AI large model used last time to process a problem request, the model number of the AI large model called for the first time is determined as the historical model number corresponding to the AI large model provided by each model provider, and the model number is stored in the number storage area. After the AI large model is called for the second time to process a problem request, the model number of the AI large model called for the second time is the model number of the AI large model used last time to process a problem request, the model number of the AI large model called for the second time is determined as the historical model number corresponding to the AI large model provided by each model provider, and the model number stored in the number storage area is updated to the model number of the AI large model called for the second time. Similarly, after the AI large model is called to process a problem request each time, the model number of the called AI large model is determined as the historical model number corresponding to the AI large model provided by each model provider, and the model number stored in the number storage area is updated to the model number of the called AI large model.

[0052] Optionally, according to the historical model number and the number polling sequence corresponding to the AI large model provided by each model provider, the model number of the AI large model used this time to process a problem request is determined, which includes: reading the model number in the number storage area; if the number storage area is empty and no model number in the number storage area is read, the first model number in the number polling sequence corresponding to the AI large model provided by each model provider is determined as the model number of the AI large model used this time to process a problem request; if the number storage area is not empty, the arrangement position of the read model number in the number polling sequence corresponding to the AI large model provided by each model provider is determined, and the model number arranged after the read model number is determined as the model number of the AI large model used this time to process a problem request.

[0053] Optionally, if the arrangement position of the read model number in the number polling sequence corresponding to the AI large model provided by each model provider is not the last one, that is, there is a model number arranged after the read model number, the model number arranged after the read model number is determined as the model number of the AI large model used this time to process a problem request.

[0054] Optionally, if the read model number is the last in the polling order corresponding to the AI large model provided by each model provider, i.e. there is no model number after the read model number, the model number in the first position is determined as the model number of the AI large model used to process the problem request this time.

[0055] Optionally, after determining the model number of the AI large model used to process the problem request this time, the AI large model provided by each model provider with the same model number as the model number of the AI large model used to process the problem request this time is determined as the target AI large model corresponding to the problem request.

[0056] Therefore, based on the model numbers of the AI large models provided by each model provider, the calling process of the AI large models is load balanced in a polling manner, and the problem request is distributed to the AI large models provided by each model provider for processing.

[0057] Optionally, according to the current load balancing mode and the model application information of the AI large models provided by each model provider, the target AI large model corresponding to the problem request is determined from the AI large models provided by each model provider, including: when the current load balancing mode is the traffic balancing mode, obtaining the traffic record information of the AI large models provided by each model provider; determining the traffic quota of the AI large models provided by each model provider according to the traffic limit information and the traffic record information of the AI large models provided by each model provider; selecting each available AI large model from the AI large models provided by each model provider according to the traffic quota of the AI large models provided by each model provider; and selecting one available AI large model as the target AI large model corresponding to the problem request from the available AI large models.

[0058] Optionally, for each AI large model provided by each model provider, the traffic record information of the AI large model can be the count value in the counter corresponding to the AI large model, i.e. the total number of times the AI large model has been called by the business server in the current hour. The count value in the counter corresponding to each AI large model provided by each model provider can be read to obtain the traffic record information of the AI large model provided by each model provider. The traffic limit information of the AI large model can be the maximum number of calls per hour of the AI large model. The traffic quota of the AI large model can be the current available number of calls of the AI large model, i.e. the total number of times the AI large model can be called remaining in the hour at the current time. The difference between the traffic limit information and the traffic record information of the AI large model is calculated, i.e. the traffic quota of the AI large model is obtained.

[0059] Optionally, the traffic quota of the AI large model provided by each model provider is determined according to the traffic limit information and the traffic record information of the AI large model provided by each model provider, including: performing the following operation on each AI large model provided by each model provider: calculating the difference between the traffic limit information and the traffic record information of the AI large model to obtain the traffic quota of the AI large model.

[0060] Optionally, the available AI large model can refer to an AI large model with a traffic quota greater than 0. Each available AI large model in the AI large models provided by each model provider is screened according to the traffic quota of the AI large model provided by each model provider, including: performing the following operation on each AI large model provided by each model provider: if the traffic quota of the AI large model is greater than 0, determining that the AI large model is an available AI large model; if the traffic quota of the AI large model is less than or equal to 0, determining that the AI large model is not an available AI large model.

[0061] Optionally, one of the available AI large models is selected as the target AI large model corresponding to the problem request, including: randomly selecting one of the available AI large models as the target AI large model corresponding to the problem request.

[0062] Optionally, one of the available AI large models is selected as the target AI large model corresponding to the problem request, including: determining the maximum value in the traffic quotas of the available AI large models, and determining the available AI large model to which the maximum value belongs as the target AI large model corresponding to the problem request. If there is an available AI large model with a traffic quota equal to the maximum value, the available AI large model with the traffic quota equal to the maximum value is selected as the target AI large model corresponding to the problem request. If there are multiple available AI large models with traffic quotas equal to the maximum value, one of the available AI large models with the traffic quotas equal to the maximum value is randomly selected as the target AI large model corresponding to the problem request.

[0063] In this way, the calling process of the AI large model is load balanced based on the traffic quotas of the AI large models provided by each model provider, and the problem request is distributed to each available AI large model in the AI large models provided by each model provider for processing.

[0064] Optionally, the target AI large model corresponding to the problem request is determined from the AI large models provided by the model providers according to the current load balancing mode and the model application information of the AI large models provided by the model providers, including: when the current load balancing mode is the traffic weight balancing mode, obtaining traffic record information of the AI large models provided by the model providers; determining a traffic quota of the AI large models provided by the model providers according to the traffic limit information and the traffic record information of the AI large models provided by the model providers; screening each available AI large model from the AI large models provided by the model providers according to the traffic quota of the AI large models provided by the model providers; obtaining weight information of each available AI large model; and screening one available AI large model as the target AI large model corresponding to the problem request from the available AI large models according to the weight information of each available AI large model.

[0065] Optionally, the service server is provided with a weight storage area corresponding to each AI large model provided by the model providers. For each AI large model provided by the model providers, the weight storage area corresponding to the AI large model can be a memory or a register for storing weight information of the AI large model. The weight information of each available AI large model can be obtained by reading data in the weight storage area corresponding to each available AI large model.

[0066] Optionally, the target AI large model corresponding to the problem request is determined from the available AI large models according to the weight information of each available AI large model, including: determining a maximum value in the weight information of each available AI large model, and determining the available AI large model to which the maximum value belongs as the target AI large model corresponding to the problem request. If the weight information of one available AI large model is the maximum value, the available AI large model with the maximum value is determined as the target AI large model corresponding to the problem request. If the weight information of multiple available AI large models is the maximum value, one available AI large model is randomly selected from the multiple available AI large models with the maximum value as the target AI large model corresponding to the problem request.

[0067] Therefore, the calling process of the AI large model is balanced based on the traffic quota and the weight information of the AI large models provided by the model providers, and the problem request is distributed to each available AI large model provided by the model providers for processing.

[0068] Optionally, the target AI large model corresponding to the problem request is determined from the AI large models provided by the model providers according to the current load balancing mode and the model application information of the AI large models provided by the model providers, including: when the current load balancing mode is the traffic proportion load balancing mode, obtaining the traffic record information and the number of processed requests of the AI large models provided by the model providers; determining the total number of current requests according to the number of processed requests of the AI large models provided by the model providers; determining the traffic quota of the AI large models provided by the model providers according to the traffic limit information and the traffic record information of the AI large models provided by the model providers; screening each available AI large model from the AI large models provided by the model providers according to the traffic quota of the AI large models provided by the model providers; obtaining the standard request processing proportion of each available AI large model; determining the current request processing proportion of each available AI large model according to the total number of current requests and the number of processed requests of each available AI large model; and screening one available AI large model as the target AI large model corresponding to the problem request from the available AI large models according to the standard request processing proportion and the current request processing proportion of each available AI large model.

[0069] Optionally, the service server is provided with a request number storage area corresponding to each AI large model provided by the model providers. For each AI large model provided by the model providers, the request number storage area corresponding to the AI large model can be a memory or a register for storing the number of processed requests of the AI large model. The number of processed requests of the AI large model is the total number of problem requests that have been processed by the service server by calling the AI large model. In the initial state, the service server has not called the AI large model to process a problem request, and the number of processed requests stored in the request number storage area corresponding to the AI large model is 0. After the service server calls the AI large model to process a problem request each time, the number of processed requests stored in the request number storage area corresponding to the AI large model is incremented by 1. The number of processed requests of each AI large model provided by the model providers can be obtained by reading the data in the request number storage area corresponding to each AI large model provided by the model providers.

[0070] Optionally, the total number of current requests can refer to the total number of problem requests that have been processed by the service server by calling all AI large models provided by the model providers. The sum of the number of processed requests of each AI large model provided by the model providers is the total number of current requests. The total number of current requests is determined according to the number of processed requests of each AI large model provided by the model providers, including: summing the number of processed requests of each AI large model provided by the model providers to obtain the total number of current requests.

[0071] Optionally, the service server is provided with a processing proportion storage area corresponding to each AI large model provided by each model provider. For each AI large model provided by each model provider, the processing proportion storage area corresponding to the AI large model can be a memory or a register for storing the standard request processing proportion of the AI large model. The standard request processing proportion of the AI large model is the ratio of the total number of problem requests that the service server has invoked the AI large model to process in an ideal state to the total number of problem requests that the service server has invoked all AI large models provided by all model providers to process, according to the capability of the AI large model to process problem requests. The standard request processing proportion of each available AI large model can be obtained by reading the data in the processing proportion storage area corresponding to each available AI large model.

[0072] Optionally, the current request processing proportion of each available AI large model is determined according to the total number of current requests and the number of processed requests of each available AI large model, including: for each available AI large model, calculating the ratio of the number of processed requests of the available AI large model to the total number of current requests to obtain the current request processing proportion of the available AI large model.

[0073] Optionally, one of the available AI large models is selected as the target AI large model corresponding to the problem request according to the standard request processing proportion and the current request processing proportion of each available AI large model, including: for each available AI large model, determining whether the current request processing proportion of the available AI large model is less than the standard request processing proportion of the available AI large model; and randomly selecting one of the available AI large models whose current request processing proportion is less than the standard request processing proportion as the target AI large model corresponding to the problem request.

[0074] Therefore, the calling process of the AI large model is load balanced based on the traffic quota and the current request processing proportion of each AI large model provided by each model provider, and the problem request is distributed to each available AI large model provided by each model provider for processing.

[0075] Step 104, invoking the target AI large model to process the problem request to obtain an answer text corresponding to the problem request.

[0076] Optionally, the target AI large model is invoked to process the question request to obtain an answer text corresponding to the question request, including: generating a question request processing prompt instruction corresponding to the question request and the target AI large model according to the question request and a question request processing prompt instruction template of the target AI large model; inputting the question request processing prompt instruction into the target AI large model to obtain an answer text corresponding to the question request output by the target AI large model.

[0077] Optionally, the question request processing prompt instruction corresponding to the question request and the AI large model can be a text for instructing the AI large model to analyze the question request and generate an answer text corresponding to the question request. For each AI large model provided by a model provider, the question request processing prompt instruction template of the AI large model can be a question request processing prompt instruction without containing a question request. The question request processing prompt instruction template of the AI large model contains a question request filling position. The question request filling position is a position for filling a question request that needs to be processed by the AI large model. After filling the question request into the question request filling position in the question request processing prompt instruction template of the AI large model, a question request processing prompt instruction corresponding to the question request and the AI large model can be obtained. The question request processing prompt instruction templates of the AI large models provided by various model providers are stored in the business server.

[0078] Optionally, the question request processing prompt instruction corresponding to the question request and the target AI large model is generated according to the question request and a question request processing prompt instruction template of the target AI large model, including: copying the question request processing prompt instruction template of the target AI large model stored in the business server; filling the question request into a question request filling position in the copied question request processing prompt instruction template of the target AI large model to obtain a question request processing prompt instruction corresponding to the question request and the target AI large model. The question request processing prompt instruction corresponding to the question request and the target AI large model is a text for instructing the target AI large model corresponding to the question request to analyze the question request and generate an answer text corresponding to the question request.

[0079] Optionally, after the question request processing prompt instruction corresponding to the question request and the target AI large model is input into the target AI large model, the target AI large model analyzes the question request contained in the question request processing prompt instruction to generate an answer text corresponding to the question request, and outputs the answer text corresponding to the question request. The answer text corresponding to the question request output by the target AI large model can be obtained.

[0080] The technical scheme of the embodiment of the application comprises the following steps: obtaining model application information of AI large models provided by each model provider, wherein the model application information comprises a model number and flow limitation information; determining a current load balancing mode according to a key load balancing parameter and a selectable load balancing mode, wherein the selectable load balancing mode comprises a polling balancing mode, a flow balancing mode, a flow weight comprehensive balancing mode and a flow proportion comprehensive balancing mode; when a problem request is detected, determining a target AI large model corresponding to the problem request from the AI large models provided by each model provider according to the current load balancing mode and the model application information of the AI large models provided by each model provider, and calling the target AI large model to process the problem request to obtain an answer text corresponding to the problem request. The AI large model calling scheme in the related art cannot utilize the AI large models provided by multiple different model providers to perform load balancing on the AI large model calling process, and the problem request is distributed to the AI large models provided by each model provider for processing, and the efficiency and stability of the AI large model calling process are difficult to guarantee. However, based on the model application information of the AI large models provided by each model provider, the AI large model calling process can be load balanced, the problem request can be distributed to the AI large models provided by each model provider for processing, and the efficiency and stability of the AI large model calling process are improved.

[0081] Embodiment two

[0082] Figure 2 A flowchart of an AI large model calling load balancing method provided by the embodiment two of the application. The embodiment of the application can be combined with each optional scheme in one or more of the above embodiments. As shown in the figure, the method comprises the following steps: Figure 2

[0083] Step 201: obtaining model application information of AI large models provided by each model provider.

[0084] The model application information comprises a model number and flow limitation information.

[0085] Step 202: when the key load balancing parameter is the model number, determining that the current load balancing mode is the polling balancing mode.

[0086] Step 203: when a problem request is detected, determining the model number of the AI large model used to process the problem request this time according to the number polling order corresponding to the AI large models provided by each model provider and the historical model number.

[0087] The historical model number is the model number of the AI large model used to process the problem request last time.

[0088] ​Step 204, determining an AI large model with the same model number as the model number of the AI large model used for processing the problem request this time from the AI large models provided by each model provider as a target AI large model corresponding to the problem request.

[0089] Step 205, calling the target AI large model to process the problem request to obtain an answer text corresponding to the problem request.

[0090] The technical scheme of the embodiment of the application can load balance the calling process of the AI large model in a polling manner based on the model numbers of the AI large models provided by each model provider, and distribute problem requests to the AI large models provided by each model provider for processing, thereby improving the efficiency and stability of the AI large model calling process.

[0091] Embodiment three

[0092] Figure 3 A flowchart of an AI large model calling load balancing method provided for the third embodiment of the application. The embodiment of the application can be combined with each optional scheme in one or more of the above embodiments. As shown in the figure, the method comprises: Figure 3

[0093] Step 301, obtaining model application information of AI large models provided by each model provider.

[0094] The model application information comprises a model number and flow limit information.

[0095] Step 302, when the key load balancing parameter is a flow quota, determining that the current load balancing mode is a flow balancing mode.

[0096] Step 303, when a problem request is detected, obtaining flow record information of the AI large models provided by each model provider.

[0097] Step 304, determining a flow quota of the AI large models provided by each model provider according to the flow limit information and the flow record information of the AI large models provided by each model provider.

[0098] Step 305, screening each available AI large model from the AI large models provided by each model provider according to the flow quota of the AI large models provided by each model provider.

[0099] Step 306, selecting an available AI large model as a target AI large model corresponding to the problem request from the available AI large models.

[0100] Step 307, calling the target AI large model to process the problem request to obtain an answer text corresponding to the problem request.​

[0101] The technical solution of the embodiment of the application can perform load balancing on the calling process of the AI large model based on the traffic quota of the AI large model provided by each model provider, distribute problem requests to each available AI large model in the AI large model provided by each model provider for processing, and improve the efficiency and stability of the AI large model calling process.

[0102] Embodiment Four

[0103] Figure 4 A flowchart of an AI large model calling load balancing method provided by Embodiment Four of the application. The embodiment of the application can be combined with each optional scheme in one or more of the above embodiments. As shown in the figure, the method comprises the following steps. Figure 4

[0104] Step 401: Obtain model application information of AI large models provided by each model provider.

[0105] The model application information comprises a model number and traffic limit information.

[0106] Step 402: When the key load balancing parameter is traffic quota and weight information, determine that the current load balancing mode is a traffic weight comprehensive balancing mode.

[0107] Step 403: When a problem request is detected, obtain traffic record information of AI large models provided by each model provider.

[0108] Step 404: Determine the traffic quota of AI large models provided by each model provider according to the traffic limit information and the traffic record information of AI large models provided by each model provider.

[0109] Step 405: According to the traffic quota of AI large models provided by each model provider, screen each available AI large model in AI large models provided by each model provider.

[0110] Step 406: Obtain weight information of each available AI large model.

[0111] Step 407: According to the weight information of each available AI large model, screen one available AI large model from the available AI large models as a target AI large model corresponding to the problem request.

[0112] Step 408: Call the target AI large model to process the problem request, and obtain an answer text corresponding to the problem request.

[0113] ​The technical scheme of the embodiment of the present application can perform load balancing on the calling process of the AI large model based on the traffic quota and weight information of the AI large model provided by each model provider, and distribute problem requests to each available AI large model in the AI large model provided by each model provider for processing, thereby improving the efficiency and stability of the AI large model calling process.

[0114] Embodiment five

[0115] Figure 5 A flowchart of an AI large model calling load balancing method provided by the fifth embodiment of the present application. The embodiment of the present application can be combined with each optional scheme in one or more of the above embodiments. As shown in the figure, the method comprises the following steps: Figure 5

[0116] Step 501: Obtain model application information of the AI large model provided by each model provider.

[0117] The model application information comprises a model number and traffic limit information.

[0118] Step 502: When the key load balancing parameter is the traffic quota and the current request processing ratio, determine that the current load balancing mode is the traffic ratio load balancing mode.

[0119] Step 503: When a problem request is detected, obtain traffic record information and the number of processed requests of the AI large model provided by each model provider.

[0120] Step 504: Determine the total number of current requests according to the number of processed requests of the AI large model provided by each model provider.

[0121] Step 505: Determine the traffic quota of the AI large model provided by each model provider according to the traffic limit information and the traffic record information of the AI large model provided by each model provider.

[0122] Step 506: Screen each available AI large model in the AI large model provided by each model provider according to the traffic quota of the AI large model provided by each model provider.

[0123] Step 507: Obtain the standard request processing ratio of each available AI large model.

[0124] Step 508: Determine the current request processing ratio of each available AI large model according to the total number of current requests and the number of processed requests of each available AI large model.

[0125] ​Step 509, screening an available AI large model from the available AI large models as a target AI large model corresponding to the question request according to the standard request processing proportion and the current request processing proportion of each available AI large model.

[0126] Step 510, calling the target AI large model to process the question request to obtain an answer text corresponding to the question request.

[0127] The technical scheme of the embodiment of the application can balance the load of the calling process of the AI large model based on the traffic quota and the current request processing proportion of the AI large model provided by each model provider, and distribute the question request to each available AI large model provided by each model provider for processing, thereby improving the efficiency and stability of the calling process of the AI large model.

[0128] Embodiment six

[0129] Figure 6 A structural schematic diagram of an AI large model calling load balancing device provided by the embodiment six of the application is provided. The device can be configured in an electronic device. As shown in the figure, the device comprises an information acquisition module 601, a mode determination module 602, a model determination module 603, and a model calling module 604. Figure 6

[0130] The information acquisition module 601 is configured to acquire model application information of the AI large model provided by each model provider, wherein the model application information comprises a model number and traffic limit information. The mode determination module 602 is configured to determine a current load balancing mode according to a key load balancing parameter and a selectable load balancing mode. The selectable load balancing mode comprises a round robin balancing mode, a traffic balancing mode, a traffic weight comprehensive balancing mode, and a traffic proportion comprehensive balancing mode. The model determination module 603 is configured to determine a target AI large model corresponding to a question request from the AI large models provided by each model provider according to the current load balancing mode and the model application information of the AI large model provided by each model provider when the question request is detected. The model calling module 604 is configured to call the target AI large model to process the question request to obtain an answer text corresponding to the question request.

[0131] ​The technical scheme of the embodiment of the application acquires model application information of AI large models provided by each model provider, the model application information including a model number and flow limit information; then determines a current load balancing mode according to a key load balancing parameter and a selectable load balancing mode, the selectable load balancing mode including: a polling balancing mode, a flow balancing mode, a flow weight comprehensive balancing mode, and a flow proportion comprehensive balancing mode; when a problem request is detected, determines a target AI large model corresponding to the problem request from the AI large models provided by each model provider according to the current load balancing mode and the model application information of the AI large models provided by each model provider, calls the target AI large model to process the problem request, and obtains an answer text corresponding to the problem request, solving the problem that the AI large model calling scheme in the related art cannot utilize AI large models provided by multiple different model providers to perform load balancing on an AI large model calling process, divert problem requests to AI large models provided by each model provider for processing, and the efficiency and stability of the AI large model calling process are difficult to guarantee, and the AI large model calling process can be load balanced based on the model application information of the AI large models provided by each model provider, problem requests are diverted to AI large models provided by each model provider for processing, and the efficiency and stability of the AI large model calling process are improved.

[0132] In an optional implementation of the embodiment of the application, the mode determination module 602 is specifically configured to: when the key load balancing parameter is the model number, determine that the current load balancing mode is the polling balancing mode; when the key load balancing parameter is the flow quota, determine that the current load balancing mode is the flow balancing mode; when the key load balancing parameter is the flow quota and the weight information, determine that the current load balancing mode is the flow weight comprehensive balancing mode; and when the key load balancing parameter is the flow quota and the current request processing proportion, determine that the current load balancing mode is the flow proportion comprehensive balancing mode.

[0133] In an optional implementation of the embodiment of the application, when the mode determination module 602 performs the operation of determining a target AI large model corresponding to the problem request from the AI large models provided by each model provider according to the current load balancing mode and the model application information of the AI large models provided by each model provider, the mode determination module 602 is specifically configured to: when the current load balancing mode is the polling balancing mode, determine the model number of the AI large model used to process the problem request this time according to the number polling order corresponding to the AI large models provided by each model provider and the historical model number; wherein the historical model number is the model number of the AI large model used to process the problem request last time; and determine the AI large model with the same model number as the model number of the AI large model used to process the problem request this time from the AI large models provided by each model provider as the target AI large model corresponding to the problem request.

[0134] In an optional implementation of the embodiment of the application, optionally, the mode determination module 602, when performing the operation of determining the target AI large model corresponding to the problem request from the AI large models provided by the model suppliers according to the current load balancing mode and the model application information of the AI large models provided by the model suppliers, is specifically configured to: when the current load balancing mode is the traffic balancing mode, obtain the traffic record information of the AI large models provided by the model suppliers; determine the traffic quota of the AI large models provided by the model suppliers according to the traffic limit information and the traffic record information of the AI large models provided by the model suppliers; and select each available AI large model from the AI large models provided by the model suppliers according to the traffic quota of the AI large models provided by the model suppliers; and select one available AI large model from the available AI large models as the target AI large model corresponding to the problem request.

[0135] In an optional implementation of the embodiment of the application, optionally, the mode determination module 602, when performing the operation of determining the target AI large model corresponding to the problem request from the AI large models provided by the model suppliers according to the current load balancing mode and the model application information of the AI large models provided by the model suppliers, is specifically configured to: when the current load balancing mode is the traffic weight balancing mode, obtain the traffic record information of the AI large models provided by the model suppliers; determine the traffic quota of the AI large models provided by the model suppliers according to the traffic limit information and the traffic record information of the AI large models provided by the model suppliers; select each available AI large model from the AI large models provided by the model suppliers according to the traffic quota of the AI large models provided by the model suppliers; obtain the weight information of each available AI large model; and select one available AI large model from the available AI large models as the target AI large model corresponding to the problem request according to the weight information of each available AI large model.

[0136] In an optional implementation of the embodiment of the application, optionally, the mode determination module 602, when performing the operation of determining the target AI large model corresponding to the problem request from the AI large models provided by the respective model providers according to the current load balancing mode and the model application information of the AI large models provided by the respective model providers, is specifically configured to: when the current load balancing mode is the traffic proportion load balancing mode, obtain the traffic record information and the number of processed requests of the AI large models provided by the respective model providers; determine the total number of current requests according to the number of processed requests of the AI large models provided by the respective model providers; determine the traffic quota of the AI large models provided by the respective model providers according to the traffic limit information and the traffic record information of the AI large models provided by the respective model providers; filter out each available AI large model in the AI large models provided by the respective model providers according to the traffic quota of the AI large models provided by the respective model providers; obtain the standard request processing proportion of each available AI large model; determine the current request processing proportion of each available AI large model according to the total number of current requests and the number of processed requests of each available AI large model; and filter out one available AI large model as the target AI large model corresponding to the problem request from the available AI large models according to the standard request processing proportion and the current request processing proportion of each available AI large model.

[0137] In an optional implementation of the embodiment of the application, optionally, the model calling module 604 is specifically configured to: generate the problem request processing prompt instruction corresponding to the problem request and the target AI large model according to the problem request and the problem request processing prompt instruction template of the target AI large model; and input the problem request processing prompt instruction into the target AI large model to obtain the answer text corresponding to the problem request output by the target AI large model.

[0138] The AI large model calling load balancing device provided in the embodiment of the application can perform the AI large model calling load balancing method provided in any embodiment of the application, and has the corresponding function modules and beneficial effects of the execution method.

[0139] Embodiment Seven

[0140] Figure 7A structural diagram of an electronic device 10 that can be used to implement the AI large model invocation load balancing method of embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, electronic devices, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0141] As shown in Figure 7 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., communicatively connected to the at least one processor 11, where the memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0142] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, speakers, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0143] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the AI large model invocation load balancing method.

[0144] In some embodiments, the AI ​​large model call load balancing method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on a heterogeneous hardware accelerator via a ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the processor, one or more steps of the AI ​​large model call load balancing method described above may be performed. Alternatively, in other embodiments, the processor may be configured to execute the AI ​​large model call load balancing method by any other appropriate means (e.g., by means of firmware).

[0145] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0146] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or electronic device.

[0147] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0148] To provide for interaction with a user, the systems and techniques described here can be implemented on a heterogeneous hardware accelerator having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the heterogeneous hardware accelerator. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0149] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0150] A computing system can include a client and an electronic device. The client and the electronic device are typically in different locations, and are often interacted with over a communications network. The relationship of the client and the electronic device is created by computer programs running on the respective computers and having a client-electronic device relationship with each other. The electronic device can be a cloud electronic device, also known as a cloud computing electronic device or a cloud host, which is a host product in a cloud computing service system, and solves the defects of large management difficulty and weak business scalability in traditional physical hosts and VPS services.

[0151] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.

[0152] The above detailed description does not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A load balancing method for calling a large AI model, characterized in that: include: Obtain model application information of the AI ​​large model provided by each model supplier; wherein the model application information includes model number and flow limit information; Determine the current load balancing mode based on key load balancing parameters and optional load balancing modes; wherein the optional load balancing modes include: polling balancing mode, flow balancing mode, flow weighted comprehensive balancing mode, and flow ratio comprehensive balancing mode; When a problem request is detected, determining a target AI big model corresponding to the problem request from the AI ​​big models provided by each model supplier based on the current load balancing mode and the model application information of the AI ​​big models provided by each model supplier; The target AI big model is called to process the question request to obtain an answer text corresponding to the question request.

2. The AI ​​large model call load balancing method according to claim 1 is characterized in that: Determine the current load balancing mode based on key load balancing parameters and optional load balancing modes, including: When the key load balancing parameter is the model number, the current load balancing mode is determined to be the round-robin balancing mode; When the key load balancing parameter is the traffic quota, the current load balancing mode is determined to be the traffic balancing mode; When the key load balancing parameters are traffic quota and weight information, the current load balancing mode is determined to be traffic weight comprehensive balancing mode; When the key load balancing parameters are traffic quota and current request processing ratio, the current load balancing mode is determined to be traffic ratio comprehensive balancing mode.

3. The AI ​​large model call load balancing method according to claim 1 is characterized in that: Determining a target AI big model corresponding to the question request from among the AI ​​big models provided by each model supplier based on the current load balancing mode and the model application information of the AI ​​big models provided by each model supplier, including: When the current load balancing mode is the round-robin balancing mode, the model number of the AI ​​large model used to process the problem request is determined based on the polling order of the numbers corresponding to the AI ​​large models provided by each model supplier and the historical model number; wherein the historical model number is the model number of the AI ​​large model used to process the problem request last time; The AI ​​big model provided by each model supplier and having the same model number as the AI ​​big model used to process the question request this time is determined as the target AI big model corresponding to the question request.

4. The AI ​​large model call load balancing method according to claim 1 is characterized in that: Determining a target AI big model corresponding to the question request from among the AI ​​big models provided by each model supplier based on the current load balancing mode and the model application information of the AI ​​big models provided by each model supplier, including: When the current load balancing mode is the traffic balancing mode, traffic record information of the AI ​​large model provided by each model supplier is obtained; Determine the traffic quota for the AI ​​large models provided by each model supplier based on the traffic limit information and traffic record information of the AI ​​large models provided by each model supplier; Based on the traffic quota of the AI ​​big models provided by each model supplier, filter out the available AI big models from the AI ​​big models provided by each model supplier; An available AI big model is selected from various available AI big models as a target AI big model corresponding to the question request.

5. The AI ​​large model call load balancing method according to claim 1 is characterized in that: Determining a target AI big model corresponding to the question request from among the AI ​​big models provided by each model supplier based on the current load balancing mode and the model application information of the AI ​​big models provided by each model supplier, including: When the current load balancing mode is the traffic weighted comprehensive balancing mode, traffic record information of the AI ​​large model provided by each model supplier is obtained; Determine the traffic quota for the AI ​​large models provided by each model supplier based on the traffic limit information and traffic record information of the AI ​​large models provided by each model supplier; Based on the traffic quota of the AI ​​big models provided by each model supplier, filter out the available AI big models from the AI ​​big models provided by each model supplier; Obtain weight information for each available AI model; According to the weight information of each available AI big model, an available AI big model is selected from each available AI big model as the target AI big model corresponding to the question request.

6. The AI ​​large model call load balancing method according to claim 1 is characterized in that: Determining a target AI big model corresponding to the question request from among the AI ​​big models provided by each model supplier based on the current load balancing mode and the model application information of the AI ​​big models provided by each model supplier, including: When the current load balancing mode is the traffic proportion comprehensive balancing mode, the traffic record information and the number of processed requests of the AI ​​large model provided by each model supplier are obtained; Determine the total number of current requests based on the number of processed requests for the AI ​​large model provided by each model provider; Determine the traffic quota for the AI ​​large models provided by each model supplier based on the traffic limit information and traffic record information of the AI ​​large models provided by each model supplier; Based on the traffic quota of the AI ​​big models provided by each model supplier, filter out the available AI big models from the AI ​​big models provided by each model supplier; Get the standard request processing ratio of each available AI model; Determine a current request processing ratio of each available AI large model based on the total number of current requests and the number of processed requests of each available AI large model; According to the standard request processing ratio and current request processing ratio of each available AI big model, an available AI big model is selected from each available AI big model as the target AI big model corresponding to the problem request.

7. The AI ​​large model call load balancing method according to claim 1 is characterized in that: The target AI model is called to process the question request to obtain an answer text corresponding to the question request, including: Generate a question request processing prompt instruction corresponding to the question request and the target AI large model according to the question request and the question request processing prompt instruction template of the target AI large model; The question request processing prompt instruction is input into the target AI big model to obtain the answer text corresponding to the question request output by the target AI big model.

8. A load balancing device for calling a large AI model, characterized in that: include: An information acquisition module is used to obtain model application information of AI large models provided by various model suppliers; wherein the model application information includes model number and flow limit information; A mode determination module is used to determine the current load balancing mode based on key load balancing parameters and optional load balancing modes; wherein the optional load balancing modes include: polling balancing mode, flow balancing mode, flow weighted comprehensive balancing mode and flow ratio comprehensive balancing mode; a model determination module configured to, when a problem request is detected, determine a target AI big model corresponding to the problem request from among the AI ​​big models provided by each model supplier based on the current load balancing mode and model application information of the AI ​​big models provided by each model supplier; The model calling module is used to call the target AI large model to process the question request and obtain the answer text corresponding to the question request.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the AI ​​large model call load balancing method described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the AI ​​large model call load balancing method according to any one of claims 1 to 7 when executed.