AI large model calling load balancing method, device, equipment and medium
By identifying the domain labels of AI large models and classifying the problem requests, the load balancing problem during the invocation of AI large models was solved, improving processing efficiency and stability.
Patent Information
- Application Number
- CN202510922706.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-17
AI Technical Summary
The existing AI large model calling solution cannot effectively utilize multiple AI large models for load balancing, making it difficult to ensure processing efficiency and stability.
By using a pre-trained model information processing model, the domain labels of each candidate AI large model are determined, and the question requests are grouped and distributed to the AI large models in their respective domains for processing using a question request classification model.
This system enables efficient routing of requests to multiple large AI models for processing after a batch of requests are received, thereby improving the efficiency and stability of the large AI model invocation process.
Smart Images

Figure CN120803719A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric digital data processing, and in particular to an AI large model calling load balancing method and device, equipment and a medium. BACKGROUND
[0002] With the development of artificial intelligence (AI) technology, more and more enterprises begin to call AI large models to process obtained user question requests to obtain answer texts corresponding to the question requests. After obtaining a batch of question requests to be processed each time, the business server of the enterprise calls the AI large model to process each question request to obtain an answer text corresponding to each question request.
[0003] In related technologies, a commonly used AI large model calling scheme is to select an AI large model from multiple callable AI large models as an AI large model for processing question requests obtained by a business server. After obtaining a batch of question requests to be processed each time, the business server uniformly calls the AI large model for processing question requests obtained by the business server to process each question request to obtain an answer text corresponding to each question request. The AI large model calling scheme in related technologies calls a fixed AI large model to process question requests, cannot balance the load of the AI large model calling process by using multiple callable AI large models, and cannot distribute each question request to multiple callable AI large models for processing after obtaining a batch of question requests to be processed each time, so that the efficiency and stability of the AI large model calling process are difficult to guarantee. SUMMARY
[0004] The present application provides an AI large model calling load balancing method, device, equipment and medium to solve the problem that the AI large model calling scheme in related technologies cannot balance the load of the AI large model calling process by using multiple callable AI large models, cannot distribute each question request to multiple callable AI large models for processing after obtaining a batch of question requests to be processed each time, and the efficiency and stability of the AI large model calling process are difficult to guarantee.
[0005] According to an aspect of the present application, an AI large model calling load balancing method is provided, comprising:
[0006] By means of the pre-trained model information processing model, the field characteristic text of each candidate AI large model is used to determine the field tag at which each candidate AI large model is good at, and the pre-set question field category that matches the field tag at which each candidate AI large model is good at is selected from each pre-set question field category; wherein the field characteristic text is attribute text, evaluation text or use record text.
[0007] When a set of to-be-processed problem requests is detected, a problem domain category of each problem request in the set of to-be-processed problem requests is screened from each preset problem domain category by a pre-trained problem request classification model;
[0008] Each problem request is grouped according to the problem domain category of the problem request, to obtain at least one problem request group; wherein the problem requests in each problem request group have the same problem domain category;
[0009] Each problem request group is distributed to each candidate AI large model for processing according to the problem domain category of each problem request group and the proficient domain label of each candidate AI large model.
[0010] According to another aspect of the present application, an AI large model calling load balancing device is provided, comprising:
[0011] A label processing module is configured to determine a proficient domain label of each candidate AI large model according to a domain representation text of each candidate AI large model by a pre-trained model information processing model, and screen a preset problem domain category matching the proficient domain label of each candidate AI large model from each preset problem domain category; wherein the domain representation text is an attribute text, an evaluation text or a use record text;
[0012] A category processing module is configured to screen a problem domain category of each problem request in a set of to-be-processed problem requests from each preset problem domain category by a pre-trained problem request classification model when the set of to-be-processed problem requests is detected;
[0013] A request grouping module is configured to group each problem request according to the problem domain category of the problem request, to obtain at least one problem request group; wherein the problem requests in each problem request group have the same problem domain category;
[0014] A distribution processing module is configured to distribute each problem request group to each candidate AI large model for processing according to the problem domain category of each problem request group and the proficient domain label of each candidate AI large model.
[0015] According to another aspect of the present application, an electronic device is provided, comprising:
[0016] at least one processor;
[0017] and a memory in communication connection with the at least one processor;
[0018] The memory stores a computer program executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the AI large model calling load balancing method according to any one of the embodiments of the application.
[0019] According to another aspect of the application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the AI large model calling load balancing method according to any one of the embodiments of the application when executed by the processor.
[0020] According to another aspect of the application, a computer program product is provided, which comprises a computer program for implementing the AI large model calling load balancing method according to any one of the embodiments of the application when executed by a processor.
[0021] The technical scheme of the embodiment of the application determines the strong field label of each candidate AI large model according to the field representation text of each candidate AI large model through the pre-trained model information processing model, and filters the preset problem field category matching the strong field label of each candidate AI large model from each preset problem field category; wherein the field representation text is attribute text, evaluation text or use record text; when detecting the to-be-processed problem request set, the problem field category of each problem request in the to-be-processed problem request set is filtered from each preset problem field category through the pre-trained problem request classification model; then each problem request is grouped according to the problem field category of each problem request, and at least one problem request group is obtained, and each problem request group is distributed to each candidate AI large model for processing according to the strong field label of each candidate AI large model and the problem field category of each problem request group; wherein the problem field category of the problem request in each problem request group is the same, which solves the problem in the related art that the AI large model calling scheme cannot load balance the AI large model calling process by utilizing multiple callable AI large models, cannot distribute each problem request to multiple callable AI large models for processing after obtaining a batch of problem requests to be processed each time, and the efficiency and stability of the AI large model calling process are difficult to guarantee, can determine the strong field label of each callable AI large model based on the attribute text, evaluation text or use record text of each callable AI large model, can determine the problem field category of each problem request through the problem request classification model and group each problem request based on the problem field category of each problem request after obtaining a batch of problem requests to be processed each time, and then distributes each problem request group to each callable AI large model for processing based on the strong field label of each callable AI large model and the problem field category of each problem request group, thereby distributing each problem request to multiple callable AI large models for processing after obtaining a batch of problem requests to be processed each time, and improving the efficiency and stability of the AI large model calling process.
[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the application, nor is it used to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0024] Figure 1 A flow chart of an AI large model calling load balancing method provided for the first embodiment of the present application.
[0025] Figure 2 A flow chart of an AI large model calling load balancing method provided for the second embodiment of the present application.
[0026] Figure 3 A flow chart of an AI large model calling load balancing method provided for the third embodiment of the present application.
[0027] Figure 4 A flow chart of an AI large model calling load balancing method provided for the fourth embodiment of the present application.
[0028] Figure 5 A structural schematic diagram of an AI large model calling load balancing device provided for the fifth embodiment of the present application.
[0029] Figure 6 A structural schematic diagram of an electronic device for implementing the AI large model calling load balancing method of the embodiments of the present application. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of the present application. It should be noted that the terms “target”, “first”, “second” and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms “contain”, “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device containing a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] Embodiment one
[0032] Figure 1A flowchart of an AI large model calling load balancing method provided for Embodiment One of the present application. This embodiment can be applicable to the case of load balancing of the AI large model calling process in the scenario of calling the AI large model to process the problem request. The method can be executed by an AI large model calling load balancing device, which can be realized in the form of hardware and / or software, and can be configured in an electronic device, such as a business server set in an enterprise, which is used in cooperation with the client sending the problem request. The business server can be a server for processing the business of the enterprise. As shown in Figure 1 The method comprises the following steps:
[0033] Step 101: determining the field of expertise label of each candidate AI large model according to the field representation text of each candidate AI large model by a pre-trained model information processing model, and screening the pre-set problem field category matching the field of expertise label of each candidate AI large model from each pre-set problem field category.
[0034] The field representation text can be attribute text, evaluation text, or usage record text.
[0035] Optionally, the AI large model can refer to a large language model (LLM) trained using AI technology and capable of performing inference processing such as problem answering, machine translation, and text generation. The number of candidate AI large models is greater than or equal to 2. Each candidate AI large model is an AI large model that can be called by a business server to process the problem request obtained by the business server and obtain the answer text corresponding to the problem request. The problem request can be a problem text input by a user. The problem text can be a text describing a problem related to the enterprise that the user is concerned about. The answer text corresponding to the problem request can be a text describing the answer to the problem described in the problem request. Each candidate AI large model is a different AI large model that can be used to process the problem request obtained by the business server. Each candidate AI large model can run in the business server, and can also run in other servers connected to the business server.
[0036] Optionally, for each candidate AI large model, the field of expertise label of the candidate AI large model can refer to one or more words representing the field of the problem described by the problem request that the candidate AI large model is good at processing, which is determined according to the field representation text of the candidate AI large model. Exemplarily, the field of the problem includes but is not limited to the fields of recruitment, production technology, daily life, law, and learning. The number of fields of the problem described by the problem request that the candidate AI large model is good at processing is greater than or equal to 1.
[0037] Optionally, for each candidate AI large model, the domain representation text of the candidate AI large model can refer to a text that contains words representing the domain of the problem request description that the candidate AI large model is good at processing, which is pre-collected and stored in the business server. The domain representation text of the candidate AI large model is attribute text, evaluation text, or usage record text. The attribute text of the candidate AI large model can be text provided by the developer of the candidate AI large model to describe the candidate AI large model. The evaluation text of the candidate AI large model can be text provided by the user of the candidate AI large model to evaluate the candidate AI large model. The usage record text of the candidate AI large model can be text recorded by the technical personnel of the business server or the enterprise, which describes the advantages and problems of the candidate AI large model found in the process of using the candidate AI large model. Generally, the attribute text, evaluation text, and usage record text of the candidate AI large model will contain words representing the domain of the problem request description that the candidate AI large model is good at processing.
[0038] Optionally, the number of preset problem domain categories is greater than or equal to 2. Each preset problem domain category can be a pre-set word used to describe a domain. Different preset problem domain categories are words used to describe different domains. Generally, the domain of the problem request description obtained by the business server will be contained in each domain described by each preset problem domain category. Each preset problem domain category is stored in the business server.
[0039] Optionally, for each candidate AI large model, the preset problem domain category matching the proficient domain tag of the candidate AI large model can refer to one or more preset problem domain categories whose described domain is contained in the domain represented by the proficient domain tag of the candidate AI large model.
[0040] Optionally, the pre-trained model information processing model can be a large language model pre-trained by the technical personnel and set in the business server, which is used to process information related to the AI large model according to the prompt instruction. The prompt instruction input to the pre-trained model information processing model is a text used to instruct the pre-trained model information processing model to perform a specified processing operation on the specified information related to the AI large model. When the prompt instruction is input to the pre-trained model information processing model, the pre-trained model information processing model will perform the specified processing operation on the specified information and output the processing result. The prompt instruction input to the pre-trained model information processing model includes but is not limited to: domain keyword extraction prompt instruction, related category screening prompt instruction.
[0041] Optionally, the domain keyword extraction prompt instruction is a text for instructing the pre-trained model information processing model to extract domain keywords from the domain representation text of the specified alternative AI large model. The domain keyword extraction prompt instruction contains the domain representation text of the specified alternative AI large model. The domain keyword is a word used to describe a domain in the domain representation text of the alternative AI large model.
[0042] Optionally, the related category screening prompt instruction is a text for instructing the pre-trained model information processing model to screen out the pre-set problem domain categories related to the proficient domain label of the specified alternative AI large model from the various pre-set problem domain categories. The domain keyword extraction prompt instruction contains the various pre-set problem domain categories and the proficient domain label of the specified alternative AI large model. Each pre-set problem domain category is a word used to describe a domain. The proficient domain label of the alternative AI large model is one or more words used to represent the domain of the problem that the alternative AI large model is proficient in processing. The proficient domain label of the alternative AI large model contains at least one word. The pre-set problem domain category related to the proficient domain label of the alternative AI large model can refer to the pre-set problem domain category with a correlation greater than a pre-set value with any one of the words in the proficient domain label of the alternative AI large model. The correlation between words can refer to the probability of two words appearing in the same context. The pre-set value can be a pre-set value.
[0043] Optionally, by the pre-trained model information processing model, the proficient domain label of each alternative AI large model is determined according to the domain representation text of each alternative AI large model, and the pre-set problem domain category matching the proficient domain label of each alternative AI large model is screened out from the various pre-set problem domain categories, including: inputting the domain keyword extraction prompt instruction corresponding to the domain representation text of each alternative AI large model into the pre-trained model information processing model to obtain the domain keywords extracted from the domain representation text of each alternative AI large model output by the model information processing model; determining the domain keywords extracted from the domain representation text of each alternative AI large model as the proficient domain label of each alternative AI large model; inputting the related category screening prompt instruction corresponding to the proficient domain label of each alternative AI large model into the model information processing model to obtain the pre-set problem domain categories related to the proficient domain label of each alternative AI large model screened out from the various pre-set problem domain categories output by the model information processing model; and determining the pre-set problem domain categories related to the proficient domain label of each alternative AI large model as the pre-set problem domain categories matching the proficient domain label of each alternative AI large model.
[0044] Optionally, the domain keyword extraction prompt instruction corresponding to the domain representation text of each alternative AI large model is input into the pre-trained model information processing model, and the domain keyword extracted from the domain representation text of each alternative AI large model output by the model information processing model is obtained, including: for each alternative AI large model, the following operations are performed: according to the domain representation text of the alternative AI large model and the domain keyword extraction prompt instruction template, the domain keyword extraction prompt instruction corresponding to the domain representation text of the alternative AI large model is generated; the domain keyword extraction prompt instruction is input into the pre-trained model information processing model, and the domain keyword extracted from the domain representation text of the alternative AI large model output by the model information processing model is obtained.
[0045] Optionally, the business memory stores a domain keyword extraction prompt instruction template. The domain keyword extraction prompt instruction template is a domain keyword extraction prompt instruction that does not include the domain representation text of the specified alternative AI large model. The domain keyword extraction prompt instruction template includes a text filling position. The text filling position is a position for filling the domain representation text of the alternative AI large model that needs to extract the domain keyword. After filling the domain representation text of the alternative AI large model that needs to extract the domain keyword into the text filling position in the domain keyword extraction prompt instruction template, the domain keyword extraction prompt instruction corresponding to the domain representation text of the alternative AI large model that needs to extract the domain keyword can be obtained. The domain keyword extraction prompt instruction corresponding to the domain representation text of the alternative AI large model that needs to extract the domain keyword is a text for instructing the pre-trained model information processing model to extract the domain keyword from the domain representation text of the alternative AI large model.
[0046] Optionally, according to the domain representation text of the alternative AI large model and the domain keyword extraction prompt instruction template, the domain keyword extraction prompt instruction corresponding to the domain representation text of the alternative AI large model is generated, including: copying the domain keyword extraction prompt instruction template stored in the business server; filling the domain representation text of the alternative AI large model into the text filling position in the copied domain keyword extraction prompt instruction template, thereby obtaining the domain keyword extraction prompt instruction corresponding to the domain representation text of the alternative AI large model. The domain keyword extraction prompt instruction corresponding to the domain representation text of the alternative AI large model is a text for instructing the pre-trained model information processing model to extract the domain keyword from the domain representation text of the alternative AI large model.
[0047] Optionally, the domain keyword extraction prompt instruction corresponding to the domain representation text of the alternative AI large model is input into the pre-trained model information processing model, the pre-trained model information processing model extracts the domain keyword from the domain representation text of the alternative AI large model contained in the domain keyword extraction prompt instruction, and outputs the domain keyword extracted from the domain representation text of the alternative AI large model. For each alternative AI large model, the domain keyword extracted from the domain representation text of the alternative AI large model is usually a word that can be used to represent the domain of the problem request description that the alternative AI large model is good at processing. The domain keyword extracted from the domain representation text of the alternative AI large model output by the pre-trained model information processing model can be obtained, and the domain keyword extracted from the domain representation text of the alternative AI large model is determined as the domain label of the alternative AI large model.
[0048] Optionally, the relevant category screening prompt instruction corresponding to the domain label of each alternative AI large model is input into the model information processing model, and the preset problem domain category related to the domain label of each alternative AI large model screened from each preset problem domain category output by the model information processing model is obtained. The operation includes: for each alternative AI large model, generating a relevant category screening prompt instruction corresponding to the domain label of the alternative AI large model according to each preset problem domain category, the domain label of the alternative AI large model, and the relevant category screening prompt instruction template; inputting the relevant category screening prompt instruction into the model information processing model, and obtaining the preset problem domain category related to the domain label of the alternative AI large model screened from each preset problem domain category output by the model information processing model.
[0049] Optionally, the business memory stores a related category screening prompt instruction template. The related category screening prompt instruction template is a related category screening prompt instruction without containing each preset problem field category and the proficient field label of the specified alternative AI large model. The related category screening prompt instruction template contains a category filling position and a label filling position. The category filling position is a position for filling each preset problem field category. The label filling position is a position for filling the proficient field label of the alternative AI large model that needs to screen out the related preset problem field category from each preset problem field category. After filling each preset problem field category into the category filling position in the related category screening prompt instruction template and filling the proficient field label of the alternative AI large model that needs to screen out the related preset problem field category from each preset problem field category into the label filling position in the related category screening prompt instruction template, the related category screening prompt instruction corresponding to the proficient field label of the alternative AI large model that needs to screen out the related preset problem field category from each preset problem field category can be obtained. The related category screening prompt instruction corresponding to the proficient field label of the alternative AI large model that needs to screen out the related preset problem field category from each preset problem field category is a text for instructing the pre-trained model information processing model to screen out the preset problem field category related to the proficient field label of the alternative AI large model from each preset problem field category.
[0050] Optionally, according to each preset problem field category, the proficient field label of the alternative AI large model, and the related category screening prompt instruction template, the related category screening prompt instruction corresponding to the proficient field label of the alternative AI large model is generated, including: copying the related category screening prompt instruction template stored in the business server; filling each preset problem field category into the category filling position in the copied related category screening prompt instruction template, and filling the proficient field label of the alternative AI large model into the label filling position in the copied related category screening prompt instruction template to obtain the related category screening prompt instruction corresponding to the proficient field label of the alternative AI large model. The related category screening prompt instruction corresponding to the proficient field label of the alternative AI large model is a text for instructing the pre-trained model information processing model to screen out the preset problem field category related to the proficient field label of the alternative AI large model from each preset problem field category.
[0051] Optionally, the related category screening prompt instruction corresponding to the expertise field label of the alternative AI large model is input into the pre-trained model information processing model, the pre-trained model information processing model screens the preset problem field category related to the expertise field label of the alternative AI large model from each preset problem field category contained in the related category screening prompt instruction, and outputs the preset problem field category related to the expertise field label of the alternative AI large model screened from each preset problem field category.
[0052] Optionally, for each alternative AI large model, the preset problem field category related to the expertise field label of the alternative AI large model screened from each preset problem field category is usually a preset problem field category whose described field is included in the field represented by the expertise field label of the alternative AI large model. The preset problem field category related to the expertise field label of the alternative AI large model screened from each preset problem field category output by the pre-trained model information processing model can be obtained, and the preset problem field category related to the expertise field label of the alternative AI large model screened from each preset problem field category is determined as the preset problem field category matched with the expertise field label of the alternative AI large model. The preset problem field category matched with the expertise field label of the alternative AI large model usually contains at least one preset problem field category.
[0053] Step 102, when detecting a set of to-be-processed problem requests, screening the problem field category of each problem request in the set of to-be-processed problem requests from each preset problem field category through the pre-trained problem request classification model.
[0054] Optionally, the set of to-be-processed problem requests can be a set composed of a plurality of problem requests that need to be processed by the business server. The problem field category of the problem request can be a preset problem field category in each preset problem field category that can be used to describe the field of the problem described in the problem request description. It can be detected whether the business server receives the set of to-be-processed problem requests, and each time the business server receives the set of to-be-processed problem requests is detected, the problem field category of each problem request in the set of to-be-processed problem requests is screened from each preset problem field category through the pre-trained problem request classification model.
[0055] Optionally, the pre-trained question request classification model can be a large language model pre-trained by the technical personnel and set in the business server, and used to determine the question domain category of the question request according to a question request classification prompt instruction. The question request classification prompt instruction can be text used to instruct the pre-trained question request classification model to filter the question domain category of the specified question request from the various preset question domain categories, so as to determine the question domain category of the specified question request. The question request classification prompt instruction contains the various preset question domain categories and the specified question request. The question request classification prompt instruction is input into the pre-trained question request classification model, and the pre-trained question request classification model filters the question domain category of the specified question request from the various preset question domain categories, and outputs the question domain category of the specified question request.
[0056] Optionally, filtering the question domain category of each question request in the set of to-be-processed question requests from the various preset question domain categories by the pre-trained question request classification model comprises: inputting a question request classification prompt instruction corresponding to each question request in the set of to-be-processed question requests into the pre-trained question request classification model to obtain the question domain category of each question request filtered from the various preset question domain categories by the question request classification model.
[0057] Optionally, inputting the question request classification prompt instruction corresponding to each question request in the set of to-be-processed question requests into the pre-trained question request classification model to obtain the question domain category of each question request filtered from the various preset question domain categories by the question request classification model comprises: for each question request in the set of to-be-processed question requests, performing the following operations: generating a question request classification prompt instruction corresponding to the question request according to the various preset question domain categories, the question request, and a question request classification prompt instruction template; inputting the question request classification prompt instruction into the pre-trained question request classification model to obtain the question domain category of the question request filtered from the various preset question domain categories by the question request classification model.
[0058] Optionally, the business storage stores a question request classification prompt instruction template. The question request classification prompt instruction template is a question request classification prompt instruction without a specified question request. The question request classification prompt instruction template includes a category filling position and a question request filling position. The category filling position is a position for filling each preset problem domain category. The question request filling position is a position for filling a question request that needs to be filtered from each preset problem domain category. After filling each preset problem domain category into the category filling position of the question request classification prompt instruction template and filling the question request that needs to be filtered from each preset problem domain category into the question request filling position of the question request classification prompt instruction template, a question request classification prompt instruction corresponding to the question request that needs to be filtered from each preset problem domain category can be obtained. The question request classification prompt instruction corresponding to the question request that needs to be filtered from each preset problem domain category is used to instruct the pre-trained question request classification model to filter the problem domain category of the question request from each preset problem domain category, so as to determine the problem domain category of the question request.
[0059] Optionally, according to the question request and the question request classification prompt instruction template, the question request classification prompt instruction corresponding to the question request is generated, including: copying the question request classification prompt instruction template stored in the business server; filling each preset problem domain category into the category filling position of the copied question request classification prompt instruction template, and filling the question request into the question request filling position of the copied question request classification prompt instruction template to obtain the question request classification prompt instruction corresponding to the question request. The question request classification prompt instruction corresponding to the question request is used to instruct the pre-trained question request classification model to filter the problem domain category of the question request from each preset problem domain category, so as to determine the problem domain category of the question request.
[0060] Optionally, the question request classification prompt instruction corresponding to the question request is input into the pre-trained question request classification model, and the pre-trained question request classification model filters the problem domain category of the question request in the question request classification prompt instruction from each preset problem domain category and outputs the problem domain category of the question request filtered from each preset problem domain category. For each question request in the set of to-be-processed question requests, the problem domain category of the question request filtered from each preset problem domain category output by the pre-trained question request classification model is obtained, so as to determine the problem domain category of the question request in the set of to-be-processed question requests.
[0061] Step 103, grouping each question request according to the question field category of each question request to obtain at least one question request group.
[0062] In each question request group, the question field categories of the question requests are the same.
[0063] Optionally, the question requests with the same question field category in each question request are divided into a question request group to obtain at least one question request group. Each question request group contains at least one question request. The question field categories of the question requests in each question request group are the same. For each question request, if there is another question request with the same question field category in each other question request, the question request and the other question request with the same question field category are divided into a question request group. If there is no other question request with the same question field category in each other question request, the question request is determined as a question request group, and the question request group contains only one question request.
[0064] Step 104, according to the field of expertise tags of each candidate AI large model and the question field categories of each question request group, distributing each question request group to each candidate AI large model for processing.
[0065] Optionally, according to the field of expertise tags of each candidate AI large model and the question field categories of each question request group, distributing each question request group to each candidate AI large model for processing, including: determining each candidate AI large model whose matched preset question field category contains the question field category of each question request group as each optional AI large model of each question request group; for each question request group, according to the weight information, request processing proportion information or traffic information of each optional AI large model, screening one optional AI large model as a target AI large model corresponding to the question request group in each optional AI large model, and calling the target AI large model to process each question request in the question request group to obtain an answer text corresponding to each question request.
[0066] Optionally, each optional AI large model of the question request group can refer to an alternative AI large model in the alternative AI large models that is good at processing the question request in the question request group. The question domain category of the question request group refers to the question domain category of the question request in the question request group. The question domain category of the question request refers to a preset question domain category in the preset question domain categories that can be used to describe the domain of the question described in the question request description. For each preset question domain category, there is usually at least one preset question domain category matched by the expertise tag of the alternative AI large model that contains the preset question domain category. Usually, if the preset question domain category matched by the expertise tag of the alternative AI large model contains the question domain category of the question request group, it can be determined that the alternative AI large model is an alternative AI large model that is good at processing the question request in the question request group, and is an optional AI large model of the question request group.
[0067] Optionally, determining each alternative AI large model whose matched preset question domain category contains the question domain category of each question request group as the optional AI large model of each question request group includes: for each question request group, performing the following operation: detecting whether the preset question domain category matched by the expertise tag of each alternative AI large model contains the question domain category of the question request group; determining each alternative AI large model whose matched preset question domain category contains the question domain category of the question request group as the optional AI large model of the question request group.
[0068] Optionally, detecting whether the preset question domain category matched by the expertise tag of each alternative AI large model contains the question domain category of the question request group includes: for each alternative AI large model, performing the following operation: determining whether there is a preset question domain category that is the same as the question domain category of the question request group in the preset question domain category matched by the expertise tag of the alternative AI large model; if yes, determining that the preset question domain category matched by the expertise tag of the alternative AI large model contains the question domain category of the question request group; if no, determining that the preset question domain category matched by the expertise tag of the alternative AI large model does not contain the question domain category of the question request group.
[0069] Optionally, according to the weight information, request processing ratio information or traffic information of each optional AI large model, screening one optional AI large model from the optional AI large models as a target AI large model corresponding to the question request group includes: determining a maximum value in the weight information of each optional AI large model, and determining the optional AI large model to which the maximum value belongs as the target AI large model corresponding to the question request group.
[0070] Optionally, the target AI large model corresponding to the question request packet can be an AI large model selected from the plurality of optional AI large models in the question request packet for processing the question request in the question request packet to obtain an answer text corresponding to the question request.
[0071] Optionally, the business server stores weight information of each of the optional AI large models. The weight information of the optional AI large model can be a numerical value for representing a degree of ability of the optional AI large model to process the question request. The greater the weight information of the optional AI large model, the higher the ability of the optional AI large model to process the question request. The smaller the weight information of the optional AI large model, the lower the ability of the optional AI large model to process the question request.
[0072] Optionally, the maximum value in the weight information of each of the optional AI large models is determined, and the optional AI large model to which the maximum value belongs is determined as the target AI large model corresponding to the question request packet, including: obtaining the weight information of each of the optional AI large models of the question request packet stored in the business server; determining the maximum value in the weight information of each of the optional AI large models, and determining the optional AI large model to which the maximum value belongs as the target AI large model corresponding to the question request packet. If there is an optional AI large model whose weight information is the maximum value, the optional AI large model whose weight information is the maximum value is determined as the target AI large model corresponding to the question request packet. If there are a plurality of optional AI large models whose weight information is the maximum value, one of the plurality of optional AI large models whose weight information is the maximum value is randomly selected as the target AI large model corresponding to the question request packet.
[0073] Optionally, the request processing ratio information includes a standard request processing ratio and a current request processing ratio; and the target AI large model corresponding to the question request packet is selected from the plurality of optional AI large models according to the weight information of each of the optional AI large models, the request processing ratio information, or the traffic information, including: selecting one of the optional AI large models in which the current request processing ratio is less than the standard request processing ratio as the target AI large model corresponding to the question request packet.
[0074] Optionally, the service server stores request processing ratio information of each candidate AI large model. The request processing ratio information includes a standard request processing ratio and a current request processing ratio. The standard request processing ratio of the candidate AI large model can be a ratio of a total number of problem requests that the service server has invoked the candidate AI large model to process in an ideal state to a total number of problem requests that the service server has invoked all candidate AI large models to process, which is set according to the ability of the candidate AI large model to process problem requests. The current request processing ratio of the candidate AI large model can be a ratio of a total number of problem requests that the service server has invoked the candidate AI large model to process to a total number of problem requests that the service server has invoked all candidate AI large models to process, which is counted by the service server.
[0075] Optionally, screening one candidate AI large model as a target AI large model corresponding to the problem request group from the candidate AI large models whose current request processing ratio is less than the standard request processing ratio includes: obtaining the standard request processing ratio and the current request processing ratio of each candidate AI large model of the problem request group stored in the service server; judging whether the current request processing ratio of each candidate AI large model is less than the standard request processing ratio of the candidate AI large model; and randomly selecting one candidate AI large model as the target AI large model corresponding to the problem request group from the candidate AI large models whose current request processing ratio is less than the standard request processing ratio.
[0076] Optionally, the traffic information includes traffic limit information and traffic record information; and screening one candidate AI large model as a target AI large model corresponding to the problem request group from the candidate AI large models according to the weight information, the request processing ratio information or the traffic information of each candidate AI large model includes: determining a traffic quota of each candidate AI large model according to the traffic limit information and the traffic record information of each candidate AI large model; and screening one candidate AI large model as the target AI large model corresponding to the problem request group from the candidate AI large models according to the traffic quota of each candidate AI large model.
[0077] Optionally, the traffic information of each candidate AI large model is stored in the business server. The traffic information includes traffic limit information and traffic record information. The traffic limit information of the candidate AI large model can be the maximum number of calls per hour of the candidate AI large model. The maximum number of calls per hour of the candidate AI large model can refer to the maximum number of calls allowed by the business server to the candidate AI large model within 1 hour while ensuring the normal operation of the candidate AI large model. The traffic record information of the candidate AI large model can be the total number of calls of the candidate AI large model by the business server within the current hour. The traffic quota of the candidate AI large model can refer to the total number of calls of the candidate AI large model remaining in the current hour by the business server. The traffic quota of the candidate AI large model is equal to the difference between the traffic limit information and the traffic record information of the candidate AI large model.
[0078] Optionally, according to the traffic limit information and the traffic record information of each optional AI large model, the traffic quota of each optional AI large model is determined, and according to the traffic quota of each optional AI large model, one optional AI large model is selected from the optional AI large models as the target AI large model corresponding to the problem request group, which includes: obtaining the traffic limit information and the traffic record information of each optional AI large model of the problem request group stored in the business server; for each optional AI large model, calculating the difference between the traffic limit information and the traffic record information of the optional AI large model; randomly selecting one optional AI large model from the optional AI large models whose difference between the traffic limit information and the traffic record information is greater than 0 as the target AI large model corresponding to the problem request group.
[0079] Optionally, for each problem request group, if there is only one optional AI large model whose preset problem domain category matched by the domain label contains the problem domain category of the problem request group, the optional AI large model is directly determined as the target AI large model corresponding to the problem request group.
[0080] Optionally, the target AI large model is called to process each problem request in the problem request group to obtain an answer text corresponding to each problem request, which includes: for each problem request in the problem request group, the following operation is performed: generating a problem request processing prompt instruction corresponding to the problem request and the target AI large model according to the problem request and the problem request processing prompt instruction template of the target AI large model; inputting the problem request processing prompt instruction into the target AI large model to obtain the answer text corresponding to the problem request output by the target AI large model.
[0081] Optionally, the question request processing prompt instruction corresponding to the question request and the candidate AI large model can be a text for instructing the candidate AI large model to analyze the question request and generate an answer text corresponding to the question request. For each candidate AI large model, the question request processing prompt instruction template of the candidate AI large model can be a question request processing prompt instruction without containing the question request. The question request processing prompt instruction template of the candidate AI large model contains a question request filling position. The question request filling position is a position for filling the question request that needs to be processed by the candidate AI large model. After filling the question request into the question request filling position in the question request processing prompt instruction template of the candidate AI large model, the question request processing prompt instruction corresponding to the question request and the candidate AI large model can be obtained. The question request processing prompt instruction template of each candidate AI large model is stored in the business server.
[0082] Optionally, the question request processing prompt instruction corresponding to the question request and the target AI large model is generated according to the question request and the question request processing prompt instruction template of the target AI large model, which includes copying the question request processing prompt instruction template of the target AI large model stored in the business server, and filling the question request into the question request filling position in the copied question request processing prompt instruction template of the target AI large model to obtain the question request processing prompt instruction corresponding to the question request and the target AI large model. The question request processing prompt instruction corresponding to the question request and the target AI large model is a text for instructing the target AI large model corresponding to the question request group where the question request is located to analyze the question request and generate an answer text corresponding to the question request.
[0083] Optionally, after the question request processing prompt instruction corresponding to the question request and the target AI large model is input into the target AI large model, the target AI large model analyzes the question request contained in the question request processing prompt instruction, generates an answer text corresponding to the question request, and outputs the answer text corresponding to the question request. The answer text corresponding to the question request output by the target AI large model can be obtained.
[0084] The technical scheme of the embodiment of the present application determines the field label at which each candidate AI large model is good at according to the field representation text of each candidate AI large model through the pre-trained model information processing model, and filters the preset problem field categories matching the field label at which each candidate AI large model is good at from each preset problem field category; wherein the field representation text is attribute text, evaluation text or use record text; when detecting the set of problem requests to be processed, the problem request classification model is used to filter the problem field category of each problem request in the set of problem requests to be processed from each preset problem field category; then, each problem request is grouped according to the problem field category of each problem request, and at least one problem request group is obtained, and each problem request group is distributed to each candidate AI large model for processing according to the field label at which each candidate AI large model is good at and the problem field category of each problem request group; wherein the problem field category of each problem request in each problem request group is the same, which solves the problem that the AI large model calling scheme in the related art cannot load balance the AI large model calling process by using multiple callable AI large models, cannot distribute each problem request to multiple callable AI large models for processing after obtaining a batch of problem requests to be processed each time, and the efficiency and stability of the AI large model calling process are difficult to guarantee, and can determine the field label at which each callable AI large model is good at based on the attribute text, evaluation text or use record text of each callable AI large model, can determine the problem field category of each problem request through the problem request classification model and group each problem request based on the problem field category of each problem request after obtaining a batch of problem requests to be processed each time, and then distribute each problem request group to each callable AI large model for processing based on the field label at which each callable AI large model is good at and the problem field category of each problem request group, so that each problem request is distributed to multiple callable AI large models for processing after obtaining a batch of problem requests to be processed each time, and the efficiency and stability of the AI large model calling process are improved.
[0085] Embodiment two
[0086] Figure 2 A flowchart of an AI large model calling load balancing method provided for the second embodiment of the present application. The second embodiment of the present application can be combined with one or more optional schemes in the above-mentioned embodiments. As shown in the figure, the method comprises: Figure 2
[0087] Step 201, determining a field label at which each candidate AI large model is good at according to field representation text of each candidate AI large model by a pre-trained model information processing model, and screening a preset problem field category matching the field label at which each candidate AI large model is good at from each preset problem field category.
[0088] The field representation text is attribute text, evaluation text or use record text.
[0089] Step 202, when a set of problem requests to be processed is detected, screening a problem field category of each problem request in the set of problem requests to be processed from each preset problem field category by a pre-trained problem request classification model.
[0090] Step 203, grouping each problem request according to the problem field category of the problem request, to obtain at least one problem request group.
[0091] The problem field categories of the problem requests in each problem request group are the same.
[0092] Step 204, determining each candidate AI large model whose field label at which the AI large model is good at matches the problem field category of each problem request group as each selectable AI large model of each problem request group.
[0093] Step 205, for each problem request group, determining a maximum value in weight information of each selectable AI large model, determining a selectable AI large model to which the maximum value belongs as a target AI large model corresponding to the problem request group, and calling the target AI large model to process each problem request in the problem request group to obtain an answer text corresponding to each problem request.
[0094] The technical scheme of the embodiment of the application can determine the problem field category of each problem request by a problem request classification model and group each problem request based on the problem field category of the problem request after obtaining a batch of problem requests to be processed each time, and then distribute each problem request group to each callable AI large model for processing based on the field label at which each callable AI large model is good at and weight information of the AI large model, and the problem field category of each problem request group, so that each problem request is distributed to multiple callable AI large models for processing after a batch of problem requests to be processed are obtained each time, and the efficiency and stability of the AI large model calling process are improved.
[0095] Embodiment three
[0096] Figure 3A flowchart of an AI large model calling load balancing method provided for Embodiment Three of the present application. Embodiment Three of the present application can be combined with each optional scheme in one or more of the above embodiments. As shown in FIG. 8, the method comprises: Figure 3
[0097] Step 301: determining, by a pre-trained model information processing model, a field label at which each candidate AI large model is good at according to a field representation text of each candidate AI large model, and filtering, from each preset problem field category, a preset problem field category matching the field label at which each candidate AI large model is good at.
[0098] The field representation text is attribute text, evaluation text, or usage record text.
[0099] Step 302: when detecting a set of problem requests to be processed, filtering, by a pre-trained problem request classification model, a problem field category of each problem request in the set of problem requests to be processed from each preset problem field category.
[0100] Step 303: grouping each problem request according to the problem field category of the problem request, to obtain at least one problem request group.
[0101] The problem field categories of the problem requests in each problem request group are the same.
[0102] Step 304: determining each candidate AI large model whose field label at which the candidate AI large model is good at matches the problem field category of each problem request group as each optional AI large model of each problem request group.
[0103] Step 305: for each problem request group, filtering one optional AI large model from each optional AI large model whose current request processing ratio is less than a standard request processing ratio as a target AI large model corresponding to the problem request group, and calling the target AI large model to process each problem request in the problem request group to obtain an answer text corresponding to each problem request.
[0104] The technical scheme of the embodiment of the present application can determine the problem domain categories of each problem request through a problem request classification model after obtaining a batch of problem requests to be processed each time, group each problem request based on the problem domain categories of each problem request, and then distribute each problem request group to each callable AI large model for processing based on the field of expertise tags and request processing proportion information of each callable AI large model, and the problem domain categories of each problem request group, so as to distribute each problem request to multiple callable AI large models for processing after obtaining a batch of problem requests to be processed each time, thereby improving the efficiency and stability of the AI large model calling process.
[0105] Embodiment four
[0106] Figure 4 A flowchart of an AI large model calling load balancing method provided by the fourth embodiment of the present application. The embodiment of the present application can be combined with each optional scheme in one or more of the above embodiments. As shown in the figure, the method comprises: Figure 4
[0107] Step 401: determining the field of expertise tags of each candidate AI large model according to the domain representation text of each candidate AI large model through a pre-trained model information processing model, and screening the preset problem domain categories that match the field of expertise tags of each candidate AI large model from each preset problem domain category.
[0108] Among them, the domain representation text is attribute text, evaluation text or use record text.
[0109] Step 402: when detecting a set of problem requests to be processed, screening the problem domain categories of each problem request in the set of problem requests to be processed from each preset problem domain category through a pre-trained problem request classification model.
[0110] Step 403: grouping each problem request according to the problem domain category of each problem request to obtain at least one problem request group.
[0111] Among them, the problem domain categories of the problem requests in each problem request group are the same.
[0112] Step 404: determining each candidate AI large model whose matched preset problem domain category contains the problem domain categories of each problem request group as each optional AI large model of each problem request group.
[0113] In step 405, for each question request packet, the traffic quota of each optional AI large model is determined according to the traffic limit information and the traffic record information of each optional AI large model, one optional AI large model is selected as a target AI large model corresponding to the question request packet from each optional AI large model according to the traffic quota of each optional AI large model, the target AI large model is called to process each question request in the question request packet to obtain an answer text corresponding to each question request.
[0114] The technical scheme of the embodiment of the application can determine the question domain category of each question request by the question request classification model and group each question request based on the question domain category of each question request after obtaining a batch of question requests to be processed each time, and then distribute each question request packet to each callable AI large model for processing based on the field of expertise tag and the traffic information of each callable AI large model and the question domain category of each question request packet, so as to distribute each question request to a plurality of callable AI large models for processing after obtaining a batch of question requests to be processed each time, thereby improving the efficiency and stability of the AI large model calling process.
[0115] Embodiment five
[0116] Figure 5 A structural schematic diagram of an AI large model calling load balancing device provided by the embodiment five of the application is provided. The device can be configured in an electronic device. As shown in the figure, the device comprises a label processing module 501, a category processing module 502, a request grouping module 503, and a distribution processing module 504. Figure 5
[0117] The label processing module 501 is configured to determine the field label at which each candidate AI large model is good at according to the field representation text of each candidate AI large model by using a pre-trained model information processing model, and filter a preset problem field category matching the field label at which each candidate AI large model is good at from each preset problem field category; the field representation text is attribute text, evaluation text, or usage record text; the category processing module 502 is configured to filter the problem field category of each problem request in the set of problem requests to be processed from each preset problem field category by using a pre-trained problem request classification model when the set of problem requests to be processed is detected; the request grouping module 503 is configured to group each problem request according to the problem field category of each problem request to obtain at least one problem request group; the problem field category of each problem request in each problem request group is the same; and the shunt processing module 504 is configured to shunt each problem request group to each candidate AI large model for processing according to the field label at which each candidate AI large model is good at and the problem field category of each problem request group.
[0118] The technical scheme of the embodiment of the present application determines the field label at which each candidate AI large model is good at according to the field representation text of each candidate AI large model through the pre-trained model information processing model, and filters the preset problem field category matching the field label at which each candidate AI large model is good at from each preset problem field category; wherein the field representation text is attribute text, evaluation text or use record text; when detecting the to-be-processed problem request set, the problem field category of each problem request in the to-be-processed problem request set is filtered from each preset problem field category through the pre-trained problem request classification model; then each problem request is grouped according to the problem field category of each problem request, and at least one problem request group is obtained, and each problem request group is distributed to each candidate AI large model for processing according to the field label at which each candidate AI large model is good at and the problem field category of each problem request group; wherein the problem field category of each problem request in each problem request group is the same, which solves the problem in the related art that the AI large model calling scheme cannot utilize multiple callable AI large models to load balance the AI large model calling process, cannot distribute each problem request to multiple callable AI large models for processing after obtaining a batch of problem requests to be processed each time, and the efficiency and stability of the AI large model calling process are difficult to guarantee, and can determine the field label at which each callable AI large model is good at based on the attribute text, evaluation text or use record text of each callable AI large model, can determine the problem field category of each problem request through the problem request classification model and group each problem request based on the problem field category of each problem request after obtaining a batch of problem requests to be processed each time, and then distribute each problem request group to each callable AI large model for processing based on the field label at which each callable AI large model is good at and the problem field category of each problem request group, so that each problem request is distributed to multiple callable AI large models for processing after obtaining a batch of problem requests to be processed each time, and the efficiency and stability of the AI large model calling process are improved.
[0119] In an optional implementation of the embodiment of the present application, optionally, the label processing module 501 is specifically configured to: input a domain keyword extraction prompt instruction corresponding to the domain representation text of each candidate AI large model into a pre-trained model information processing model to obtain domain keywords extracted from the domain representation text of each candidate AI large model output by the model information processing model; determine the domain keywords extracted from the domain representation text of each candidate AI large model as the domain expertise label of each candidate AI large model; input a related category filtering prompt instruction corresponding to the domain expertise label of each candidate AI large model into the model information processing model to obtain a preset problem domain category related to the domain expertise label of each candidate AI large model filtered from each preset problem domain category output by the model information processing model; and determine the preset problem domain category related to the domain expertise label of each candidate AI large model as the preset problem domain category matched with the domain expertise label of each candidate AI large model.
[0120] In an optional implementation of the embodiment of the present application, optionally, the category processing module 502, when performing the operation of filtering the problem domain category of each problem request in the set of problem requests to be processed from each preset problem domain category through a pre-trained problem request classification model, is specifically configured to: input a problem request classification prompt instruction corresponding to each problem request in the set of problem requests to be processed into a pre-trained problem request classification model to obtain the problem domain category of each problem request filtered from each preset problem domain category output by the problem request classification model.
[0121] In an optional implementation of the embodiment of the present application, optionally, the shunt processing module 504 is specifically configured to: determine each candidate AI large model whose matched preset problem domain category contains the problem domain category of each problem request grouping as each optional AI large model of each problem request grouping; for each problem request grouping, filter an optional AI large model as a target AI large model corresponding to the problem request grouping according to the weight information, request processing proportion information or traffic information of each optional AI large model, and call the target AI large model to process each problem request in the problem request grouping to obtain an answer text corresponding to each problem request.
[0122] In an optional implementation of the embodiment of the application, optionally, when performing the operation of screening one optional AI large model from the plurality of optional AI large models as the target AI large model corresponding to the question request group according to the weight information, the request processing proportion information or the traffic information of the plurality of optional AI large models, the shunt processing module 504 is specifically configured to: determine the maximum value in the weight information of the plurality of optional AI large models, and determine the optional AI large model to which the maximum value belongs as the target AI large model corresponding to the question request group.
[0123] In an optional implementation of the embodiment of the application, optionally, the request processing proportion information includes a standard request processing proportion and a current request processing proportion; when performing the operation of screening one optional AI large model from the plurality of optional AI large models as the target AI large model corresponding to the question request group according to the weight information, the request processing proportion information or the traffic information of the plurality of optional AI large models, the shunt processing module 504 is specifically configured to: screen one optional AI large model from the optional AI large models in which the current request processing proportion is less than the standard request processing proportion as the target AI large model corresponding to the question request group.
[0124] In an optional implementation of the embodiment of the application, optionally, the traffic information includes traffic limit information and traffic record information; when performing the operation of screening one optional AI large model from the plurality of optional AI large models as the target AI large model corresponding to the question request group according to the weight information, the request processing proportion information or the traffic information of the plurality of optional AI large models, the shunt processing module 504 is specifically configured to: determine the traffic quota of each optional AI large model according to the traffic limit information and the traffic record information of each optional AI large model, and screen one optional AI large model from the plurality of optional AI large models as the target AI large model corresponding to the question request group according to the traffic quota of each optional AI large model.
[0125] The AI large model calling load balancing device provided by the embodiment of the application can perform the AI large model calling load balancing method provided by any embodiment of the application, and has the corresponding function modules and beneficial effects of the execution method.
[0126] Embodiment six
[0127] Figure 6A structural diagram of an electronic device 10 that can be used to implement the AI large model invocation load balancing method of embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, electronic devices, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.
[0128] As shown in Figure 6 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., communicatively connected to the at least one processor 11, where the memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0129] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, speakers, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0130] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the AI large model invocation load balancing method.
[0131] In some embodiments, the AI large model invocation load balancing method can be implemented as a computer program tangibly embodied in a computer readable storage medium, e.g., a memory unit. In some embodiments, portions or all of the computer program can be loaded and / or installed onto the heterogeneous hardware accelerator via a ROM and / or communication unit. When the computer program is loaded onto the RAM and executed by the processor, one or more steps of the AI large model invocation load balancing method described above can be performed. Alternatively, in other embodiments, the processor can be configured to perform the AI large model invocation load balancing method by other any suitable means, e.g., by means of firmware.
[0132] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0133] Computer programs used to implement the processes of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as part of a standalone software package, or entirely on a remote machine or electronic device.
[0134] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0135] To provide for interaction with a user, the systems and techniques described here can be implemented on a heterogeneous hardware accelerator having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the heterogeneous hardware accelerator. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0136] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), blockchain network, and the Internet.
[0137] A computing system can include a client and an electronic device. The client and the electronic device are typically in different locations, and are often interacted with over a communications network. The relationship of the client and the electronic device is created by computer programs running on the respective computers and having a client-electronic device relationship with each other. The electronic device can be a cloud electronic device, also known as a cloud computing electronic device or a cloud host, which is a host product in a cloud computing service system, and solves the defects of large management difficulty and weak business scalability in traditional physical hosts and VPS services.
[0138] It should be understood that the various forms of flow shown above can be reordered, steps added or deleted. For example, the steps described in the present application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which are not limited herein.
[0139] The above detailed description does not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A load balancing method for calling a large AI model, characterized in that: include: Through the pre-trained model information processing model, the domain labels of each candidate AI large model are determined based on the domain representation texts of each candidate AI large model, and preset problem domain categories that match the domain labels of each candidate AI large model are screened from each preset problem domain category; wherein the domain representation text is attribute text, evaluation text, or usage record text; When a set of pending question requests is detected, the problem domain category of each question request in the set of pending question requests is screened out from various preset problem domain categories using a pre-trained problem request classification model; Grouping each question request according to the problem domain category of each question request to obtain at least one question request group; wherein the problem domain category of the question requests in each question request group is the same; According to the labels of the areas of expertise of each candidate AI model and the problem area categories of each question request group, each question request group is diverted to each candidate AI model for processing.
2. The AI large model call load balancing method according to claim 1 is characterized in that: Through the pre-trained model information processing model, the domain labels of each candidate AI model are determined based on the domain representation text of each candidate AI model. Preset problem domain categories that match the domain labels of each candidate AI model are screened from various preset problem domain categories, including: Inputting the domain keyword extraction prompt instruction corresponding to the domain representation text of each candidate AI large model into the pre-trained model information processing model, and obtaining the domain keywords extracted from the domain representation text of each candidate AI large model output by the model information processing model; The domain keywords extracted from the domain representation texts of each candidate AI model are determined as the domain labels of each candidate AI model; Inputting relevant category screening prompt instructions corresponding to the expertise labels of each candidate AI large model into the model information processing model, obtaining preset problem domain categories related to the expertise labels of each candidate AI large model screened from various preset problem domain categories and output by the model information processing model; The preset problem domain category related to the domain-of-excellence label of each candidate AI big model is determined as the preset problem domain category that matches the domain-of-excellence label of each candidate AI big model.
3. The AI large model call load balancing method according to claim 1 is characterized in that: The problem domain category of each problem request in the set of problem requests to be processed is screened out from various preset problem domain categories using a pre-trained problem request classification model, including: The question request classification prompt instruction corresponding to each question request in the set of pending question requests is input into a pre-trained question request classification model to obtain the question domain category of each question request screened from each preset question domain category output by the question request classification model.
4. The AI large model call load balancing method according to claim 1 is characterized in that: Based on the domain labels of each candidate AI model and the problem domain categories of each question request group, each question request group is divided into different candidate AI models for processing, including: Determine each candidate AI big model whose preset problem domain category matched by the proficiency domain label includes the problem domain category of each problem request group as each optional AI big model of each problem request group; For each question request group, based on the weight information, request processing ratio information or traffic information of each optional AI big model, an optional AI big model is selected from each optional AI big model as the target AI big model corresponding to the question request group, and the target AI big model is called to process each question request in the question request group to obtain the answer text corresponding to each question request.
5. The AI large model call load balancing method according to claim 4 is characterized in that: Based on the weight information, request processing ratio information, or traffic information of each optional AI big model, one optional AI big model is selected from each optional AI big model as the target AI big model corresponding to the problem request group, including: The maximum value among the weight information of each optional AI big model is determined, and the optional AI big model to which the maximum value belongs is determined as the target AI big model corresponding to the question request group.
6. The AI large model call load balancing method according to claim 4 is characterized in that: The request processing ratio information includes a standard request processing ratio and a current request processing ratio; Based on the weight information, request processing ratio information, or traffic information of each optional AI big model, one optional AI big model is selected from each optional AI big model as the target AI big model corresponding to the problem request group, including: Among the optional AI big models whose current request processing ratio is less than the standard request processing ratio, one optional AI big model is selected as the target AI big model corresponding to the problem request group.
7. The AI large model call load balancing method according to claim 4 is characterized in that: The flow information includes flow restriction information and flow record information; Based on the weight information, request processing ratio information, or traffic information of each optional AI big model, one optional AI big model is selected from each optional AI big model as the target AI big model corresponding to the problem request group, including: Based on the traffic limit information and traffic record information of each optional AI big model, the traffic quota of each optional AI big model is determined. Based on the traffic quota of each optional AI big model, an optional AI big model is selected from each optional AI big model as the target AI big model corresponding to the question request group.
8. A load balancing device for calling a large AI model, characterized in that: include: A label processing module is used to determine the domain labels of each candidate AI model based on the domain representation text of each candidate AI model through a pre-trained model information processing model, and to screen out preset problem domain categories that match the domain labels of each candidate AI model from each preset problem domain category; wherein the domain representation text is attribute text, evaluation text, or usage record text; A category processing module is configured to, when a set of pending question requests is detected, filter out the problem domain category of each question request in the set of pending question requests from various preset problem domain categories using a pre-trained question request classification model; a request grouping module, configured to group each question request according to the problem domain category of each question request to obtain at least one question request group; wherein the problem domain category of the question requests in each question request group is the same; The diversion processing module is used to divert each question request group to each alternative AI large model for processing based on the domain labels of each alternative AI large model and the problem domain categories of each question request group.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the AI large model call load balancing method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the AI large model call load balancing method according to any one of claims 1 to 7 when executed.