Large language model access method, system and device and storage medium
By extracting the target parameters in the user access request and distributing them to the matching executor, the problem of incompatibility of interfaces of different large language models is solved, and efficient interface compatibility and access efficiency is achieved.
Patent Information
- Application Number
- CN202411982132.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-06-03
AI Technical Summary
There are differences in the interfaces of different large language models in the prior art, which leads to incompatibility and affects efficient use.
By extracting the target parameters in the user access request, including access protocol, model identification and model source, the request is distributed to the matching target executor, and the executor executes the target operation process for dialogue and interaction with the target model.
It realizes API interfaces that are fast and efficiently compatible with many large language models, handles user access requests to arbitrary models, solves interface incompatibility and differences, reduces docking work, and improves access efficiency.
Smart Images

Figure CN120086323A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, system, device, and storage medium for accessing large language models. Background Art
[0002] With the rapid development of artificial intelligence technology, current large language models emerge in an endless stream, and the large language models supported by each cloud platform are also changing with each passing day. Differences in one or more of manufacturers, platforms, model types, model versions, etc. may lead to interface differences between different large language models.
[0003] In the prior art, users may face the problem of incompatible interfaces when accessing large language models with different API interfaces, which may further cause users to be unable to access different large language models smoothly and efficiently, bringing certain usage troubles to users. Summary of the Invention
[0004] The main purpose of this application is to provide a method, system, device, and storage medium for accessing large language models, aiming to solve the technical problem in the prior art that different interfaces of large language models are incompatible, affecting the efficient use of large language models.
[0005] In the first aspect of this application, a method for accessing a large language model is provided. The method for accessing a large language model includes:
[0006] Extract a first target parameter from the obtained user access request, where the first target parameter includes at least one of an access protocol, a model identifier of a target model to be accessed, and a model source;
[0007] Distribute the user access request to a target executor that matches the first target parameter;
[0008] Use the target executor to execute a target operation process to perform a conversation interaction with the target model.
[0009] This application also provides a system for accessing a large language model. The system for accessing a large language model includes: a request distribution module, an execution module, and a conversation engine. The execution module includes at least one executor;
[0010] The request distribution module is configured to extract a first target parameter from the obtained user access request, where the first target parameter includes at least one of an access protocol, a model identifier of a target model to be accessed, and a model source; and distribute the user access request to a target executor that matches the first target parameter;
[0011] The execution module is configured to use the target executor to execute a target operation process to perform a conversation interaction with the target model through the conversation engine.
[0012] A third aspect of the present application provides a computer device, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory to enable the computer device to execute the above-mentioned access method for large language models.
[0013] A fourth aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it enables the computer to execute the above-mentioned access method for large language models.
[0014] The present application provides an access solution for large language models. By setting executors corresponding to different models, it can quickly and efficiently be compatible with the API interfaces of numerous large language models, can process access requests of users to any model, solves the problems of incompatibility and interface differences of different model interfaces, reduces the repetitive docking work in the prior art, and improves the access efficiency to various large language models. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic flowchart of the first embodiment of the access method for large language models in the embodiments of the present application;
[0016] Figure 2 It is a schematic diagram of the functional modules of an embodiment of the access system for large language models in the embodiments of the present application;
[0017] Figure 3 It is a schematic diagram of the functional modules of another embodiment of the access system for large language models in the embodiments of the present application;
[0018] Figure 4 It is a schematic diagram of an embodiment of the computer device in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0020] Currently, large language models are emerging in an endless stream, and the large language models supported by each cloud platform are also changing with each passing day. In order to quickly and efficiently be compatible with numerous large language model API interfaces, a general framework is needed to improve the docking efficiency and reduce repetitive docking work.
[0021] Reference Figure 1 , this application provides a method for accessing a large language model, and the method for accessing the large language model includes:
[0022] S100: Extract a first target parameter from the obtained user access request, where the first target parameter includes at least one of an access protocol, a model identifier of a target model to be accessed, and a model source.
[0023] Specifically, the method for accessing a large language model in this embodiment can be applied to an access system for a large language model built based on a general framework for accessing a large language model.
[0024] Users can input parameters according to open or general API interfaces and send user access requests.
[0025] At the front end or client of the access system for a large language model, there is a human-computer interaction interface. For example, users can input dialogue text at the human-computer interaction interface. For example, the dialogue text is a question text: "What's the weather like today? Can you recommend a travel guide suitable for today's weather?"
[0026] There can also be model options for multiple large language models at the human-computer interaction interface. These model options include multiple large language models, and the model options can be displayed in the form of a drop-down list or tiled display, etc., not limited to this form. Users can select or specify a target model to be accessed from multiple different models in the drop-down list or tiled display.
[0027] The system generates a user access request according to the user's operations on the human-computer interaction interface.
[0028] The following are examples of two different user access requests:
[0029] User access request 1:
[0030] {
[0031] "prompt": "Search for the weather in Shenzhen on December 18, 2024",
[0032] "options": {
[0033] "parentMessageId": "64b1310a81aa2f610a88d6d2"
[0034] },
[0035] "systemMessage":null,
[0036] "botId":"6478385ab9443b0966111fed",
[0037] "projectId":"649cfbdf41756556bc427241",
[0038] "id":"64b5068836740c3ad2092c6e",
[0039] "accountId":"649cfbdf41756556bc427240",
[0040] "questionId":"64b5068836740c3ad2092c6e",
[0041] "answerId":"64b5068836740c3ad2092c6f",
[0042] "contextList":["I am Xiaohua","Who are you","I am very happy today",""],
[0043] "projectCreatorId":null,
[0044] "coinAccountId":"649cfbdf41756556bc427240",
[0045] "noCorrelationResponse":null,
[0046] "createTime":1689585288520,
[0047] "conversationId":null,
[0048] "openAiApiKey":
[0049] "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",
[0050] "chatKey":"AKIDra6RZabnriMmCPOYdhlT8f6cG2Km9NiF",
[0051] "chatSecretKey":"fTJ6pTE2lZL1YzQm6PrWBh9Xqz4KWd6F",
[0052] "activate":"GPT_BOTS",
[0053] "coinProjectId":"649cfbdf41756556bc427241",
[0054] "operateAccountId":"649cfbdf41756556bc427240",
[0055] "topK":5,
[0056] "shortMemoryToken":163,
[0057] "longMemoryToken":0,
[0058] "datasetToken":6717,
[0059] "botInfo":{
[0060] "projectId":"649cfd9041756556bc427243",
[0061] "creativityLevel":0.8,
[0062] "docRelevance":0.1,
[0063] "botId":"6478385ab9443b0966111fed",
[0064] "botName":"Java Technical Expert",
[0065] "mode":"excellent",
[0066] "botType":"QuestionAnswer",
[0067] "prompt":"You are a Java expert",
[0068] "docCorrelationSwitch":true,
[0069] "noCorrelationResponse":null,
[0070] "longMemorySwitch": true,
[0071] "aiSource": "Tencent",
[0072] "aiModel": "hunyuan-pro",
[0073] "endpoint": "https: / / hunyuan.tencentcloudapi.com"
[0074] },
[0075] "endpoint": "https: / / hunyuan.tencentcloudapi.com",
[0076] "aiSource": "Tencent",
[0077] "debugSwitch": false,
[0078] "locale": "en_US",
[0079] "docSwitch": true,
[0080] "maxTokens": 1024,
[0081] "modelName": "hunyuan-pro"
[0082] }
[0083] User access request 2:
[0084] {
[0085] "prompt": "Hello",
[0086] "botId": "655da1d3977a1107a028b746",
[0087] "projectId": "gen-lang-client-0801872983",
[0088] "id": "0P4En4h2Agr0rXT",
[0089] "accountId": "ding_group:cid5CI40XUM4qCjVKYg4H / 2Eg==",
[0090] "questionId":"656dbf07249a893046ad4fb0",
[0091] "answerId":"656dbf07249a893046ad4fb1",
[0092] "contextList":null,
[0093] "coinAccountId":"64759a27acf95b3638a7f2d6",
[0094] "noCorrelationResponse":"",
[0095] "conversationId":"656985be122f090d745734e3",
[0096] "openAiApiKey":
[0097] "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",
[0098] "activate":"GPT_BOTS",
[0099] "coinProjectId":"6530d8c83250964271fbd25e",
[0100] "topK":5,
[0101] "shortMemoryToken":0,
[0102] "longMemoryToken":0,
[0103] "datasetToken":7700,
[0104] "botInfo":{
[0105] "projectId":"6530d8c83250964271fbd25e",
[0106] "creativityLevel":0.8,
[0107] "docRelevance":0.8,
[0108] "botId":"655da1d3977a1107a028b746",
[0109] "botName":"DingTalk Debugging",
[0110] "mode":"excellent",
[0111] "botType":"QuestionAnswer",
[0112] "prompt":"",
[0113] "docCorrelationSwitch":true,
[0114] "noCorrelationResponse":"",
[0115] "longMemorySwitch":false,
[0116] "aiModel":"Vertex"
[0117] },
[0118] "debugSwitch":false,
[0119] "locale":"en_US",
[0120] "blocking":true,
[0121] "docSwitch":false,
[0122] "embeddingRate":1.0,
[0123] "chatKey":"AIzaSyBtxawj-SPIydwvt3fWJUt1xxTb1irwzUw",
[0124] "aiSource":"Google",
[0125] "modelName":"gemini-1.5-pro",
[0126] "chatSecretKey":null,
[0127] "maxTokens":1024,
[0128] "endpoint":"https: / / generativelanguage.googleapis.com / v1beta / models / ","protocol":"Google"
[0129] The first target parameter indicating the model type of the target model to be accessed is carried in the user access request. Based on the first target parameter, the type of the target model to be accessed can be determined.
[0130] For example, in the above two user access requests, "protocol":"Google", "aiSource":
[0131] "Google", "modelName":"hunyuan-pro", "modelName":"gemini-1.5-pro" can all be used as the first target parameter. "Google" and "hunyuan-pro" are the access protocols or model names or model sources of the target model to be accessed.
[0132] S200: Distribute the user access request to the target executor that matches the first target parameter.
[0133] Specifically, the executor is the executor. Each model corresponds to an executor. It can be that all models of the same type correspond to the same executor, or it can be that models of the same type are classified according to their models and correspond to different executors, and different types of models correspond to different executors.
[0134] Each executor has corresponding functions according to the performance and interfaces of the corresponding model. Therefore, the models of the corresponding interfaces can be accessed targeted according to different executors.
[0135] Based on the correspondence between the access protocol, model name, model source and the model, the target model to be accessed can be determined.
[0136] For example, in the above example, according to the values of any one or more of the first target parameters "protocol", "aiSource", "modelName", the target model to be accessed can be determined as "Google", or "hunyuan-pro", or "gemini-1.5-pro".
[0137] Then, based on the correspondence between the first target parameter and the executor, which target executor needs to be used currently can be determined from multiple executors.
[0138] S300: Use the target executor to execute the target operation process and conduct a conversation interaction with the target model.
[0139] Specifically, the same model can handle various different user requirements, and different data may need to be sent to the model for different user access requests. To ensure that users can use various functions of the model, in this embodiment, the same executor can have one running process or multiple running processes, which are specifically configured according to the actual application scenario, and this application does not limit this. The steps or running rules executed by each running process may vary. For example, there are differences in the requests sent to the model, the data returned by the model, the processing flow of the executor, etc., so as to achieve the maximum utilization of the model.
[0140] The target executor's dialogue interaction with the target model is equivalent to calling the target model through the API interface.
[0141] This solution solves the problem of API interface differences between major manufacturers and cloud platforms' large language models (LLMs). Using the general framework, users can use a set of protocol specifications to access the large language model interfaces of any manufacturer or cloud platform, achieving efficient docking with various large language models.
[0142] In this embodiment, by setting executors corresponding to different models, it can handle user access requests for any model, solve the problems of incompatibility and differences in different model interfaces, and improve the access efficiency to various large language models. This solution is applicable to scenarios of docking various large language model interface APIs of different manufacturers, different platforms, different models, and different versions.
[0143] In one embodiment, the target running process is one of a non-streaming without plug-in running process, a non-streaming with plug-in running process, a streaming without plug-in running process, and a streaming with plug-in running process;
[0144] Before using the target executor to execute the target running process to conduct dialogue interaction with the target model in step S300, the access method for the large language model further includes:
[0145] Obtain the access path corresponding to the user access request;
[0146] If the second target parameter indicating the use of streaming is not extracted from the access path, and the third target parameter indicating the use of a plug-in is not extracted from the user access request or the value of the third target parameter is empty, then determine to execute the non-streaming without plug-in running process;
[0147] If the second target parameter indicating the use of streaming is not extracted from the access path, and the third target parameter indicating the use of a plug-in is extracted from the user access request and the value of the third target parameter is not empty, then determine to execute the non-streaming with plug-in running process;
[0148] If a second target parameter for indicating the use of streaming is extracted from the access path, and a third target parameter for indicating the use of a plug-in is not extracted from the user access request or the value of the third target parameter is empty, then it is determined to execute a streaming without plug-in operation process;
[0149] If a second target parameter for indicating the use of streaming is extracted from the access path, and a third target parameter for indicating the use of a plug-in is extracted from the user access request and the value of the third target parameter is not empty, then it is determined to execute a streaming with plug-in operation process.
[0150] Specifically, the second target parameter is used to indicate whether to use streaming processing. If streaming processing is used, then the second target parameter exists in the access path. For example, the second target parameter is "streaming". If streaming processing is not used, then the second target parameter does not exist in the access path, for example, "streaming" does not exist.
[0151] The user can specify a plug-in in the human-computer interaction interface, or, if the user does not specify a plug-in, then the default plug-in is used or no plug-in is used.
[0152] If a plug-in is used, then a third target parameter exists in the user access request. For example, the third target parameter is the field "plugins". If no plug-in is used, then the third target parameter does not exist in the user access request. For example, the field "plugins" does not exist.
[0153] Or, if a plug-in is used, then a third target parameter exists in the user access request and the value of the third target parameter is not empty; if no plug-in is used, then the value of the third target parameter in the user access request is empty.
[0154] By comprehensively judging the second target parameter in the access path and the third target parameter in the user access request, it can be determined which one of all the operation processes supported by the target executor the target operation process corresponding to the current user access request is. The target operation process is specifically one of a non-streaming without plug-in operation process, a non-streaming with plug-in operation process, a streaming without plug-in operation process, and a streaming with plug-in operation process.
[0155] It should be noted that each executor includes at least one of a non-streaming without plug-in operation process, a non-streaming with plug-in operation process, a streaming without plug-in operation process, and a streaming with plug-in operation process. It is specifically configured or updated according to the type, version, or model of the model. This application does not limit this.
[0156] Each operation process is determined according to the actual user access request to meet different user needs and achieve the maximum utilization of the model.
[0157] In one embodiment, in step S300, using the target executor to execute the target operation process to conduct a dialogue interaction with the target model includes:
[0158] Using the target executor to take the request parameters in the user access request as the original request parameters;
[0159] According to the original request parameters and the operation rules corresponding to the target operation process, reconstruct the request parameters, and call the dialogue engine to initiate a dialogue request to the target model based on the reconstructed request parameters, where the dialogue request satisfies the operation rules corresponding to the target operation process;
[0160] Obtain the response result returned by the target model based on the dialogue request, and take the response result as the data to be parsed;
[0161] Parse the data to be parsed to obtain a parsing result;
[0162] If it is determined according to the parsing result that the data to be parsed carries a target identifier indicating the end of the request, then integrate the parsing result and return it to the user.
[0163] Specifically, the user access request is an access request sent by the user to the access system of the large language model for any model according to the open interface or the general interface. The access system of the large language model is connected to various different models, and the access interfaces (APIs) of each model may be different. Therefore, it is necessary to convert the general user access request into a dialogue request for a specific model.
[0164] Therefore, the access system of the large language model needs to reconstruct the request parameters and initiate a dialogue request to the target model based on the reconstructed request parameters.
[0165] Each operation process corresponds to a dialogue method, and each operation process has corresponding operation rules. According to the target operation rules of the target operation process, it is possible to determine which parameters need to be reconstructed to obtain the reconstructed request parameters that meet the access requirements of the target model, and based on the reconstructed request parameters, call the dialogue engine to initiate a dialogue request to the target model.
[0166] Among them, calling the dialogue engine to initiate a dialogue request to the target model is actually to call the API interface of the target model.
[0167] More specifically, for example, several dialogue request templates are preset in the target operation rules. According to the request parameters in the user access request, the current target dialogue request template can be determined from multiple dialogue request templates.
[0168] The target dialogue request template records which key parameters are required to generate a complete dialogue request. These key parameters include known key parameters that can be directly extracted from the user access request, and may also include new key parameters that need to be reconstructed.
[0169] Based on this, known key parameters that can be directly extracted and the values of the known key parameters can be matched from the request parameters of the user access request.
[0170] The target running rule may further include a reconstruction rule on how to construct new key parameters according to the known key parameters in the user access request, and / or a reconstruction rule on how to assign values to the new key parameters.
[0171] For key parameters that do not exist in the user access request, new key parameters are constructed according to the reconstruction rule and the known key parameters.
[0172] The obtained key parameters are reconstructed (i.e., concatenated) and the dialogue engine is called to send a dialogue request to the target model. Alternatively, the obtained key parameters are transmitted to the dialogue engine, and the dialogue engine reconstructs or concatenates the key parameters to obtain a dialogue request and sends the dialogue request to the target model.
[0173] After receiving the dialogue request, the target model responds to the dialogue request and returns a response result to the target executor through the dialogue engine.
[0174] Since user access requests are diverse, the target model may return the final response result after receiving the first dialogue request, or may need to go through multiple rounds of dialogue interactions with the target executor to obtain the final response result.
[0175] Based on this, the response result is parsed. If it is determined according to the obtained parsing result that the response result carries a target identifier, it indicates that the response ends or the request ends, and the response result is the final result given by the target model.
[0176] Among them, the target identifier is, for example, the "end" identifier. If the "end" identifier is carried in the response result, it indicates that the request ends, and the response result given by the target model this time is the final response result.
[0177] The target executor can integrate the parsing result of the response result and return it to the user.
[0178] The user can view the final reply returned by the target model on the human-computer interaction interface.
[0179] In this embodiment, a dialogue request that conforms to the API interface call rule of the target model can be sent to the target model through parameter reconstruction, realizing a barrier-free dialogue with the target model, solving the compatibility problem of model interface calls. For users, the differences between various large language model interfaces are shielded, enabling users to efficiently dock and use various large language models.
[0180] In one embodiment, when using the target executor to execute the target operation process to conduct a dialogue interaction with the target model in step S300, it further includes:
[0181] If it is determined according to the parsing result that the target identifier is not carried in the data to be parsed, relevant operations are performed according to the parsing result to obtain an intermediate result;
[0182] Taking the parameters in the intermediate result as the original request parameters, performing the steps of reconstructing the request parameters according to the original request parameters and the operation rules corresponding to the target operation process, calling the dialogue engine to initiate a dialogue request to the target model, and subsequent steps to parse the data to be parsed to obtain a parsing result until it is determined according to the parsing result that the target identifier indicating the end of the request is carried in the data to be parsed;
[0183] The obtained parsing result is integrated and then returned to the user.
[0184] Specifically, for the target model, although its function is powerful, inevitably, sometimes the target model needs the executor to cooperate to provide some intermediate data so that the target model can accurately give the final response result.
[0185] For example, the user asks the target model "What's the weather like today? Can you recommend a travel strategy suitable for today's weather?"
[0186] The weather is changing all the time. Without the assistance of external tools, the target model may not know what the weather is like today, may not know the user's city or specific location, and may not know the surrounding environment of the user's location. Therefore, before outputting a travel strategy to the user, the target model may need the target executor to help query today's weather, obtain the user's location, and obtain information such as the surrounding environment of the user's location.
[0187] Thus, it can be seen that the response result returned by the target model may not necessarily be the final response result, but may be an intermediate response result.
[0188] If the response result is parsed and it is determined according to the obtained parsing result that the response result does not carry the target identifier, for example, does not carry the "end" identifier, it means that the response or the request is not ended, and this response result is the intermediate response result given by the target model.
[0189] After the target model returns the intermediate response result to the target executor through the dialogue engine, the target executor assists the target model to perform corresponding operations according to the intermediate response result to obtain an intermediate result.
[0190] The intermediate result contains some parameters that need to be transmitted to the target model. Therefore, the parameters in the intermediate result need to be used as the original request parameters, and the running rules corresponding to the original request parameters and the target running process need to be re-executed to reconstruct the request parameters. Then, the dialogue engine is called to initiate a dialogue request to the target model. Among them, the dialogue request satisfies the running rules corresponding to the target running process. The response result returned by the target model based on the dialogue request is obtained and used as the data to be parsed. The data to be parsed is parsed to obtain a parsing result. If it is determined according to the parsing result that the target identifier is not carried in the data to be parsed, relevant operations are performed according to the parsing result to obtain an intermediate result, and the loop is returned to be re-executed until it is determined according to the parsing result that the target identifier indicating the end of the request is carried in the data to be parsed. Then, the parsing result is integrated and returned to the user.
[0191] In this embodiment, when it is determined that the response result is not the final response result, the dialogue engine is re-called to send an intermediate dialogue request to the target model to return the intermediate result to the target model. In this way, through multiple rounds of dialogue, the executor assists the target model to obtain the final accurate response result.
[0192] In one embodiment, performing relevant operations according to the parsing result to obtain an intermediate result includes:
[0193] If it is determined according to the obtained parsing result that a tool call needs to be executed, the corresponding target tool is called, and the obtained tool call result is used as the intermediate result.
[0194] Specifically, if it is determined according to the parsing result that the response result is an intermediate response result, the target model can indicate to the target executor through the intermediate response result what additional data it needs, or what additional data it needs and which target tools to call respectively.
[0195] The target executor can determine what operations need to be performed for the target model according to the parsing result, and then obtain the corresponding intermediate result. And performing relevant operations requires the target executor to call the corresponding target tool. The target tool can be some plugins, etc., such as browser plugins, location plugins, calendar plugins, etc. This application does not limit this.
[0196] Among them, the target tool can be specified by the user in advance. For example, the user specifies the plugin type of the plugin that may be used in the human-computer interaction interface. More specifically, for example, the user can specify the type of the browser plugin, etc. For plugins not specified by the user, default plugins can be used. This application does not limit this.
[0197] The target executor calls the corresponding target tool according to the parsing result, and then returns the tool call result as the intermediate result to the target model by re-calling the dialogue engine.
[0198] For example, if the target model needs to accurately recommend travel strategies suitable for today's weather to the user, it needs to obtain information such as today's weather conditions, the user's specific location, and the environment around the user's location. After the target model obtains the initial conversation request, it returns an intermediate response result to the target executor through the conversation engine. The target executor parses the intermediate response result, calls plugins or tools (such as browser plugins and location plugins) to obtain information such as today's weather conditions, the user's specific location, and the environment around the user's location, and returns the obtained information as an intermediate result to the target model through the conversation engine.
[0199] Based on the information such as today's weather conditions, the user's specific location, and the environment around the user's location obtained by the target model, the algorithm can be started to accurately recommend travel strategies to the user.
[0200] In this embodiment, by calling tools, the target model can be provided with the required additional data to assist the target model in quickly and accurately outputting the final response result.
[0201] In one embodiment, when using the target executor to execute the target operation process for dialogue interaction with the target model in step S300, it further includes:
[0202] If it is determined according to the parsing result that the response result is an abnormal return, corresponding exception handling is performed according to the abnormally returned response result;
[0203] If it is monitored that the exception handling is completed, the steps of using the target executor to use the request parameters in the user access request as the original request parameters and subsequent steps are re-executed.
[0204] Specifically, during the model call process, errors will inevitably occur. For example, problems such as request model failure, tool or plugin call failure, executor internal error, and model error.
[0205] If a tool or plugin call fails, the target executor can send a corresponding error alert to the administrator to indicate that the administrator troubleshoots the problem.
[0206] If it is an internal error of the executor, the execution module where the executor is located can send a corresponding error alert to the administrator to indicate that the administrator troubleshoots the problem.
[0207] If a request model failure or model error occurs, the target executor can re-attempt to execute the target operation process to conduct dialogue interaction with the target model by calling the conversation engine. If the continuous attempt count exceeds the count threshold, a corresponding error alert can be sent to the administrator to indicate that the administrator troubleshoots the problem.
[0208] The target actuator or execution module also monitors the completion of exception handling. When it is determined that the exception handling is completed, it will re-execute the step of using the target actuator to use the request parameters in the user access request as the original request parameters; according to the original request parameters and the running rules corresponding to the target running process, reconstruct the request parameters, and call the dialogue engine to initiate a dialogue request to the target model, where the dialogue request meets the running rules corresponding to the target running process; obtain the response result returned by the target model based on the dialogue request, and use the response result as the data to be parsed; parse the data to be parsed to obtain a parsing result; if it is determined according to the parsing result that the data to be parsed carries a target identifier indicating the end of the request, then integrate the parsing result and return it to the user, etc.
[0209] In addition, if an error occurs, the target model may also return a response result containing an error code (or error type), etc.; according to the error code in this response result, the target actuator can determine the error type and can also find the corresponding solution.
[0210] Among them, the solutions are sorted and classified for errors in advance and stored in the actuator. The actuator can match and find the solutions through the error code or error type.
[0211] Alternatively, the target actuator can also obtain the solution corresponding to the error code from the target model through dialogue interaction, and the target model has been pre-trained according to the error type or error code and the corresponding solution.
[0212] In addition, after multiple attempts to handle exceptions, if it still reports an error, a feedback result indicating the exception or error can be returned to the user, and this feedback result can also be used to indicate to the user to re-select and use other large language models. This enables the user to select a suitable and normal model from multiple models for invocation, and timely meets the user's usage requirements.
[0213] This embodiment can automatically and timely perform corresponding exception handling on the response result returned by the exception, ensuring the timely response to the user request.
[0214] Reference Figure 2 , this application also provides an access system for a large language model. The access system for the large language model includes: a request distribution module 100, an execution module 200, and a dialogue engine 300. The execution module 200 includes at least one actuator 210;
[0215] The request distribution module 100 is used to extract the first target parameters from the obtained user access request, where the first target parameters include at least one of an access protocol, a model identifier of the target model to be accessed, and a model source;
[0216] The request distribution module 100 is further configured to distribute the user access request to a target executor that matches the first target parameter;
[0217] The execution module 200 is configured to execute a target operation process by using the target executor and perform a dialogue interaction with the target model through the dialogue engine 300.
[0218] Specifically, referring to Figure 2 , the access system of the large language model includes: a request distribution module 100, an execution module 200, and a dialogue engine 300 that are connected in sequence. The execution module 200 includes at least one executor 210 (i.e., executor 1, executor 2... executor n); each executor 210 can interact with the corresponding model through the dialogue engine 300. Among them, the models include model 1, model 2... model m. Some models may share an executor.
[0219] In a specific embodiment, the request distribution module 100 may be a router. The router is essentially a factory class, and it can distribute the user access request to the corresponding executor according to the model type or protocol of the user access request.
[0220] The executor 210 (Executor): is responsible for processing specific requests and dialogue returns. As Figure 3 shown, the executor 210 includes four operation processes, and each corresponds to a type of request method.
[0221] In an embodiment, the target operation process is one of a non-streaming and plug-in-free operation process, a non-streaming and plug-in operation process, a streaming and plug-in-free operation process, and a streaming and plug-in operation process;
[0222] The target executor is specifically configured to:
[0223] Obtain the access path corresponding to the user access request;
[0224] If the second target parameter indicating the use of streaming is not extracted from the access path, and the third target parameter indicating the use of a plug-in is not extracted from the user access request, then determine to execute a non-streaming and plug-in-free operation process;
[0225] If the second target parameter indicating the use of streaming is not extracted from the access path, and the third target parameter indicating the use of a plug-in is extracted from the user access request, then determine to execute a non-streaming and plug-in operation process;
[0226] If the second target parameter indicating the use of streaming is extracted from the access path, and the third target parameter indicating the use of a plug-in is not extracted from the user access request, then determine to execute a streaming and plug-in-free operation process;
[0227] If a second target parameter for indicating the use of streaming is extracted from the access path, and a third target parameter for indicating the use of a plug-in is extracted from the user access request, then it is determined to execute the streaming with plug-in operation process.
[0228] Specifically, Figure 3 It is a schematic diagram of the functional modules of another embodiment of the access system of the large language model in the embodiments of the present application; refer to Figure 3 , an actuator 210 includes a total of 4 operation processes: a non-streaming without plug-in operation process, a non-streaming with plug-in operation process, a streaming without plug-in operation process, and a streaming with plug-in operation process. After receiving the user access request distributed by the request distribution module 100, the actuator 210 determines a target operation process from the 4 operation processes, and executes the target operation process to interact with the target model through the dialogue engine 300 (Engine).
[0229] In one embodiment, the target actuator is specifically used for:
[0230] Take the request parameters in the user access request as the original request parameters;
[0231] According to the original request parameters and the operation rules corresponding to the target operation process, reconstruct the request parameters, and call the dialogue engine 300 to initiate a dialogue request to the target model, where the dialogue request satisfies the operation rules corresponding to the target operation process;
[0232] Obtain the response result returned by the target model based on the dialogue request, and take the response result as the data to be parsed;
[0233] Parse the data to be parsed to obtain a parsing result;
[0234] If it is determined according to the parsing result that the data to be parsed carries a target identifier indicating the end of the request, then integrate the parsing result and return it to the user.
[0235] In one embodiment, the target actuator is further used for:
[0236] If it is determined according to the parsing result that the data to be parsed does not carry the target identifier, then perform relevant operations according to the parsing result to obtain an intermediate result;
[0237] Take the parameters in the intermediate result as the original request parameters, execute the steps of reconstructing the request parameters according to the original request parameters and the operation rules corresponding to the target operation process, calling the dialogue engine 300 to initiate a dialogue request to the target model, and subsequent steps, until it is determined according to the parsing result that the data to be parsed carries a target identifier indicating the end of the request;
[0238] Integrate the obtained parsing result and return it to the user.
[0239] In one embodiment, the target executor or execution module is specifically configured to:
[0240] If it is determined according to the obtained parsing result that a tool call needs to be executed, then call the corresponding target tool, and use the obtained tool call result as an intermediate result.
[0241] In one embodiment, the target executor is specifically configured to:
[0242] If it is determined according to the parsing result that the response result is an abnormal return, then perform corresponding exception handling according to the abnormally returned response result;
[0243] If it is monitored that the exception handling is completed, then re-execute the steps of using the target executor with the request parameters in the user access request as the original request parameters and subsequent steps.
[0244] For the functions of the access system of the large language model, please refer to the description of the above-mentioned access method of the large language model, which will not be elaborated here.
[0245] Figure 4 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 700 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 710 (for example, one or more processors) and a memory 720, and one or more storage media 730 for storing application programs 733 or data 732 (for example, one or more mass storage devices). Among them, the memory 720 and the storage media 730 may be transient storage or persistent storage. The program stored in the storage media 730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device 700. Further, the processor 710 may be configured to communicate with the storage media 730 and execute a series of instruction operations in the storage media 730 on the computer device 700.
[0246] The computer device 700 may further include one or more power supplies 740, one or more wired or wireless network interfaces 750, one or more input / output interfaces 760, and / or one or more operating systems 731, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 4 The shown computer device structure does not limit the computer device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0247] The present application also provides a computer device, which includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor is caused to execute the steps of the method for accessing the large language model in the above embodiments.
[0248] The present application also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the method for accessing the large language model.
[0249] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0250] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0251] The above embodiments are only used to illustrate the technical solution of the present application and are not intended to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A method for accessing a large language model, characterized in that: The access method of the large language model includes: Extracting a first target parameter from the acquired user access request, wherein the first target parameter includes at least one of an access protocol, a model identifier of a target model to be accessed, and a model source; Distributing the user access request to a target executor matching the first target parameter; The target executor is used to execute the target operation process and to interact with the target model.
2. The method for accessing a large language model according to claim 1, characterized in that: The target running process is one of a non-streaming running process without plug-ins, a non-streaming running process with plug-ins, a streaming running process without plug-ins, and a streaming running process with plug-ins; Before using the target executor to execute the target operation process and to interact with the target model, the method for accessing the large language model further includes: Obtaining an access path corresponding to the user access request; If the second target parameter for indicating the use of streaming is not extracted from the access path, and the third target parameter for indicating the use of plug-in is not extracted from the user access request or the value of the third target parameter is empty, it is determined to execute the non-streaming plug-in-free operation process; If the second target parameter for indicating the use of streaming is not extracted from the access path, and the third target parameter for indicating the use of plug-ins is extracted from the user access request and the value of the third target parameter is not empty, it is determined to execute the non-streaming plug-in running process; If a second target parameter for indicating the use of streaming is extracted from the access path, and a third target parameter for indicating the use of a plug-in is not extracted from the user access request or the value of the third target parameter is empty, it is determined to execute the streaming plug-in-free operation process; If a second target parameter for indicating the use of streaming is extracted from the access path, and a third target parameter for indicating the use of a plug-in is extracted from the user access request and the value of the third target parameter is not empty, it is determined that the streaming plug-in is executed.
3. The method for accessing a large language model according to claim 1, characterized in that: The step of utilizing the target executor to execute the target operation process and to interact with the target model in dialogue includes: Using the target executor to use the request parameters in the user access request as original request parameters; Reconstructing the request parameters according to the original request parameters and the operation rules corresponding to the target operation process, and calling the dialog engine to initiate a dialog request to the target model, wherein the dialog request satisfies the operation rules corresponding to the target operation process; Obtaining a response result returned by the target model based on the dialogue request, and using the response result as data to be parsed; Parse the data to be parsed and obtain the parsing result; If it is determined according to the analysis result that the data to be analyzed carries a target identifier indicating the end of the request, the analysis result is integrated and returned to the user.
4. The method for accessing a large language model according to claim 3, characterized in that: The using the target executor to execute the target operation process and to interact with the target model in dialogue also includes: If it is determined according to the parsing result that the data to be parsed does not carry the target identifier, then performing relevant operations according to the parsing result to obtain an intermediate result; The parameters in the intermediate result are used as original request parameters, the operation rules corresponding to the original request parameters and the target operation process are executed, the request parameters are reconstructed, and the dialog engine is called to initiate a dialog request to the target model, and subsequent steps, until the target identifier indicating the end of the request is determined in the data to be parsed according to the parsing result; The obtained analysis results are integrated and returned to the user.
5. The method for accessing a large language model according to claim 4, characterized in that: The performing of related operations according to the analysis result to obtain an intermediate result includes: If it is determined that a tool call needs to be executed based on the obtained parsing result, the corresponding target tool is called and the obtained tool call result is used as an intermediate result.
6. The method for accessing a large language model according to claim 3, characterized in that: The using the target executor to execute the target operation process and to interact with the target model in dialogue also includes: If the response result is determined to be an abnormal return according to the analysis result, corresponding exception handling is performed according to the abnormal response result; If it is monitored that the exception handling is completed, the step of using the target executor to use the request parameters in the user access request as original request parameters and subsequent steps are re-executed.
7. A system for accessing a large language model, characterized in that: The access system of the large language model includes: a request distribution module, an execution module and a dialogue engine, wherein the execution module includes at least one executor; The request distribution module is used to extract a first target parameter from the acquired user access request, wherein the first target parameter includes at least one of an access protocol, a model identifier of a target model to be accessed, and a model source; The request distribution module is further used to distribute the user access request to a target executor matching the first target parameter; The execution module is used to utilize the target executor to execute the target operation process and to interact with the target model through the dialogue engine.
8. The large language model access system according to claim 7, characterized in that: The execution module is further used for: Obtaining an access path corresponding to the user access request; If the second target parameter for indicating the use of streaming is not extracted from the access path, and the third target parameter for indicating the use of plug-ins is not extracted from the user access request, it is determined to execute the non-streaming plug-in-free operation process; If the second target parameter for indicating the use of streaming is not extracted from the access path, and the third target parameter for indicating the use of plug-ins is extracted from the user access request, it is determined to execute the non-streaming plug-in operation process; If a second target parameter for indicating the use of streaming is extracted from the access path, and a third target parameter for indicating the use of a plug-in is not extracted from the user access request, determining to execute a streaming plug-in-free operation process; If a second target parameter for indicating the use of streaming is extracted from the access path, and a third target parameter for indicating the use of a plug-in is extracted from the user access request, it is determined that the streaming plug-in running process is executed.
9. A computer device, characterized in that: The computer device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory to cause the computer device to execute the method for accessing a large language model according to any one of claims 1 to 6.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by a processor, a method for accessing a large language model according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Crosslinker composition including synthetic layered silicate
US10240081B2
Cited By
Server access method for large model and actual service scene
CN120873060A
A server access method for large models and actual business scenarios
CN120873060B