Interaction method and device based on artificial intelligence, intelligent agent, equipment and medium

Large models are routed in the interactive system through an intermediate server, and the target large model is selected based on input information and device configuration. This solves the stability and efficiency problems caused by model routing operations in the interactive system and achieves efficient load balancing and system stability.

CN120744362APending Publication Date: 2025-10-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510837955.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In interactive systems, the large model routing operations affect service stability and efficiency, making it difficult to efficiently select and call multiple large models of different scales or functions.

Method used

Large model routing is performed through an intermediate server. Based on the attributes of the input information and the model configuration information, a candidate large model is determined from multiple large models, and the target large model is selected according to the device configuration information, thereby reducing the routing operations of the interactive system, achieving load balancing and improving system stability.

Benefits of technology

It improves the system stability and efficiency of the interaction process, reduces the burden on the interaction system, and achieves efficient routing and load balancing for multiple large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744362A_ABST
    Figure CN120744362A_ABST
Patent Text Reader

Abstract

The invention provides an interaction method and device based on artificial intelligence, an intelligent agent, equipment, a medium and a product, and relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, large models, man-machine interaction and the like. The specific implementation scheme is as follows: obtaining an interaction request from a client, wherein the interaction request comprises to-be-processed input information and a system identifier for indicating a required interaction system; acquiring model configuration information and equipment configuration information of the interactive system based on the system identifier; determining at least one candidate large model from the plurality of large models based on attribute information of the input information and model configuration information; and based on the device configuration information, determining a target large model from the at least one candidate large model, so that the interaction system calls the target large model to process the input information, and feedback information for the interaction request is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as deep learning, large models, and human-computer interaction, and specifically to interaction methods, devices, intelligent agents, equipment, media, and products based on artificial intelligence. Background Art

[0002] With the continuous development of artificial intelligence technology, large models are widely used. With the rapid development of large models, their types are increasing, including large language models, large visual models, and large multimodal models. When interactive systems integrate multiple large models of different types, model routing technology becomes an important technical direction for optimizing large model inference. Summary of the Invention

[0003] The present disclosure provides an artificial intelligence-based interaction method, device, intelligent agent, electronic device, and storage medium.

[0004] According to one aspect of the present disclosure, an artificial intelligence-based interaction method is provided, including: obtaining an interaction request from a client, the interaction request including input information to be processed and a system identifier indicating a required interaction system; the interaction system is configured with multiple large models; based on the system identifier, model configuration information and device configuration information of the interaction system are obtained; the model configuration information represents the model configuration of each of the multiple large models, and the device configuration information represents the device configuration for running the large model device; based on the attribute information and the model configuration information of the input information, at least one candidate large model is determined from the multiple large models; and based on the device configuration information, a target large model is determined from the at least one candidate large model, so that the interaction system calls the target large model to process the input information and obtains feedback information for the interaction request.

[0005] According to another aspect of the present disclosure, an artificial intelligence-based interaction device is provided, including: a request acquisition module, used to obtain an interaction request from a client, the interaction request including input information to be processed and a system identifier indicating the required interaction system; the interaction system is configured with multiple large models; a configuration acquisition module, used to obtain model configuration information and device configuration information of the interaction system based on the system identifier; the model configuration information represents the model configuration of each of the multiple large models, and the device configuration information represents the device configuration for running the large model device; a candidate determination module, used to determine at least one candidate large model from the multiple large models based on the attribute information and model configuration information of the input information; a target determination module, used to determine a target large model from the at least one candidate large model based on the device configuration information, so that the interaction system calls the target large model to process the input information and obtain feedback information for the interaction request.

[0006] According to another aspect of the present disclosure, an artificial intelligence agent is provided, comprising: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a large model based on the target task, and obtaining output information by calling the large model to execute the method described above; and an output module for outputting the output information obtained by the processing module.

[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described above.

[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method described above when executed by a processor.

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0012] Figure 1 Schematically illustrates an exemplary system architecture to which an artificial intelligence-based interaction method and apparatus according to an embodiment of the present disclosure may be applied;

[0013] Figure 2 The following schematically shows a flow chart of an artificial intelligence-based interaction method according to an embodiment of the present disclosure;

[0014] Figure 3 The following schematically illustrates an interaction process diagram of an artificial intelligence-based interaction method according to an embodiment of the present disclosure;

[0015] Figure 4 The following schematically shows an interface diagram of a client according to an embodiment of the present disclosure;

[0016] Figure 5A The following schematically illustrates a method for determining a candidate large model according to an embodiment of the present disclosure;

[0017] Figure 5B Schematically shows a schematic diagram of determining a candidate large model according to another embodiment of the present disclosure;

[0018] Figure 6A The following schematically shows a schematic diagram of determining a target macro model according to an embodiment of the present disclosure;

[0019] Figure 6B Schematically shows a schematic diagram of determining a target macro model according to another embodiment of the present disclosure;

[0020] Figure 7 Schematically shows a flow chart for determining a target macro model according to an embodiment of the present disclosure;

[0021] Figure 8 A block diagram of an artificial intelligence-based interaction device according to an embodiment of the present disclosure is schematically shown;

[0022] Figure 9 A block diagram schematically illustrates a structure of an artificial intelligence agent according to an embodiment of the present disclosure; and

[0023] Figure 10 A block diagram of an electronic device suitable for implementing an artificial intelligence-based interaction method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0024] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0025] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0026] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0027] Currently, large models are experiencing an unprecedented explosion in popularity and have become one of the core drivers of the AI ​​industry. With the continuous introduction of large models, they are widely used in a variety of fields, including intelligent question-answering, content generation, programming assistance, search enhancement, and multimodal interaction. Model routing technology, a key area of ​​optimization for large model inference, primarily addresses the efficient selection and invocation of multiple large models of varying scale and functionality, thereby improving performance, reducing costs, and enhancing responsiveness.

[0028] Model routing strategies typically involve the interactive system routing large models based on rules such as input length, keywords, user type, and task type or difficulty. While implementing the inventive concepts of this disclosure, the inventors discovered that routing by the interactive system places a significant burden on the system, potentially impacting service stability and efficiency.

[0029] In view of this, an embodiment of the present disclosure provides an artificial intelligence-based interaction method, including: obtaining an interaction request from a client, the interaction request including input information to be processed and a system identifier indicating the required interaction system; the interaction system is configured with multiple large models; based on the system identifier, model configuration information and device configuration information of the interaction system are obtained; the model configuration information represents the model configuration of each of the multiple large models, and the device configuration information represents the device configuration for running the large model device; based on the attribute information and model configuration information of the input information, at least one candidate large model is determined from the multiple large models; and based on the device configuration information, a target large model is determined from the at least one candidate large model, so that the interaction system calls the target large model to process the input information and obtain feedback information for the interaction request. The embodiment of the present disclosure routes the large model through an intermediate server, avoids the routing operation of the interaction system, reduces the burden of the interaction system, and achieves a balanced load of the interaction system, thereby improving the system stability and interaction efficiency of the entire interaction process.

[0030] Figure 1 An exemplary system architecture to which an artificial intelligence-based interaction method and apparatus can be applied according to an embodiment of the present disclosure is schematically illustrated.

[0031] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not imply that the embodiments of the present disclosure may not be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the artificial intelligence-based interaction method and apparatus may be applied may include a terminal device, but the terminal device may implement the artificial intelligence-based interaction method and apparatus provided by the embodiments of the present disclosure without interacting with a server.

[0032] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, an intermediate server 105, and interactive systems 106 and 107. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, and 103 and the intermediate server 105. The network 104 may include various connection types, such as wired and / or wireless communication links. The interactive system 106 may include large models 1061 and 1062, and the interactive system 107 may include a large model 1071.

[0033] Users can use terminal devices 101, 102, and 103 to interact with the intermediate server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).

[0034] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0035] Intermediary server 105 can provide interactive services provided by various interactive systems and can route multiple large models configured for different interactive systems. For example, it can provide interactive services provided by interactive systems 106 and 107 and route large models 1061 and 1062 configured for interactive system 106, as well as large model 1072 configured for interactive system 107 (for example only). Intermediary server 105 can route interactive requests from terminal devices 101, 102, and 103 and send them to large models 1061, 1062, or 1072. It can also return feedback information generated by large models 1061, 1062, or 1072 in response to the interactive requests to terminal devices 101, 102, and 103.

[0036] The interactive systems 106 and 107 can provide interactive services through the configured large models. The large models 1061 and 1062 configured for the interactive system 106 and the large model 1072 configured for the interactive system 107 can be a combination of one or more of a large language model (LLM), a large vision model (LVM), or a multimodal large model (MLM).

[0037] The artificial intelligence-based interaction method provided in the embodiment of the present disclosure can also generally be executed by the intermediate server 105. Accordingly, the artificial intelligence-based interaction device provided in the embodiment of the present disclosure can generally be set in the intermediate server 105. The artificial intelligence-based interaction method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the intermediate server 105 and can communicate with the terminal devices 101, 102, 103 and / or the intermediate server 105 and / or the interaction systems 106, 107. Accordingly, the artificial intelligence-based interaction device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the intermediate server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105 and / or the interaction systems 106, 107.

[0038] For example, when a user needs to interact with an interactive system, terminal devices 101, 102, and 103 can generate an interaction request based on the user's input information and the system identifier of the desired interactive system, and then send the interaction request to the intermediate server 105. The intermediate server 105 obtains the model configuration information and device configuration information of the interactive system based on the system identifier; determines at least one candidate large model from multiple large models based on the attribute information and model configuration information of the input information; and determines a target large model from the at least one candidate large model based on the device configuration information, so that the interactive systems 106 and 107 call the target large model to process the input information and obtain feedback information regarding the interaction request. After the interactive systems 106 and 107 obtain the feedback information regarding the interaction request, the intermediate large model 105 receives the feedback information regarding the interaction request from the interactive systems 106 and 107 and sends it to the terminal devices 101, 102, and 103.

[0039] It should be understood that Figure 1 The number of terminal devices, networks, intermediate servers and interactive systems in the embodiment is merely illustrative. Any number of terminal devices, networks, intermediate servers and interactive systems may be provided as required.

[0040] It should be noted that the sequence numbers of the operations in the following method are only used to indicate the operation for the purpose of description, and should not be regarded as indicating the order in which the operations should be performed. Unless explicitly stated, the method does not need to be performed in the order shown.

[0041] Figure 2 The flowchart of the artificial intelligence-based interaction method according to an embodiment of the present disclosure is schematically shown.

[0042] like Figure 2 As shown, the method includes operations S210 to S240.

[0043] In operation S210, an interaction request from a client is obtained, where the interaction request includes input information to be processed and a system identifier indicating a required interaction system; the interaction system is configured with a plurality of large models.

[0044] In operation S220, based on the system identifier, model configuration information and device configuration information of the interactive system are obtained; the model configuration information represents the model configuration of each of the multiple large models, and the device configuration information represents the device configuration for running the large model device.

[0045] In operation S230, at least one candidate large model is determined from a plurality of large models based on the attribute information and the model configuration information of the input information.

[0046] In operation S240 , a target large model is determined from at least one candidate large model based on the device configuration information, so that the interactive system calls the target large model to process input information and obtain feedback information for the interactive request.

[0047] In embodiments of the present disclosure, an artificial intelligence-based interaction method can be applied to an intermediate server, which can provide a client with interactive services provided by multiple different interactive systems. Accordingly, the client can provide an interactive interface corresponding to each of the multiple interactive systems, and the interactive interface corresponding to different interactive systems can be different.

[0048] When the user needs to interact with the interactive system, he can enter the client in the terminal device where the client is installed, enter the interactive interface for interacting with the desired interactive system in the client interface, and enter the input information to be processed in the interactive interface.

[0049] After the user completes the input, the client can generate an interaction request based on the input information and the system identifier of the interactive system, and send the interaction request to the intermediate server so that the intermediate server can interact with the interactive system required by the user and process the input information to be processed through multiple large models configured in the interactive system.

[0050] Since the model configurations and device configurations of multiple large models in the interactive system may be different, different large models are suitable for processing input information with different attributes. Therefore, an intermediate server can be used to screen out the target large model that is most suitable for processing the input information from multiple large models to implement routing operations on multiple large models, thereby avoiding routing operations on the interactive system and reducing the burden on the interactive system.

[0051] The attribute information of the input information may include the intent of the input information and the length of the information. Correspondingly, the model configuration information may include the functions of the large model and the length of information that it is suitable for processing. When routing multiple large models using an intermediate server, the attribute information of the input information and the model configuration information of the large models may be used to first filter out candidate large models from the multiple large models that can process the input information.

[0052] For example, when the attribute information indicates that the user's intention includes image recognition, the candidate large model determined should have image recognition capabilities. For another example, when the attribute information indicates that the input information is long, the candidate large model determined should be suitable for processing long input information.

[0053] After determining the candidate large models, the device configuration information of at least one candidate large model can be used to screen out the target large model with better device status, lower load or higher resource configuration from the candidate large models, thereby achieving load balancing of multiple large models in the interactive system, thereby making the target large model more efficient in processing input information and improving the overall efficiency of the interactive process.

[0054] For example, the hardware resource configuration of the target large model may be higher than the hardware resource configuration of other candidate large models, or the load of the target large model may be lower than the load of other candidate large models.

[0055] After determining the target large model, the intermediate server may send the input information and the model identifier of the target large model to the interactive system to instruct the interactive system to call the target large model to process the input information and obtain feedback information.

[0056] According to the embodiments of the present disclosure, an intermediate server is utilized to provide clients with interactive services provided by a variety of different interactive systems, and multiple large models configured for each interactive system can be routed, thereby improving the breadth of services through multiple interactive systems and multiple large models provided by each interactive system. The large models are also routed through the intermediate server to avoid routing operations of the interactive system, reduce the burden on the interactive system, and achieve balanced load on the interactive system, thereby improving the system stability and interaction efficiency of the entire interactive process.

[0057] Reference below Figures 3 to 7 , combined with specific embodiments Figure 2 The method shown is further explained.

[0058] Figure 3 The interactive flow diagram of the artificial intelligence-based interactive method according to an embodiment of the present disclosure is schematically shown.

[0059] like Figure 3 As shown, the interaction process of the interaction method based on artificial intelligence includes operations S310 to S340.

[0060] In operation S310 , an interaction request is sent.

[0061] The client can send an interaction request to the intermediate server, where the interaction request includes input information to be processed and a system identifier indicating the desired interaction system. The interaction system is configured with multiple large models.

[0062] In operation S320, a target macro model is determined from among a plurality of macro models configured in the interactive system.

[0063] After receiving the interaction request, the intermediate server can determine the target large model from multiple large models.

[0064] In operation S330 , the input information and the model identification of the target large model are sent.

[0065] After determining the target large model, the intermediate server may send the input information and the model identifier of the target large model to the interactive system.

[0066] In operation S340, the target large model is called to process the input information and obtain feedback information.

[0067] After obtaining the input information and the model identifier of the target large model, the interactive system calls the target large model to process the input information and obtain feedback information.

[0068] According to an embodiment of the present disclosure, the artificial intelligence-based interaction method further includes: receiving feedback information for the interaction request from the interaction system and sending it to the client.

[0069] Since the interaction request is sent to the interaction system by the intermediate server, the interaction system will send the feedback information to the intermediate server after generating the feedback information for the interaction request. At this time, the intermediate server can send the feedback information to the client so that the user can view the feedback information.

[0070] After receiving the feedback information sent by the intermediate server, the client can display the feedback information on the interactive interface. In some embodiments, the input information and the feedback information can be displayed on the interactive interface in the form of a dialogue.

[0071] Figure 4 The following schematically shows an interface diagram of a client according to an embodiment of the present disclosure.

[0072] like Figure 4 As shown, the client interface includes a system selection interface 410 and an interaction interface 420 .

[0073] like Figure 4As shown, in the system selection interface 410, multiple interactive system selection buttons may be displayed, such as a first interactive system selection button 411, a second interactive system selection button 412, etc. The user may click the selection button of the desired interactive system to enter the interactive interface for interacting with the interactive system.

[0074] For example, the user clicks the selection button 411 to instruct interaction with the first interactive system and enters the interactive interface for interacting with the first interactive system.

[0075] like Figure 4 As shown, the interactive interface 420 of the first interactive system may include an interactive process sub-interface 421, in which the user's input information 4211 "What day is today?" and the corresponding feedback information 4212 "Today is Friday." are displayed.

[0076] like Figure 4 As shown, the interactive interface 420 may further include an input box 422, and the user may input information in the input box 422. After the user completes the input, the user may click a send button 423 to process the input information using the first interactive system.

[0077] According to an embodiment of the present disclosure, the attribute information includes intent information, and the model configuration information includes function information. Determining at least one candidate large model from a plurality of large models based on the attribute information and the model configuration information of the input information includes: determining a plurality of initial candidate large models that can process the input information from the plurality of large models based on the intent information and the function information; and determining at least one candidate large model from the plurality of initial candidate large models based on the model routing strategy and the model configuration information.

[0078] Among the multiple large models in an interactive system, different large models may have different functions. For example, multiple large models may include a large model for image recognition and a large model for content polishing. Therefore, to ensure that the final selected target large model matches the user's needs, the functional information of the large models can be used to perform a preliminary screening of multiple large models using the input information's intent information to obtain initial candidate large models whose functional information matches the input information's intent information.

[0079] The initial candidate large models can be large models whose functional information matches the intent of the input information. Different initial candidate large models can have different model configuration information. For example, if the intent information indicates that the user needs to polish the article, the multiple initial candidate large models determined accordingly should all have content polishing capabilities.

[0080] Since the model configuration information of different initial candidate large models may be different, such as model structure, training data, and training rounds, which leads to different processing efficiency and resource consumption of different initial candidate large models, after determining multiple initial candidate large models, the model routing strategy and the model configuration information of the multiple initial candidate large models can be used to further screen the multiple initial candidate large models to obtain the candidate large models whose model configuration information matches the model routing strategy.

[0081] The model routing strategy may include strategies such as shortest time consumption and least resource consumption. Accordingly, the model configuration information may include the processing time and resource consumption of each of the multiple initial candidate large models.

[0082] For example, when the model routing strategy is the shortest time-consuming strategy, the processing time of the determined candidate large model is less than the processing time of other initial candidate large models.

[0083] According to the embodiments of the present disclosure, by routing the large model according to the intent information and the model routing strategy, the model configuration information of the candidate large model is matched with the intent information and the model routing strategy, which not only improves the adaptability of the candidate large model to the input information, but also eliminates the need for the interactive system to perform routing operations, thereby improving the stability and efficiency of the interactive process.

[0084] According to an embodiment of the present disclosure, the model configuration information also includes time consumption information. Determining at least one candidate large model from a plurality of initial candidate large models based on the model routing strategy and the model configuration information includes: determining at least one candidate large model from the plurality of initial candidate large models based on the time consumption information of each of the plurality of initial candidate large models when the model routing strategy characterizes time-consuming routing.

[0085] Time-based routing usually considers the processing efficiency of large models. Therefore, based on the time-consuming information of multiple initial candidate large models, the initial candidate large model with the shortest time consumption can be selected as the candidate large model, or the initial candidate large model with a time consumption less than a preset time consumption threshold can be selected as the candidate large model.

[0086] According to an embodiment of the present disclosure, by selecting a candidate large model based on a time-consuming route, the processing efficiency of the determined candidate large model is high, thereby improving the efficiency of the interactive system in processing input information.

[0087] Figure 5A The figure schematically shows a schematic diagram of determining a candidate large model according to an embodiment of the present disclosure.

[0088] like Figure 5AAs shown, the initial candidate large models include a first large model 10, a second large model 20, a third large model 30, and a fourth large model 40. The time consumption information of the first large model 10 is short, the time consumption information of the second large model 20 is long, the time consumption information of the third large model 30 is short, and the time consumption information of the fourth large model 40 is short.

[0089] like Figure 5A As shown, when the model routing strategy characterization is based on time-consuming routing, the candidate large models determined from the first large model 10 , the second large model 20 , the third large model 30 and the fourth large model 40 are the first large model 10 , the third large model 30 and the fourth large model 40 .

[0090] According to an embodiment of the present disclosure, the model configuration information also includes resource consumption information. Determining at least one candidate large model from a plurality of initial candidate large models based on the model routing policy and the model configuration information includes: determining at least one candidate large model from the plurality of initial candidate large models based on the resource consumption information of each of the plurality of initial candidate large models when the model routing policy indicates routing based on resource consumption.

[0091] Resource consumption-based routing usually considers the resource consumption of large models. Therefore, based on the resource consumption information of multiple initial candidate large models, the initial candidate large model with the least resource consumption can be selected as the candidate large model, or the initial candidate large model with resource consumption less than a preset resource consumption threshold can be selected as the candidate large model.

[0092] According to an embodiment of the present disclosure, by selecting a candidate large model based on resource consumption routing, the resource consumption of the determined candidate large model is reduced, thereby reducing the resource consumption of the interactive system in processing input information.

[0093] The above Figure 5A The process of determining the candidate large model based on time-consuming routing is detailed below. Figure 5B The process of determining candidate large models based on resource consumption routing is described in detail.

[0094] Figure 5B The figure schematically shows a schematic diagram of determining a candidate large model according to another embodiment of the present disclosure.

[0095] like Figure 5B As shown, the initial candidate large models include a first large model 10, a second large model 20, a third large model 30, and a fourth large model 40. Among them, the resource consumption information of the first large model 10 is large resource consumption, the resource consumption information of the second large model 20 is large resource consumption, the resource consumption information of the third large model 30 is small resource consumption, and the resource consumption information of the fourth large model 40 is large resource consumption.

[0096] like Figure 5B As shown, when the model routing policy characterizes routing based on resource consumption, the candidate large model determined from the first large model 10 , the second large model 20 , the third large model 30 , and the fourth large model 40 is the third large model 30 .

[0097] In other embodiments, in order to improve the processing speed of large models and reduce the resource consumption of large models, and to realize time-based routing and resource consumption-based routing at the same time, a large bucket service and a small bucket service will be deployed at the same time for large models with the same function. The large bucket service supports longer user input and can output longer content, while the small bucket service supports shorter user input, consumes less time, occupies less Graphics Processing Unit (GPU) resources, and has a cost advantage.

[0098] When users interact with an interactive system, the length of their input and the length of the required feedback can vary significantly. For requests with shorter input and no need for long feedback, the intermediate processing server prioritizes the initial candidate large model from the small bucket service, reducing both processing time and resource consumption.

[0099] According to an embodiment of the present disclosure, determining a target big model from at least one candidate big model based on device configuration information includes: determining the target big model from at least one candidate big model based on the device configuration information and a device routing policy.

[0100] Since the device configuration information of different candidate large models may be different, such as different device status and device resource configuration, which leads to different processing efficiency of different candidate large models, after determining multiple candidate large models, the device routing strategy and the device configuration information of the multiple candidate large models can be used to screen the multiple candidate large models to obtain the target large model whose device configuration information matches the device routing strategy.

[0101] The device routing strategy may include strategies such as the device being in a normal state and the device having the best resource configuration. Accordingly, the device configuration information may include the states and resource configurations of the respective multiple candidate large models.

[0102] For example, when the device routing strategy is the highest resource configuration strategy, the resource configuration of the device running the determined target large model is higher than the resource configuration of other candidate large models, such as the number of GPU cores of the device running the target large model is greater than the number of GPU cores of the devices running other candidate large models.

[0103] According to the embodiments of the present disclosure, by performing further routing operations on the candidate large models based on the device routing strategy, the device configuration of the determined target large model is higher, which not only improves the processing efficiency of processing input information using the target large model, but also eliminates the need for the interactive system to perform routing operations, thereby improving the stability and efficiency of the interactive process.

[0104] According to an embodiment of the present disclosure, device configuration information includes device status information. Determining a target large model from at least one candidate large model based on the device configuration information and a device routing policy includes: obtaining device status information of a device running a candidate large model from an interactive system to obtain at least one piece of device status information, when the device routing policy indicates routing based on device status; and determining the target large model from the at least one candidate large model based on the at least one piece of device status information.

[0105] The device status information may represent the operating status of the operating device of the large model, such as fault status, normal status, etc. When obtaining the device status information, the intermediate server may send the model identifiers of the plurality of candidate large models to the interactive system to query the device status information of the plurality of candidate large models from the interactive system.

[0106] Routing based on device status usually considers whether the device running the large model is faulty, so as to avoid the interactive system calling the large model in the faulty state to process the input information, which reduces the processing efficiency. Therefore, based on the device status information of multiple candidate large models, the candidate large model whose running device is in normal state can be selected as the target large model.

[0107] According to an embodiment of the present disclosure, by selecting a target macro model based on device status routing, the operating device of the determined target macro model is in a normal state and can process input information normally, thereby ensuring the interaction efficiency of the interaction process.

[0108] According to an embodiment of the present disclosure, device status information includes fault information and load information. Based on at least one piece of device status information, a target large model is determined from at least one candidate large model, including: determining a distributed device that is scheduled to serve a client from a distributed device cluster of an interactive system, the distributed device being configured with the candidate large model; and, if fault information from the distributed device indicates an abnormality in the distributed device, determining the target large model based on load information of other distributed devices in the distributed cluster; the other distributed devices configured with the candidate large model are not the scheduled service clients.

[0109] In an embodiment of the present disclosure, multiple devices of the interactive system may be deployed in a distributed manner. For example, the same large model may be deployed on multiple distributed devices, and the geographical locations of the different distributed devices may be different.

[0110] Since the same large model can be deployed on multiple distributed devices, in order to improve the processing efficiency of the interaction process, when determining the distributed device that is scheduled to serve the client, the geographical location of the client terminal and the geographical locations of the multiple distributed devices deployed with the candidate large model can be used to select the distributed device that is closer to the client terminal as the distributed device that is scheduled to serve the client, so as to reduce the information transmission distance when interacting with the interactive system.

[0111] If the distributed device of the intended service client does not experience any anomalies, the candidate large model configured on the distributed device of the intended service client can be determined as the target large model. However, if the distributed device of the intended service client experiences an anomaly, the distributed device of the intended service client cannot normally provide interactive services. In this case, it is necessary to reselect a distributed device that can serve the client from multiple distributed devices that have candidate large models deployed, and determine the candidate large model configured on this distributed device as the target large model.

[0112] When reselecting a distributed device that can serve the client, a distributed device with a lower load among other distributed devices may be selected to serve the client based on the load information of other distributed devices, so as to improve processing efficiency.

[0113] According to an embodiment of the present disclosure, when an abnormality occurs in the distributed device of the predetermined service client, the interaction request is diverted from the distributed device of the predetermined service client to other distributed devices deployed with candidate large models to ensure user experience, and load information is also taken into consideration when reselecting the distributed device of the service client, which can ensure that the processing efficiency of the device running the target large model is high, thereby improving the efficiency of the interaction process.

[0114] Figure 6A The figure schematically shows a schematic diagram of determining a target macro model according to an embodiment of the present disclosure.

[0115] like Figure 6A As shown, the candidate large models include the first large model 10, the third large model 30 and the fourth large model 40. The distributed device scheduled to serve the client is the distributed device configured with the first large model 10, and the fault information of the distributed device indicates that the distributed device is abnormal.

[0116] like Figure 6A As shown, the load information of the third large model 30 indicates that the distributed device load of the third large model 30 is low, and the load information of the fourth large model 40 indicates that the distributed device load of the fourth large model 40 is high. When the device routing policy indicates routing based on device status, the target large model is determined to be the third large model 30.

[0117] According to an embodiment of the present disclosure, the candidate large models include multiple ones, and the device configuration information includes resource configuration information. Based on the device configuration information and the device routing policy, a target large model is determined from at least one candidate large model, including: when the device routing policy represents routing based on resource allocation, resource configuration information for running the candidate large model device is obtained from the interactive system to obtain multiple resource configuration information; determining the relationship between the multiple candidate large models to obtain a relationship identification result; and when the relationship identification result represents that the multiple candidate large models are in a homologous relationship, based on the multiple device resource configurations, determining the target large model from the multiple candidate large models; the homologous relationship indicates that the multiple candidate large models have the same service type but different training rounds.

[0118] The resource configuration information may represent the resource configuration of the device running the large model, such as the number of GPUs, etc. When obtaining the resource configuration information, the intermediate server may send the model identifiers of the multiple candidate large models to the interactive system to query the resource configuration information of the multiple candidate large models from the interactive system.

[0119] During the iterative update of a large model in an interactive system, multiple candidate large models may be of the same origin. In this case, if the candidate large models have the same service type but different training rounds, they can be considered to be old and new versions of the large model for the same service type. For example, a candidate large model with fewer training rounds is considered the old version, while a candidate large model with more training rounds is considered the new version.

[0120] In the embodiments of the present disclosure, since a batch of spare hardware resources need to be prepared for starting the new version of the service during the iterative update of the large model, if the service scale is large, the cost of spare hardware resources is high, so the rolling update method can be used to iteratively update the large model of the same service type.

[0121] During the rolling update process, the number of instances of the old version of the large model gradually decreases, and the number of instances of the new version of the large model gradually increases. At this time, the new version of the large model can be started in sequence using the hardware resources idled by the old version of the large model that has been offline to achieve rolling updates.

[0122] When routing based on resource allocation, since the resource configurations of different versions of large models may be different, for example, the new version candidate large model may have a higher resource configuration than the old version candidate large model, the number of requests sent to the new version candidate large model and the old version candidate large model can be reasonably allocated based on the resource configuration information to ensure the stability of the interactive system.

[0123] For example, the ratio of the GPU occupied by the new version large model to the GPU occupied by the old version large model is 3:2. When routing based on resource allocation, the request processing volume of the new version large model and the request processing volume of the old version large model can be maintained at 3:2.

[0124] According to an embodiment of the present disclosure, when multiple candidate large models are of the same origin, the stability of the interactive system can be ensured by determining the target large model based on resource allocation routing.

[0125] The above Figure 6A The process of determining the target large model based on device state routing is detailed below. Figure 6B The process of determining the target large model based on resource configuration routing is described in detail.

[0126] Figure 6B A schematic diagram of determining a target macro model according to another embodiment of the present disclosure is schematically shown.

[0127] like Figure 6B As shown, the candidate large models include a first large model 10, a third large model 30, and a fourth large model 40, and the first large model 10, the third large model 30, and the fourth large model 40 are of the same origin. The resource configuration information of the first large model 10 is high resource configuration, the resource configuration information of the third large model 30 is low resource configuration, and the resource configuration information of the fourth large model 40 is low resource configuration.

[0128] like Figure 6B As shown, when the device routing policy characterizes routing based on resource allocation, the first large model 10 is determined as the target large model.

[0129] According to an embodiment of the present disclosure, the artificial intelligence-based interaction method further includes: determining the amount of information to be processed of the target large model; and updating the target large model when the amount of information is greater than or equal to a preset information amount threshold.

[0130] The target large model's pending information may be information that has already been forwarded to the target large model through routing operations and has not yet been processed by the target large model. If the target large model has a large amount of pending information, such as if the amount of pending information is greater than or equal to a preset information threshold, the target large model may be unable to process the input information in a timely manner, resulting in low efficiency in processing the input information using the target large model. Therefore, the target large model may be updated and redefined to improve interaction efficiency.

[0131] In some embodiments, when updating the target large model, the target large model can be reselected from other candidate large models based on the device routing policy. For example, the target large model can be re-determined based on the load information of the running devices of other candidate large models.

[0132] According to the embodiments of the present disclosure, by judging the amount of information to be processed of the target large model and updating the target large model when the amount of information is greater than or equal to a preset information amount threshold, it is possible to avoid low processing efficiency of input information due to a large amount of information to be processed in the target large model and improve interaction efficiency.

[0133] According to an embodiment of the present disclosure, the target big model is updated, including: updating the historical big model to the target big model based on the model identifier of the historical big model indicated in the interaction request; the historical big model is used to process the above information from the client, and the above information is at least the previous round of input information in the multi-round interaction process.

[0134] Since users can have multiple rounds of interactions with the interactive system, in order to ensure the interactive experience of users in each round of interaction, the target large model can be re-determined when processing the information input by users in each round, so as to use the target large model that is most suitable for processing the input information of the current round for interaction.

[0135] However, if the target large model in the current round has a large amount of information to process, processing efficiency cannot be guaranteed. In this case, a historical large model that has processed the client's previous information can be identified as the target large model. Because the historical large model has processed input information from at least the previous round of multiple interaction rounds, using the historical large model that understands these multiple interaction rounds to process input information can improve processing efficiency.

[0136] According to an embodiment of the present disclosure, by determining the historical big model that processes the above information as the target big model, the processing efficiency of the target big model can be improved, thereby improving the efficiency of the interaction process.

[0137] According to an embodiment of the present disclosure, updating the target large model includes: updating other candidate large models in at least one candidate large model to the target large model.

[0138] Since the candidate large models are suitable for processing input information, when updating other candidate large models to the target large model, the device routing strategy used when determining the target large model can be used to screen the candidate large models again, or the device routing strategy not used when determining the target large model can be used to screen the candidate large models.

[0139] For example, after determining the target large model based on the load information of other distributed devices, if the amount of information to be processed of the target large model is greater than the preset information amount threshold, the target large model can be re-determined based on the load information of other distributed devices except the distributed device that configures the target large model.

[0140] For example, after determining the target large model based on the load information of other distributed devices, if the amount of information to be processed of the target large model is greater than the preset information amount threshold, the target large model can be re-determined based on the resource configuration information of the running devices of other candidate large models.

[0141] According to the embodiments of the present disclosure, by re-determining the target big model from the candidate big models, not only can the function of the determined target big model be guaranteed to be compatible with the intention of the input information, but also the determination efficiency of the target big model can be improved, thereby improving the overall efficiency of the interaction process.

[0142] According to an embodiment of the present disclosure, the artificial intelligence-based interaction method also includes: generating a processing request for calling the target large model based on the input information when the amount of information is less than a preset information amount threshold; and sending the processing request to the interactive system.

[0143] When the amount of information is less than the preset information amount threshold, it means that the target large model can process the input information faster. At this time, the interactive system can be directly instructed to use the target large model to process the input information.

[0144] When instructing the interactive system to process the input information using the target large model, a processing request for calling the target large model may be generated based on the input information, and the processing request may be sent to the interactive system.

[0145] In some embodiments, the model identifier of the target large model may be written into the processing request so that the interactive system can call the target large model according to the model identifier of the target large model after parsing the processing request.

[0146] According to an embodiment of the present disclosure, by generating a processing request for calling a target large model, the interactive system can directly call the target large model to process input information without requiring routing operations of the interactive system, thereby reducing the burden on the interactive system.

[0147] According to an embodiment of the present disclosure, a processing request for calling a target large model is generated based on input information, including: obtaining a request generation program that matches the system identifier; using the request generation program to process the input information and the model identifier indicating the target large model to obtain a processing request that matches the interaction mode of the interactive system; wherein the request generation program is obtained by encapsulating the configuration component.

[0148] Since different interactive systems may be developed by different manufacturers and may use different interactive protocols, which may lead to differences in the field definition, range, format, etc. of the interactive requests that can be processed, therefore, considering that different interactive systems use different protocols, when generating processing requests, it is necessary to generate processing requests that match the interactive mode of the interactive system to ensure that the interactive system can process the interactive request normally.

[0149] In the embodiments of the present disclosure, to improve the efficiency of generating processing requests, multiple configuration components for generating processing requests can be pre-encapsulated to obtain a request generation program. When a processing request needs to be generated, the request generation program can be directly called to process the input information and a model identifier indicating the target macro model to obtain the processing request. The configuration components may include, for example, a format conversion component and an information field concatenation component.

[0150] Since different interactive systems use different protocols, the request generation programs corresponding to different interactive systems are also different. It is necessary to determine the request generation program that matches the interactive system based on the system identifier, so as to use the request generation program to generate a processing request that matches the interactive mode of the interactive system.

[0151] According to an embodiment of the present disclosure, by utilizing a request generation program that matches a system identifier, a processing request that matches an interactive system is generated, thereby improving the generation efficiency and accuracy of the processing request.

[0152] According to an embodiment of the present disclosure, the request generation program is obtained by the following operations: determining the interaction mode of the interactive system; updating the parameters of the reference component based on the interaction mode to determine the configuration component; and encapsulating the configuration component to obtain the request generation program.

[0153] When generating a request generation program, the interaction mode of the interactive system can be determined first, so that the generated request generation program can generate processing requests that match the interaction mode. The interaction mode may include, for example, the interaction protocol used by the interactive system. Different interaction modes may correspond to different processing request formats, such as field definitions, ranges, and formats.

[0154] In an embodiment of the present disclosure, the reference component may be a request generation component that does not adaptively adjust parameters. For example, the reference component may include a format conversion reference component, but the format conversion reference component may not define parameters for format conversion, such as field length.

[0155] The interaction mode may have corresponding component parameters, so the parameters of the reference component may be updated according to the interaction mode to adjust the reference component to a configuration component matching the interaction mode, and then the configuration component may be encapsulated to obtain a request generation program.

[0156] According to an embodiment of the present disclosure, by updating the reference component to a configuration component based on the interaction mode, and encapsulating the configuration component into a request generation program, the obtained request generation program can generate a processing request that matches the interaction mode, thereby improving the efficiency and accuracy of processing request generation.

[0157] According to an embodiment of the present disclosure, based on the system identifier, the model configuration information and device configuration information of the interactive system are obtained, including: based on the system configuration mapping relationship and the system identifier, the system configuration information of the interactive system is determined from multiple candidate system configuration information, the system configuration information includes model configuration information and device configuration information, and the system configuration mapping relationship represents the association relationship between the system identifier and the candidate system configuration information.

[0158] In an embodiment of the present disclosure, a system configuration mapping relationship and multiple candidate system configuration information may be pre-stored in the intermediate server. When obtaining the model configuration information and device configuration information of the interactive system, the candidate system configuration information that has an association relationship with the system identifier may be determined as the system configuration information of the interactive system based on the system configuration mapping relationship and the system identifier.

[0159] According to an embodiment of the present disclosure, by pre-storing candidate system configuration information of multiple interactive systems in an intermediate server, model configuration information and device configuration information can be directly obtained based on the system configuration mapping relationship and system identification, thereby improving the efficiency of obtaining model configuration information and device configuration information.

[0160] According to an embodiment of the present disclosure, the artificial intelligence-based interaction method also includes: obtaining the model identifier of the historical big model indicated in the interaction request, the historical big model is used to process the context information from the client, and the context information is at least the previous round of input information in the multi-round interaction process; and when the intention information of the input information matches the functional information of the historical big model, the historical big model is used as the target big model.

[0161] When a user performs multiple rounds of interactions with the interactive system, the model identifier of the historical large model can be indicated in the interaction request so that the intermediate server can directly route the interaction request to the historical large model and use the historical large model to process the input information of the current round.

[0162] After receiving the interaction request, the intermediate server can judge the input information. If the intention information of the input information matches the functional information of the historical large model, the historical large model can be directly determined as the target large model, and the historical large model can be used to directly process the input information of the current round.

[0163] For example, the intention information of the input information indicates that the user's intention is to play music, and the function information of the historical large model also indicates that the historical large model is used to play music. In this case, the historical large model can be directly used to process the input information.

[0164] However, there may be a mismatch between the intent information of the input information and the functional information of the historical large model. In this case, in order to improve the user experience, the intermediate server can also re-determine the target large model.

[0165] According to an embodiment of the present disclosure, by using the historical big model as the target big model when the intention information of the input information matches the functional information of the historical big model, the processing efficiency of the target big model on the input information can be improved, thereby improving the efficiency of the interaction process.

[0166] Figure 7 The flowchart of determining the target macro model according to an embodiment of the present disclosure is schematically shown.

[0167] like Figure 7 As shown, determining the target macro model includes operations S710 to S760.

[0168] In operation S710 , an interaction request is acquired.

[0169] In operation S720, it is determined whether the interaction request indicates a model identifier of a historical large model. If the interaction request indicates a model identifier of a historical large model, operation S730 is performed; otherwise, operation S750 is performed.

[0170] In operation S730, it is determined whether the intention information of the input information matches the function information of the historical large model. If the intention information of the input information matches the function information of the historical large model, operation S740 is performed, otherwise operation S750 is performed.

[0171] In operation S740 , the historical large model is determined as a target large model.

[0172] In operation S750, at least one candidate large model is determined from a plurality of large models based on the attribute information and the model configuration information of the input information.

[0173] In operation S760, a target large model is determined from at least one candidate large model based on the device configuration information.

[0174] Figure 8 A block diagram of an artificial intelligence-based interaction device according to an embodiment of the present disclosure is schematically shown.

[0175] like Figure 8As shown, the artificial intelligence-based interaction device 800 includes a request acquisition module 810 , a configuration acquisition module 820 , a candidate determination module 830 , and a target determination module 840 .

[0176] The request acquisition module 810 is used to acquire an interaction request from a client, where the interaction request includes input information to be processed and a system identifier indicating a required interaction system; the interaction system is configured with multiple large models.

[0177] The configuration acquisition module 820 is used to obtain the model configuration information and device configuration information of the interactive system based on the system identifier; the model configuration information represents the model configuration of each of the multiple large models, and the device configuration information represents the device configuration used to run the large model device.

[0178] The candidate determination module 830 is configured to determine at least one candidate large model from a plurality of large models based on the attribute information and the model configuration information of the input information.

[0179] The target determination module 840 is used to determine a target large model from at least one candidate large model based on the device configuration information, so that the interactive system calls the target large model to process input information and obtain feedback information for the interactive request.

[0180] According to an embodiment of the present disclosure, the attribute information includes intent information; the model configuration information includes function information; and the candidate determination module 830 includes a function determination submodule and a configuration determination submodule.

[0181] The function determination submodule is used to determine multiple initial candidate large models that can process input information from multiple large models based on intention information and function information.

[0182] The configuration determination submodule is used to determine at least one candidate large model from multiple initial candidate large models based on the model routing strategy and model configuration information.

[0183] According to an embodiment of the present disclosure, the model configuration information further includes time-consuming information; and the configuration determination submodule includes a time-consuming routing unit.

[0184] The time-consuming routing unit is configured to determine at least one candidate large model from a plurality of initial candidate large models based on respective time-consuming information of the plurality of initial candidate large models when the model routing strategy representation is based on time-consuming routing.

[0185] According to an embodiment of the present disclosure, the model configuration information further includes resource consumption information; and the configuration determination submodule includes a resource routing unit.

[0186] The resource routing unit is configured to determine at least one candidate large model from a plurality of initial candidate large models based on resource consumption information of each of the plurality of initial candidate large models when the model routing strategy characterizes resource consumption-based routing.

[0187] According to an embodiment of the present disclosure, the target determination module 840 includes a device routing submodule.

[0188] The device routing submodule is configured to determine a target large model from at least one candidate large model based on device configuration information and a device routing policy.

[0189] According to an embodiment of the present disclosure, the device configuration information includes device status information. The device routing submodule includes a status acquisition unit and a status routing unit.

[0190] The state acquisition unit is used to acquire device state information of a device for running a candidate large model from an interactive system when the device routing policy characterizes routing based on device state, and obtain at least one piece of device state information.

[0191] The state routing unit is configured to determine a target large model from at least one candidate large model based on at least one device state information.

[0192] According to an embodiment of the present disclosure, the device status information includes fault information and load information; the status routing unit includes a predetermined determination subunit and a load routing subunit.

[0193] The reservation determination subunit is used to determine the distributed device that is scheduled to serve the client from the distributed device cluster of the interactive system, and the distributed device is configured with a candidate large model.

[0194] The load routing subunit is used to determine the target large model based on the load information of other distributed devices in the distributed cluster when the fault information of the distributed device indicates that the distributed device is abnormal; and configure the non-scheduled service clients of other distributed devices with candidate large models.

[0195] According to an embodiment of the present disclosure, the candidate large models include multiple ones; the device configuration information includes resource configuration information; and the device routing submodule includes a configuration acquisition unit, a relationship identification unit, and a device routing unit.

[0196] The configuration acquisition unit is used to acquire resource configuration information for running the candidate large model device from the interactive system when the device routing policy representation is based on resource allocation routing, and obtain multiple resource configuration information.

[0197] The relationship identification unit is used to determine the relationship between multiple candidate large models and obtain a relationship identification result.

[0198] The device routing unit is used to determine the target large model from multiple candidate large models based on multiple device resource configurations when the relationship recognition result indicates that multiple candidate large models are of the same source relationship; the same source relationship indicates that the service types of the multiple candidate large models are the same but the training rounds are different.

[0199] According to an embodiment of the present disclosure, the artificial intelligence-based interaction device 800 further includes an information determination module and a target update module.

[0200] The information determination module is used to determine the amount of information to be processed of the target large model.

[0201] The target updating module is used to update the target large model when the amount of information is greater than or equal to a preset information threshold.

[0202] According to an embodiment of the present disclosure, the target update module includes a history update submodule.

[0203] The history update submodule is used to update the historical large model to the target large model based on the model identifier of the historical large model indicated in the interaction request; the historical large model is used to process the above information from the client, and the above information is at least the previous round of input information in the multi-round interaction process.

[0204] According to an embodiment of the present disclosure, the target update module includes a candidate update sub-module.

[0205] The candidate updating submodule is used to update other candidate large models in at least one candidate large model to a target large model.

[0206] According to an embodiment of the present disclosure, the artificial intelligence-based interaction device 800 further includes a request generating module and a request sending module.

[0207] The request generation module is used to generate a processing request for calling the target large model based on the input information when the amount of information is less than a preset information amount threshold.

[0208] The request sending module is used to send the processing request to the interactive system.

[0209] According to an embodiment of the present disclosure, the request generation module includes a request generation submodule.

[0210] The request generation submodule is used to process the input information and the model identifier indicating the target large model using the request generation program to obtain a processing request that matches the interaction mode of the interactive system; wherein the request generation program is obtained by encapsulating the configuration component.

[0211] According to an embodiment of the present disclosure, the request generation module further includes a mode determination module, a component update module and a component packaging module.

[0212] The mode determination module is used to determine the interaction mode of the interactive system.

[0213] The component update module is used to update the parameters of the reference component based on the interactive mode and determine the configuration component.

[0214] The component encapsulation module is used to encapsulate the configuration component to obtain the request generation program.

[0215] According to an embodiment of the present disclosure, the configuration acquisition module 820 includes a configuration acquisition submodule.

[0216] The configuration acquisition submodule is used to determine the system configuration information of the interactive system from multiple candidate system configuration information based on the system configuration mapping relationship and the system identifier. The system configuration information includes model configuration information and device configuration information. The system configuration mapping relationship represents the association between the system identifier and the candidate system configuration information.

[0217] According to an embodiment of the present disclosure, the artificial intelligence-based interaction device 800 further includes a history acquisition module and a history determination module.

[0218] The history acquisition module is used to obtain the model identifier of the historical large model indicated in the interaction request. The historical large model is used to process the previous information from the client. The previous information is at least the previous round of input information in the multi-round interaction process.

[0219] The history determination module is used to use the historical big model as the target big model when the intention information of the input information matches the function information of the historical big model.

[0220] According to an embodiment of the present disclosure, the artificial intelligence-based interaction device 800 further includes a feedback sending module.

[0221] The feedback sending module is used to receive feedback information for the interaction request from the interaction system and send it to the client.

[0222] Figure 9 The structural block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is schematically shown.

[0223] In the embodiments of the present disclosure, Figure 9 As shown, the AI ​​agent 900 may include an input module 910 , a processing module 920 and an output module 930 .

[0224] Input module 910, for receiving input information;

[0225] A processing module 920 is configured to determine a target task based on input information received by the input module, determine a large language model based on the target task, and obtain output information by invoking the large language model to execute an interaction method based on a large language model according to an embodiment of the present disclosure, or by invoking the large language model to execute a training method for a large language model according to an embodiment of the present disclosure.

[0226] The output module 930 is used to output the output information obtained by the processing module.

[0227] According to an embodiment of the present disclosure, the input module 910 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., a user or the external environment) and converting it into a format that can be understood and processed by the AI ​​agent 900. The input module 910 is the primary link for the AI ​​agent 900 to interact with the outside world. It enables the AI ​​agent 900 to efficiently and accurately obtain the necessary "sensory" information from the outside world and respond to this information.

[0228] In an example, the input module 910 may input the demand speech or sample demand speech, sample demand speech features, demand speech features, etc. described above.

[0229] In this example, the processing module 920 is the core support for the AI ​​agent 900 to handle complex tasks. The processing module 920 can execute the interaction method based on the large language model and the training method of the large language model described above.

[0230] In this example, the performance of processing module 920 may be closely related to the large model underlying AI agent 900. To fully leverage the capabilities of the large model, the internal structure of processing module 920 may be designed to be highly configurable and extensible to handle a variety of different tasks and requirements in real-world scenarios.

[0231] In the example, after the AI ​​agent 900 obtains the required voice, the processing module 920 can use the large voice recognition model to process the required voice to obtain voice recognition features, and the large language model to process the voice recognition features to obtain the reply text, and pass the reply text to the output module 930.

[0232] Understandably, while large language models possess excellent language understanding and generation capabilities, like humans, they are limited in the tasks they can perform without tools. However, when AI Agent 900 is empowered with tool-based capabilities, it can perform tasks such as mathematical calculations using a calculator, data analysis using Python, and weather forecasting using search engines.

[0233] In an example, the output module 930 may output the reply text described above or the trained large language model.

[0234] The AI ​​agent 900 according to the embodiment of the present disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.

[0235] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0236] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0237] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above method.

[0238] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the above method when executed by a processor.

[0239] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0240] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. RAM 1003 may also store various programs and data required for the operation of device 1000. Computing unit 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0241] Various components in device 1000 are connected to an input / output (I / O) interface 1005, including an input unit 1006, such as a keyboard and mouse; an output unit 1007, such as various types of displays and speakers; a storage unit 1008, such as a magnetic disk and optical disk; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0242] The computing unit 1001 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the artificial intelligence-based interaction method. For example, in some embodiments, the artificial intelligence-based interaction method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the artificial intelligence-based interaction method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute an artificial intelligence-based interaction method in any other appropriate manner (eg, by means of firmware).

[0243] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0244] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0245] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0246] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0247] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0248] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0249] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0250] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. An artificial intelligence-based interaction method, comprising: Obtaining an interaction request from a client, the interaction request including input information to be processed and a system identifier indicating a required interaction system; The interactive system is configured with a plurality of large models; Based on the system identifier, model configuration information and device configuration information of the interactive system are obtained; the model configuration information represents the model configuration of each of the plurality of large models, and the device configuration information represents the device configuration for running the large model device; determining at least one candidate large model from a plurality of large models based on the attribute information of the input information and the model configuration information; as well as Based on the device configuration information, the target large model is determined from at least one of the candidate large models, so that the interactive system calls the target large model to process the input information and obtain feedback information for the interactive request.

2. The method according to claim 1, wherein The attribute information includes intent information; the model configuration information includes function information; The determining, based on the attribute information of the input information and the model configuration information, at least one candidate large model from the plurality of large models comprises: Determining, based on the intention information and the function information, a plurality of initial candidate large models that can process the input information from the plurality of large models; and At least one candidate large model is determined from a plurality of initial candidate large models based on a model routing strategy and the model configuration information.

3. The method according to claim 2, wherein: The model configuration information also includes time-consuming information; The determining, based on the model routing strategy and the model configuration information, at least one candidate large model from the plurality of initial candidate large models comprises: In the case where the model routing strategy representation is based on time-consuming routing, at least one candidate large model is determined from the multiple initial candidate large models based on the time-consuming information of each of the multiple initial candidate large models.

4. The method according to claim 2, wherein: The model configuration information also includes resource consumption information; The determining, based on the model routing strategy and the model configuration information, at least one candidate large model from the plurality of initial candidate large models comprises: In the case where the model routing strategy represents resource consumption-based routing, at least one candidate large model is determined from the plurality of initial candidate large models based on the resource consumption information of each of the plurality of initial candidate large models.

5. The method according to any one of claims 1 to 4, wherein The step of determining the target large model from at least one candidate large model based on the device configuration information includes: The target big model is determined from at least one of the candidate big models based on the device configuration information and the device routing policy.

6. The method according to claim 5, wherein: The device configuration information includes device status information; The determining the target big model from at least one of the candidate big models based on the device configuration information and the device routing policy includes: In the case where the device routing strategy representation is based on device state routing, obtaining device state information for running the candidate large model device from the interactive system to obtain at least one piece of device state information; as well as The target large model is determined from at least one of the candidate large models based on at least one of the device status information.

7. The method according to claim 6, wherein: The device status information includes fault information and load information; The determining the target large model from at least one candidate large model based on at least one piece of device status information includes: Determining a distributed device that is scheduled to serve the client from a distributed device cluster of the interactive system, wherein the distributed device is configured with the candidate large model; as well as When the fault information of the distributed device indicates that the distributed device is abnormal, the target large model is determined based on the load information of other distributed devices in the distributed cluster; the other distributed devices configured with the candidate large model do not serve the client in a scheduled manner.

8. The method according to claim 5, wherein The candidate large models include multiple ones; the device configuration information includes resource configuration information; The determining the target big model from at least one of the candidate big models based on the device configuration information and the device routing policy includes: In the case where the device routing policy representation is based on resource allocation routing, resource configuration information for running the candidate large model device is acquired from the interactive system to obtain a plurality of resource configuration information; Determining the relationships between the plurality of candidate large models to obtain a relationship recognition result; and When the relationship identification result indicates that the plurality of candidate large models are of a homologous relationship, the target large model is determined from the plurality of candidate large models based on the plurality of device resource configurations; the homologous relationship indicates that the plurality of candidate large models have the same service type but different training rounds.

9. The method according to any one of claims 1 to 8, further comprising: Determining the amount of information to be processed of the target large model; When the amount of information is greater than or equal to a preset information threshold, the target large model is updated.

10. The method according to claim 9, wherein: The updating of the target large model includes: Based on the model identifier of the historical large model indicated in the interaction request, the historical large model is updated to the target large model; the historical large model is used to process the above information from the client, and the above information is at least the previous round of information of the input information in the multi-round interaction process.

11. The method according to claim 9, wherein: The updating of the target large model includes: At least one other candidate large model in the candidate large model is updated to the target large model.

12. The method according to claim 9, further comprising: When the amount of information is less than the preset information amount threshold, generating a processing request for calling the target large model based on the input information; as well as The processing request is sent to the interactive system.

13. The method according to claim 12, wherein: The generating, based on the input information, a processing request for calling the target large model comprises: Obtaining a request generation program that matches the system identifier; Processing the input information and the model identifier indicating the target large model using a request generation program to obtain the processing request that matches the interaction mode of the interactive system; The request generation program is obtained by encapsulating the configuration component.

14. The method according to claim 13, further comprising: determining an interaction mode of the interactive system; Based on the interaction mode, updating parameters of the reference component to determine the configuration component; as well as The configuration component is encapsulated to obtain the request generation program.

15. The method according to any one of claims 1 to 14, wherein: The acquiring, based on the system identifier, model configuration information and device configuration information of the interactive system includes: Based on the system configuration mapping relationship and the system identifier, the system configuration information of the interactive system is determined from multiple candidate system configuration information, the system configuration information includes the model configuration information and the device configuration information, and the system configuration mapping relationship represents the association relationship between the system identifier and the candidate system configuration information.

16. The method according to any one of claims 1 to 15, further comprising: Obtaining a model identifier of a historical large model indicated in the interaction request, wherein the historical large model is used to process previous information from the client, wherein the previous information is information of at least a previous round of the input information in a multi-round interaction process; as well as In a case where the intention information of the input information matches the function information of the historical big model, the historical big model is used as the target big model.

17. The method according to any one of claims 1 to 16, further comprising: Feedback information for the interaction request is received from the interaction system and sent to the client.

18. An interactive device based on artificial intelligence, comprising: A request acquisition module, configured to acquire an interaction request from a client, wherein the interaction request includes input information to be processed and a system identifier indicating a desired interaction system; The interactive system is configured with a plurality of large models; A configuration acquisition module, configured to acquire model configuration information and device configuration information of the interactive system based on the system identifier; the model configuration information represents the model configuration of each of the plurality of large models, and the device configuration information represents the device configuration for running the large model device; a candidate determining module, configured to determine at least one candidate large model from the plurality of large models based on the attribute information of the input information and the model configuration information; The target determination module is used to determine the target large model from at least one of the candidate large models based on the device configuration information, so that the interactive system calls the target large model to process the input information and obtain feedback information for the interactive request.

19. An artificial intelligence agent, comprising: An input module, used for receiving input information; a processing module, configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the method according to any one of claims 1 to 17 by calling the large model to obtain output information; An output module is used to output the output information obtained by the processing module.

20. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 17.

21. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 17.

22. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 17.

Citation Information

Cited By

  • Data processing method, platform and system, computing equipment and storage medium

    CN122285234A