Agent application object generation method and apparatus, and computer device and storage medium
By using a public service pool during the generation of intelligent agent application objects and referencing public services to compile components in the topological representation graph, the problem of low efficiency in the generation of traditional intelligent agent application objects is solved, and efficient and stable resource utilization is achieved.
Patent Information
- Application Number
- PCT/CN2025/101496
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-08
- Filing Date
- 2025-06-17
- Publication Date
- 2026-02-12
AI Technical Summary
Traditional intelligent agent applications have low object generation efficiency, especially when constructing topological representation graphs, which is time-consuming and resource-intensive, and also has service stability issues.
By using a public service pool, components in the topology representation graph are compiled by referencing public services. Public services are stripped out to achieve resource sharing and optimize the compilation process.
It significantly improves the efficiency of generating intelligent agent application objects, saves computing, storage and network resources, and enhances resource utilization efficiency and service stability.
Smart Images

Figure CN2025101496_12022026_PF_FP_ABST
Abstract
Description
Method and device for generating application object of agent, computer device and storage medium
[0001] Related applications
[0002] The present application claims priority to the Chinese patent application No. 202411081852.8, filed on August 8, 2024, and entitled "Method and device for generating application object of agent, computer device and storage medium", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] The present application relates to the technical field of computer, in particular to a method and device for generating application object of agent, computer device, storage medium and computer program product. BACKGROUND
[0004] With the development of computer technology, Agent applications have emerged. Agent applications, i.e., intelligent agent applications, can be implemented based on large language models (LLM). A large language model is an artificial intelligence model designed to understand and generate human language. Agent applications based on large language models determine their behavior according to the current state of the agent and external input. They can be used to solve a specific type of problem, such as a legal consultation agent that provides legal consultation services.
[0005] In traditional technology, the generation of intelligent agent application objects can be obtained by compiling a pre-constructed topological representation graph. However, there is often a problem of low efficiency in the generation of intelligent agent application objects. SUMMARY
[0006] The present application provides a method and device for generating application object of agent, computer device, computer readable storage medium and computer program product.
[0007] In a first aspect, the present application provides a method for generating application object of agent, which is executed by a computer device, and the method comprises:
[0008] In response to a service request, a topological representation graph indicated by the service request is determined, the topological representation graph comprising a plurality of topological nodes, each topological node representing a component, and at least one of the components using a common service in a common service pool;
[0009] The topological representation graph is compiled to initialize the components represented by the compiled topological nodes, and the initialized components are obtained.
[0010] wherein, when a compiled topology node represents a common service using component using a common service in the common service pool, a function of the common service using component using the common service is compiled in a manner of referencing the common service in the common service pool used by the common service using component; and
[0011] In response to completion of the compiling of the topology representation graph, an agent application object is obtained based on the initialized components corresponding to each topology node in the topology representation graph.
[0012] In a second aspect, the present application further provides an agent application object generation device. The device comprises:
[0013] a topology representation graph determination module configured to determine a topology representation graph indicated by a service request in response to the service request, the topology representation graph comprising a plurality of topology nodes, each topology node representing a component, and at least one of the components using a common service in a common service pool;
[0014] a topology representation graph compiling module configured to compile the topology representation graph to initialize the components represented by the compiled topology nodes, and obtain initialized components; wherein, when a compiled topology node represents a common service using component using a common service in the common service pool, a function of the common service using component using the common service is compiled in a manner of referencing the common service in the common service pool used by the common service using component; and
[0015] an agent application object obtaining module configured to obtain an agent application object based on the initialized components corresponding to each topology node in the topology representation graph in response to completion of the compiling of the topology representation graph.
[0016] In a third aspect, the present application further provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the above-mentioned agent application object generation method when executing the computer program.
[0017] In a fourth aspect, the present application further provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program implements the steps of the above-mentioned agent application object generation method when executed by a processor.
[0018] In a fifth aspect, the present application further provides a computer program product. The computer program product comprises a computer program, and the computer program implements the steps of the above-mentioned agent application object generation method when executed by a processor.
[0019] The details of one or more embodiments of the application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the application will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort based on the disclosed drawings.
[0021] Fig. 1 is an application environment diagram of the agent application object generation method in some embodiments;
[0022] Fig. 2 is a flow diagram of the agent application object generation method in some embodiments;
[0023] Fig. 3 is a schematic diagram of a topology characterization graph in some embodiments;
[0024] Fig. 4 is a schematic diagram of a topology characterization graph in some other embodiments;
[0025] Fig. 5 is a schematic diagram of the overall process of the agent application object generation method in some embodiments;
[0026] Fig. 6 is a flow diagram of the response steps for a newly initiated service request in some embodiments;
[0027] Fig. 7 is a schematic diagram of an agent application service framework in some embodiments;
[0028] Fig. 8 is a schematic diagram of an agent application service framework in some other embodiments;
[0029] Fig. 9 is a schematic diagram of the compilation process of an agent application object in some embodiments;
[0030] Fig. 10 is a schematic diagram of the service process of an agent application object in some embodiments;
[0031] Fig. 11 is a structural block diagram of an agent application object generation apparatus in some embodiments;
[0032] Fig. 12 is an internal structure diagram of a computer device in some embodiments;
[0033] Fig. 13 is an internal structure diagram of a computer device in some other embodiments. DETAILED DESCRIPTION
[0034] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0035] The intelligent agent application object generation method provided by the embodiments of the present application can be applied in the application environment as shown in FIG. 1. The terminal 102 and the server 104 can communicate through a network, such as a wired or wireless network. The data storage system can store data required to be processed by the server 104. The data storage system can be separately arranged, can be integrated on the server 104, or can be placed on a cloud or other server. The terminal 102 can be, but is not limited to, various desktop computers, notebook computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be a stand-alone physical server, a server cluster composed of multiple physical servers or a distributed system, or a cloud server providing cloud computing services.
[0036] The execution subject of each step of the intelligent agent application object generation method provided by the embodiments of the present application can be a computer device, which refers to an electronic device with data calculation, processing and storage capabilities. Taking the application environment shown in FIG. 1 as an example, the intelligent agent application object generation method can be executed by the terminal 102 alone, can be executed by the server 104 alone, or can be executed by the terminal 102 and the server 104 in cooperation, and the present application does not make any limitation in this regard.
[0037] Taking the execution by the server 104 alone as an example, the server can respond to a service request sent by the terminal, determine a topology representation graph indicated by the service request, the topology representation graph including a plurality of topology nodes, each topology node representing a component, at least one component using a common service in a common service pool, compile the topology representation graph to initialize the components represented by the compiled topology nodes, and obtain the initialized components, wherein when the compiled topology node represents a common service using component using a common service in a common service pool, the function of the common service using component using the common service in the common service pool is compiled in a manner of referencing the common service used by the common service using component, and the intelligent agent application object is obtained based on the initialized components corresponding to the topology nodes in the topology representation graph.
[0038] In some embodiments, as shown in FIG. 2, an agent application object generation method is provided, which is executed by a computer device, which can be the server 104 or the terminal 102 in FIG. 1. In the embodiments of the present application, the method is applied to the server 104 in FIG. 1 as an example, which includes the following steps:
[0039] In step 202, in response to the service request, a topology characterization graph indicated by the service request is determined, the topology characterization graph includes a plurality of topology nodes, each topology node characterizes a component, and at least one component uses a common service in a common service pool.
[0040] The service request is a request for an agent application to provide a service. In different application scenarios, the service requested by the service request can be different services. For example, in a legal consultation application scenario, the service request is used to request the corresponding agent application to provide legal consultation services, and in a medical consultation application scenario, the service request is used to request the corresponding agent application to provide medical consultation services. An agent application is an application implemented based on an Agent (agent). In an agent application, the behavior of an Agent can be determined based on the current state of the Agent and external input. In specific applications, the agent application of the present application can be an AI-based agent application, for example, an LLM-based agent application.
[0041] The topology characterization graph is used to characterize an agent application. The topology characterization graph includes a plurality of topology nodes, each topology node characterizing a component, i.e., the topology characterization graph is a graph that represents the components included in the agent application with topology nodes, and the connection relationship between the topology nodes in the topology characterization graph represents the dependency relationship between the topology nodes. The topology characterization graph can be a DAG (directed acyclic graph). For example, taking an LLM-based agent application as an example, various tools, LLMs, and other components that the Agent application depends on are represented as topology nodes, and are connected according to the dependency relationship, to obtain a topology characterization graph of the LLM-based agent application. Referring to FIG. 3, which is a schematic diagram of the topology characterization graph in some embodiments. In FIG. 3, the final node of the topology characterization graph is an Agent, which depends on service nodes and various tool nodes provided by an LLM in an upstream. The LLM and various tools are components represented by each topology node in the topology characterization graph. As can be seen from FIG. 3, each tool will also depend on services such as databases and search engines, and will be compiled and built into an Agent application object to provide external services.
[0042] In the components represented by the respective topology nodes of the topology representation graph, at least one component uses a public service in a public service pool, that is, the running of the at least one component depends on one or more public services in the public service pool. The public service pool here refers to a service pool composed of multiple public services, which are services needed by the components in the agent application object in the running process. In this application, these services are separated from the agent application object to form independent services. The advantages mainly include: first, the compilation and construction time of the topology representation graph is greatly saved. In general, database connection, network resource initialization and the like are time-consuming, and separate calling exists service stability guarantee problem; second, a large amount of computing, storage and network resources are saved. Different agents only refer to the objects of the resource pool in the compilation and construction process, and do not independently occupy, use and release, thereby realizing public service resource sharing and greatly improving resource utilization efficiency. For example, referring to FIG. 4, which is a schematic diagram of a topology representation graph in another embodiment. Taking the topology representation graph in FIG. 3 as an example, the databases, search engines, large model services and the like relied on by the respective components can be separated to form a public service pool.
[0043] Specifically, the topology representation graph indicated by the service request is the topology representation graph determined according to the service request. For different application scenarios, the agent application providing the service is usually different, so different topology representation graphs can be constructed in advance for different application scenarios, and the correspondence between the application scenario identifier and the topology representation graph is stored in the database. After receiving the service request, the server first parses the service request, and extracts the application scenario identifier through a specific parsing rule (for example, from the header information of the request or a specific field of the request body). Then, the server uses the application scenario identifier to search in the database. The specific searching method is: performing an exact match query in the application scenario identifier field of the database. If a matching application scenario identifier is found, the topology representation graph associated with the application scenario identifier is obtained, which is the topology representation graph indicated by the service request. If no matching identifier is found, an error message is returned or a topology representation graph is selected according to a preset default rule. The preset default rule is: selecting a general topology representation graph preset in the database, which is suitable for most common application scenarios; if no general topology representation graph is set, selecting the topology representation graph with the highest usage frequency in the database. The application scenario identifier is an identifier used to distinguish different application scenarios. The server extracts the identifier from the request through a specific parsing rule after receiving the service request, and then uses the identifier to search for the corresponding topology representation graph in the database.
[0044] In some embodiments, among the components represented by the respective topology nodes, at least one component is a pre-trained large language model.
[0045] In step 204, the topology characterization graph is compiled to initialize the components represented by the compiled topology nodes, and initialized components are obtained; when the compiled topology nodes represent public service usage components using public services in the public service pool, the functions of the public service usage components using the public services are compiled by referencing the public services used by the public service usage components in the public service pool.
[0046] The public service usage component refers to a component represented by a topology node in the topology characterization graph, which runs in dependence on one or more public services in the public service pool. When the topology characterization graph is compiled, if such a component is compiled, the functions of the public service usage component are compiled by referencing the corresponding public services in the public service pool.
[0047] The topology characterization graph is compiled, that is, each topology node included in the topology characterization graph is compiled. By compiling each topology node, the components represented by the topology nodes can be initialized. After each topology node is compiled, the initialized component corresponding to the topology node can be obtained.
[0048] Specifically, the service request can also indicate a compilation parameter. The compilation parameter can be a parameter required in the initialization process of each component. For example, for an Agent component, the compilation parameter can be an Agent parameter, for an LLM component, the compilation parameter can be an LLM parameter, and for a tool component, the compilation parameter can be a tool parameter. After receiving the service request, the server parses the service request according to the format of the service request. If the service request adopts the JSON format, the compilation parameter is stored in the request body in the form of a key-value pair, the key is the component type (such as “Agent”, “LLM”, “Tool”), and the value is the corresponding parameter. The server extracts the compilation parameters of each component by parsing the JSON data. If the service request adopts the XML format, the server can use an XML parser to extract the compilation parameters of each component according to the pre-defined XML tag structure. If the service request is in other formats, the server can find specific fields or areas to obtain the compilation parameters according to the characteristics of the format. If it is not possible to directly obtain the compilation parameters, the default parameters set for the topology characterization graph are used as the compilation parameters.
[0049] For example, if the JSON data is {“Agent”:{“param1”:“value1”,“param2”:“value2”},“LLM”:{“param3”:“value3”}}, the parameters of the Agent component can be extracted as {“param1”:“value1”,“param2”:“value2”}, and the parameters of the LLM component can be extracted as {“param3”:“value3”}. Then, after the server determines the topology characterization graph according to the service request, the determined topology characterization graph can be compiled according to the obtained compilation parameters, and for each component represented by a topology node in the topology characterization graph, the corresponding compilation parameters are used for initialization.
[0050] In the components represented by the topology nodes, at least one component uses a common service of the common service pool, and therefore, in the process of compiling the topology characterization graph, when the common service using component represented by the compiled topology node does not use the common service, the initialization of the common service using component is completed according to the self-defined manner; when the common service using component represented by the compiled topology node uses the target common service in the common service pool, the server can compile the function of the common service using component using the common service in the manner of referencing the common service used by the common service using component in the common service pool, and for other functions in the common service using component, the initialization of the common service using component is continued according to the self-defined manner. For example, the reference to the target common service in the common service pool can be recording the access location of the target common service in the common service pool, for example, the uniform resource locator (URL) of the target common service.
[0051] Optionally, in the process of compiling the topology characterization graph, the server can also determine the dependency relationship between the topology nodes according to the connection relationship between the topology nodes in the topology characterization graph, and then determine the compilation order of the topology nodes according to the dependency relationship, and compile the topology nodes in the order to initialize the components represented by the compiled topology nodes and obtain the initialized components. Optionally, in the process of compiling the topology characterization graph, for the topology nodes that have no dependency relationship with each other, the compilation can also be performed simultaneously by calling multiple threads to further improve the efficiency of generating the agent application object.
[0052] In step 206, in response to the completion of the compilation of the topology characterization graph, the agent application object determined based on the initialized components corresponding to the topology nodes in the topology characterization graph is obtained.
[0053] Specifically, the server can determine whether each topology node in the topology characterization graph is compiled, and when all topology nodes are compiled, the server can determine that the compilation of the topology characterization graph is completed. The server can obtain an agent application object in response to the completion of the compilation of the topology characterization graph, and the agent application object is an object determined based on the initialized components corresponding to each topology node in the topology characterization graph. The specific determination method is that the server maintains a compiled node list, and adds a compiled node to the list when each topology node is compiled. When the number of nodes in the compiled node list is equal to the total number of topology nodes in the topology characterization graph, it is determined that all topology nodes are compiled.
[0054] Further, the server can obtain a response result of the service request by running the agent application object, and then return the response result to the terminal sending the service request.
[0055] The above steps 202 to 206 will be illustrated below in conjunction with FIG. 5. Referring to FIG. 5, it is a schematic diagram of the overall process of the agent application object generation method in some embodiments. After receiving the service request, the server can determine the topology characterization graph indicated by the service request. It is assumed that the topology characterization graph contains three topology nodes, namely topology node A, topology node B and topology node C, and the component characterized by the topology node C uses a public service. During the compilation of the topology characterization graph, when the topology node A is compiled, the component characterized by the topology node A can be initialized, and the initialized component corresponding to the topology node A is obtained. When the topology node B is compiled, the component characterized by the topology node B can be initialized, and the initialized component corresponding to the topology node B is obtained. When the topology node C is compiled, the part of the component characterized by the topology node C using the public service is compiled by referring to the public service, and the other part is initialized. After obtaining the initialized component corresponding to the topology node C, the compilation of the topology characterization graph is completed, and an agent application object is obtained. The agent application object contains the initialized component corresponding to the topology node A, the initialized component corresponding to the topology node B, and the initialized component corresponding to the topology node C.
[0056] In the intelligent agent application object generation method, in response to the service request, a topology representation graph indicated by the service request is determined, the topology representation graph includes a plurality of topology nodes, each topology node represents a component, at least one component uses a common service in a common service pool, the topology representation graph is compiled to initialize the components represented by the compiled topology nodes, and the initialized components are obtained. In response to the completion of the compilation of the topology representation graph, an intelligent agent application object is obtained based on the initialized components corresponding to the topology nodes in the topology representation graph. When the compiled topology nodes represent a common service usage component using a common service in the common service pool, the function of the common service usage component using the common service is compiled in a manner of referencing the common service used by the common service usage component in the common service pool. In this way, the common service can be separated from the intelligent agent application object, so that only the part of the intelligent agent application object itself needs to be compiled and constructed during the generation of the intelligent agent application object, the occupied computing and storage resources are small, the intelligent agent application object is quickly compiled and constructed, and the generation efficiency of the intelligent agent application object is improved.
[0057] In some embodiments, the service request is a request initiated by a first user under a first session, and the intelligent agent application object generation method further includes: running the intelligent agent application object in a stateless mode to obtain output information of the intelligent agent application object for the service request; taking the output information as a response result of the service request, and returning the response result to a terminal of the first user.
[0058] The session generally represents all contents of a plurality of rounds of interactions of a same user with respect to a same task and an agent application object, and can include dialogues, agent application object states and other information. The stateless mode is a working mode in which state data is not saved. The intelligent agent application object is run in the stateless mode, that is, the state data is not saved in the intelligent agent application object, and the intelligent agent application object is a stateless object. For example, when the intelligent agent application object is run in the stateless mode, the server can save the session state information in the session process in a local database, so that the server can query the session state information from the local database when a next service request under the first session is initiated. Alternatively, the server can carry the session state information in the session process in a response result of the service request and return the response result to the user terminal. The terminal can send the session state information in a service request to the server when a next service request under the first session is initiated.
[0059] Specifically, after obtaining the agent application object, the server can further determine the input information of the agent application object according to the service request, and then run the agent application object in the stateless mode to obtain the output information of the agent application object corresponding to the service request in the first session. The server can take the output information as the response result of the service request and return the response result to the terminal of the first user who initiates the service request.
[0060] It can be understood that after the first user obtains the response result, the first user can continue to initiate a service request in the first session, and the server can cache the generated agent application object. When receiving a service request in the first session again, the server can run the agent application object again to obtain a response result.
[0061] In the embodiment, since the agent application object can be run in the stateless mode, the session state information in the first session can be stripped from the agent application object, so that the agent application object can be reused between different users and different sessions.
[0062] In some embodiments, running the agent application object in the stateless mode to obtain the output information of the agent application object for the service request includes: determining the input information of the agent application object; inputting the input information into the agent application object to run the agent application object and obtain the output information of the agent application object for the service request; and determining the session state information in the first session according to the input information and the state information in the running process of the agent application object.
[0063] Specifically, the agent application object is usually generated when a new session is initiated for the first time. Since it is the first service request, the agent application object does not have historical session state information in the session. Therefore, the user input information is the input information of the agent application object. The user input information is carried in the service request. After receiving the service request, the server can extract the user input information from the service request. When running the agent application object, the server can input the user input information into the agent application object to obtain the output information of the agent application object for the service request.
[0064] The session state information under the first session includes input information of each service request, state information of the agent application object during each running process, and the like, so that after the server obtains the output information for the current service request by running the agent application object, the server can determine the session state information corresponding to the current service request according to the input information, the state information of the agent application object during the running process, and the like. The server can save the session state information in association with the session identifier of the first session to the distributed cache. Since the session state information can be saved to the distributed cache, the agent application object is decoupled from the session state information, the stateless running of the agent application object is realized, and the reuse of the agent application object between different users and different sessions can be maximized.
[0065] The distributed cache is a system for storing data, and is used to store session state information and the like in the present application, so that the agent application object is decoupled from the session state information, the stateless running of the agent application object is realized, and the reusability of the agent application object between different users and different sessions is improved.
[0066] Specifically, the session state information is defined as a JSON object, which includes an "input_info" field for storing input information, an "agent_state" field for storing state information of the agent application object during the running process, and an "output_info" field for storing output information. The server can save the session state information in association with the session identifier of the first session to the distributed cache. Since the session state information can be saved to the distributed cache, the agent application object is decoupled from the session state information, the stateless running of the agent application object is realized, and the reuse of the agent application object between different users and different sessions can be maximized.
[0067] In some embodiments, the server can also determine the session state information under the first session according to the input information, the output information of the agent application object for the service request, and the state information of the agent application object during the running process. For example, the distributed cache can be Redis, and can also be other forms, which are not limited in the present application.
[0068] In the above embodiments, the session state information under the first session can be determined according to the input information and the output information, and the session state information can be saved in association with the session identifier of the first session to the distributed cache. Not only the stateless running of the agent application object is realized, but also the sharing and management of the state data are better supported.
[0069] In some embodiments, input information is input into the agent application object to run the agent application object, and output information of the agent application object for a service request is obtained, including: inputting the input information into the agent application object; according to the topological characterization map, calling each initialized component in the agent application object to process the input information, and obtaining the output information of the agent application object for the service request.
[0070] When the called component is a component using a public service in the public service pool, a service call request is sent to the public service in the public service pool referenced by the called component, and processing of the input information is continued based on a response result of the service call request.
[0071] If an exception occurs when a component processes input information, the server will capture the exception information, record the component name, input information, exception type, and specific error information, etc. Then, the server can perform different processing according to the exception type. If it is a recoverable exception, such as a temporary network failure or a data format error, the server can try to call the component again or process the input information again after modification. If it is an unrecoverable exception, such as a serious error in the component or resource exhaustion, the server will stop the current processing flow, return error information to the user, and record the exception situation for subsequent analysis and repair. A recoverable exception refers to an exception such as a temporary network failure or a data format error that occurs in the process of the agent application object processing input information. The server can try to call the component again or process the input information again after modification. An unrecoverable exception refers to an exception such as a serious error in the component or resource exhaustion that occurs in the process of the agent application object processing input information. The server will stop the current processing flow, return error information to the user, and record the exception situation for subsequent analysis and repair.
[0072] The service call request is a request sent by the server to the public service according to the access location of the public service used by the component recorded in the compilation process when the called component is a component using a public service in the public service pool during the running of the agent application object. The purpose is to obtain the response result of the public service to continue processing the input information.
[0073] Specifically, in the generation process of the agent application object, when compiling the topology characterization graph of the agent application, if the public service used by the public service usage component is in the public service pool, the function of the public service usage component using the public service is compiled in a way of referencing the public service used by the public service usage component in the public service pool. Thus, in the obtained agent application object, the function of the initialized component using the public service corresponding to the topology node is implemented through the public service. Based on this, in the running of the agent application object in the embodiment, the server first inputs the input information into the agent application object. Then, according to the compiling order of each topology node in the topology characterization graph, each initialized component in the agent application object is called to process the input information in turn. In the calling process, if the called component is a component using the public service in the public service pool, the server sends a service calling request to the public service according to the access location (such as a uniform resource locator URL) of the public service used by the component recorded in the compiling process. After receiving the response result returned by the public service, the response result is integrated with the input information. The specific integration method is as follows: if the input information is of a text type, the response result is connected to the input information after a specific separator (such as “;”) to form new input information; if the input information is structured data (such as JSON format), the response result is added as a new field to the structure of the input information. The component is continuously called based on the integrated information to continue participating in the processing of the input information. If the called component does not use the public service, the input information is directly passed to the component for processing. Through the calling of each component in turn, the output information of the agent application object for the service request is finally obtained.
[0074] The specific logic of processing the input information in the agent application object is as follows: first, the calling order of the components is determined according to the compiling order and the dependency relationship of each topology node in the topology characterization graph. The input information is passed to the first component to be called. If the component is a component using the public service in the public service pool, the response result of the public service is obtained according to the public service calling method described above, and the response result is integrated with the input information and then passed to the component for processing. The component processes the input information according to its own function and logic, for example, if the component is an LLM component, the input information is input to the LLM model as a question, and the LLM model generates an answer according to the pre-trained knowledge and algorithm. The processing result of the component is passed to the next dependent component as new input information, and the above process is repeated until all the components to be called are processed. Finally, the processing result of the last component is taken as the output information of the agent application object for the service request.
[0075] Exemplarily, in the process of calling the initialized components, the dependency relationship between the topological nodes can be determined according to the connection relationship between the topological nodes in the topological representation graph, and then the calling order of the nodes can be determined according to the dependency relationship, and the initialized components corresponding to the topological nodes can be called in turn according to the calling order. In specific application, the components corresponding to the topological nodes without dependency relationship can be called first, and when there are multiple topological nodes without dependent topological nodes, the calling order of these topological nodes can be determined randomly. Taking the topological representation graph in FIG. 4 as an example, assuming that the topological representation graph only includes the topological nodes already shown in FIG. 4, the determined calling order can be Tool A, Tool B, Tool C, Large Model, and Agent.
[0076] In the embodiment, the initialized components in the Agent application object can be called to process the input information according to the topological representation graph, and when the called component is a component using the common service in the common service pool, a service calling request is sent to the common service in the common service pool referenced by the called component, and the processing of the input information is continued based on the response result of the service calling request, so that efficient and stable service can be realized in the case of resource minimization of the Agent application service, a large amount of computing, storage and network resources are saved, and the single-node Agent service response and concurrency capability are improved.
[0077] In some embodiments, the Agent application object generation method of the present application further includes: caching the Agent application object, the topological representation graph, and the compilation parameters corresponding to the topological representation graph, and establishing an association relationship between the Agent application object, the topological representation graph, and the compilation parameters corresponding to the topological representation graph.
[0078] The compilation parameters can be parameters required in the initialization process of each component. For example, for an Agent component, the compilation parameters can be Agent parameters, for an LLM component, the compilation parameters can be LLM parameters, and for a tool component, the compilation parameters can be tool parameters.
[0079] Specifically, for the Agent application object that has been compiled and generated, the server can cache the Agent application object, the topological representation graph used when the Agent application object is compiled, and the compilation parameters used in the compilation process of the topological representation graph, and can establish an association relationship between the Agent application object, the topological representation graph, and the compilation parameters, so that when a service request under a new session is received, the topological representation graph and the compilation parameters can be used to query whether there is a reusable Agent application object in the cache, and if there is, the Agent application object can be directly reused, thereby avoiding time and resource consumption caused by multiple compilations.
[0080] In some embodiments, the service request is a request initiated by the first user under a first session, and the method further includes a response step for the newly initiated service request, as shown in FIG. 6, which includes the following steps:
[0081] In step 602, the session identifier of the first session and the user input information are obtained from the newly initiated service request in response to the newly initiated service request by the first user under the first session.
[0082] The user input information is information input by the user to prompt the agent application object to process the task. For example, the user input information can be a question that the user wants the agent application object to answer. For example, in a legal consultation application scenario, the user input information can be a legal question that the user wants to consult.
[0083] Specifically, the server generates an agent application object based on the service request initiated by the first user under the first session, and the agent application object can be applied to the newly initiated service request by the first user under the first session. When the server creates the first session, the server can generate a session identifier for the first session, which is used to uniquely identify the first session. The server can return the session identifier to the terminal of the first user, so that the terminal of the first user can carry the session identifier in each service request initiated during the first session. Thus, when the server receives the newly initiated service request by the first user under the first session, the server can obtain the session identifier of the first session from the newly initiated service request. The server can also obtain the user input information from the newly initiated service request.
[0084] The session identifier is a unique identifier generated by the server when creating a session, which is used to distinguish different sessions. The terminal carries the session identifier in each service request initiated during the session, so that the server can identify the session. The specific generation method of the session identifier can be to generate a unique string as the session identifier using the UUID (Universally Unique Identifier) algorithm. UUID is a 128-bit identifier composed of numbers and letters, which has global uniqueness and can ensure that the identifier of each session is unique.
[0085] In step 604, the session state information under the first session is obtained according to the session identifier.
[0086] Specifically, the state information of each service request under the first session can be associated with the session identifier of the first session and stored. Thus, after the server obtains the session identifier from the service request, the server can find the session state information associated with the session identifier according to the session identifier.
[0087] Exemplarily, during the first session, the server can run the agent application object in the stateless mode, and store the session state information under the first session in the distributed cache in association with the session identifier corresponding to the first session, so that the server can find the session state information under the first session from the distributed cache according to the session identifier. The specific operation is as follows: assuming that the distributed cache is Redis, the server uses the GET command of Redis to obtain the corresponding session state information from Redis by taking the session identifier as the key. If the obtained information is a JSON string, the server needs to parse it into the corresponding object for subsequent use.
[0088] Exemplarily, during the first session, the server can run the agent application object in the stateless mode, and store the session state information under the first session in the distributed cache in association with the session identifier corresponding to the first session, so that the server can find the session state information under the first session from the distributed cache according to the session identifier. The specific operation is as follows: assuming that the distributed cache is Redis, the server uses the GET command of Redis to obtain the corresponding session state information from Redis by taking the session identifier as the key. If the obtained information is a JSON string, the server needs to parse it into the corresponding object for subsequent use.
[0089] Step 606, determining the input information of the agent application object according to the session state information and the user input information.
[0090] The user input information is information input by the user for prompting the agent application object to process the task, for example, in the legal consultation scenario, it is the legal problem that the user wants to consult, and it is used as the input content when the agent application object is running. The session state information contains the input information of each service request and the state information of the agent application object in each running process, etc. The server determines the session state information corresponding to the current service request according to these information, and saves the session state information in association with the session identifier of the corresponding session to the distributed cache, for subsequent determination of the input information of the agent application object.
[0091] Specifically, the server can concatenate the session state information and the user input information to obtain the final input information of the agent application object. First, the session state information is converted into a string form, and if the session state information is a complex data structure (such as a dictionary, a list, etc.), it can be converted into a string by using JSON serialization or the like. Then, a specific delimiter such as "### " is added between the session state information string and the user input information string, and the two are concatenated to obtain the final input information of the agent application object.
[0092] Step 608, inputting the determined input information into the agent application object to obtain the response result of the newly initiated service request.
[0093] Specifically, the server inputs the final input information of the agent application object into the agent application object, runs the agent application object, and obtains output information of the agent application object for the newly initiated service request, which is the response result of the newly initiated service request.
[0094] For example, in the process of running the agent application object, the server can call the initialized components in the agent application object to process the input information according to the topological characterization map, and obtain the output information of the agent application object for the newly initiated service request; when the called component is a component using a public service in the public service pool, a service call request is sent to the public service in the public service pool referenced by the called component, and the processing of the input information is continued based on the response result of the service call request.
[0095] In the above embodiment, the session identifier of the first session and the user input information are obtained from the newly initiated service request, the session state information under the first session is obtained according to the session identifier, and the input information of the agent application object is determined according to the session state information and the user input information, so that the response result of the newly initiated service request can be determined through the cached agent application object, and the response efficiency for the newly initiated service request is improved.
[0096] In some embodiments, the session state information under the first session is stored in the distributed cache, and the method further includes: saving the response result of the newly initiated service request to the distributed cache to update the session state information under the first session in the distributed cache; and the updated session state information is used to determine the response result of a subsequent service request initiated by the first user under the first session.
[0097] In this embodiment, the agent application object is run in a stateless mode, the session state information under the first session can be stored in the distributed cache, and after obtaining the response result of the newly initiated service request, the server can also save the response result to the distributed cache to update the session state information under the first session in the distributed cache. When receiving a subsequent service request initiated by the first user under the first session, the server can obtain the first session identifier therefrom, and then query the latest updated session state information from the distributed cache according to the first session identifier, so that the final input information of the agent application object can be determined according to the latest session state information, and then the response result of the subsequent service request can be determined through the cached agent application object.
[0098] Specifically, first, the server obtains the current session state information from the distributed cache according to the first session identifier (using the obtaining method described above). If the session state information is the user dialogue record and the agent running state, the response result is added to the user dialogue record as a new dialogue record. For example, if the user dialogue record in the current session state information is ["question 1", "answer 1"], and the response result is "answer 2", the updated user dialogue record is ["question 1", "answer 1", "answer 2"]. Then, the integrated session state information is converted into a JSON string, and the SET command of Redis is used to update the session state information in the distributed cache. When a subsequent service request initiated by the first user under the first session is received, the server can obtain the first session identifier therefrom, and then query the latest updated session state information from the distributed cache according to the first session identifier, so as to determine the final input information of the agent application object according to the latest session state information, and then determine the response result of the subsequent service request through the cached agent application object.
[0099] In the above embodiments, since the session state information can be stored in the distributed cache, and the response result of the newly initiated service request can be saved to the distributed cache, the stateless processing of the agent application object is realized, the reusability and distributed horizontal scalability of the agent application object are improved, and the overall service performance of the agent service is improved.
[0100] In some embodiments, the agent application object generation method of the present application further comprises: in response to a service request initiated by a second user under a second session, obtaining a target topology representation graph and a target compilation parameter indicated by the service request initiated by the second user under the second session; when the target topology representation graph is consistent with the topology representation graph associated with the cached agent application object, and the target compilation parameter is consistent with the compilation parameter associated with the cached agent application object, taking the cached agent application object as a target agent application object under the second session; and determining a response result of the service request initiated by the second user under the second session through the target agent application object.
[0101] The second user and the first user can be the same user or different users. The second session and the first session are different sessions. In the present application, the cached agent application object can be applied to a new session of the same user or a new session of different users. The target topology representation graph is the topology representation graph indicated by the service request initiated by the second user under the second session, which is compared with the topology representation graph associated with the cached agent application object by the server. If they are consistent and the compilation parameters are also consistent, the cached agent application object can be reused.
[0102] The target compilation parameter is a compilation parameter indicated by a service request initiated by the second user in the second session, and the server compares the compilation parameter with a compilation parameter associated with the cached agent application object. If the comparison is consistent with the target topology characterization map, the cached agent application object can be reused.
[0103] Specifically, when receiving a service request initiated by the second user in the second session, the server can obtain an application scenario identifier from the service request, and then query a topology characterization map corresponding to the application scenario identifier, and then compare the topology characterization map with a topology characterization map associated with each cached agent application object. If the application scenario identifier fails to be parsed, the server returns an error message to the terminal of the second user, prompting the user that the service request format is incorrect or lacks necessary information. At the same time, the server can record the abnormal request for subsequent problem troubleshooting and optimization of the service request parsing logic. If the server is configured with a backup parsing rule or a default application scenario identifier, the backup rule can be used to parse or the default application scenario identifier can be used to continue the subsequent process.
[0104] Further, when receiving a service request initiated by the second user in the second session, the server parses the service request according to the format (such as an HTTP request). Assuming that the service request is an HTTP request, the application scenario identifier is stored in the header information of the request, and the key is “Application-Scenario”. The server finds the value of the “Application-Scenario” field by parsing the header information of the HTTP request to obtain the application scenario identifier. If the field is not found in the header information, it is checked whether the application scenario identifier exists in the request body (determined according to the specific business logic and request format). If the application scenario identifier is found, the corresponding topology characterization map is queried according to the identifier, and then the topology characterization map is compared with a topology characterization map associated with each cached agent application object.
[0105] If the topology characterization map is consistent with the topology characterization map associated with any one of the cached agent application objects, the server further obtains a compilation parameter indicated by the service request, compares the compilation parameter indicated by the service request with a compilation parameter associated with the agent application object whose topology characterization map is consistent with the topology characterization map, and if the compilation parameter indicated by the service request is consistent with the compilation parameter associated with the agent application object, the agent application object can be reused, that is, the agent application object is used as a target agent application object in the second session, and then the response result of the service request initiated by the second user in the second session can be determined through the target agent application object.
[0106] When the target topology characterization graph is consistent with the topology characterization graph associated with the cached agent application object, and the target compilation parameter is consistent with the compilation parameter associated with the cached agent application object, the cached agent application object is taken as the target agent application object under the second session. When the cached agent application object is reused, firstly, it is checked whether the state of the agent application object is available (such as whether it has expired, whether it has been damaged, etc.).
[0107] The specific checking method is as follows: for the check of whether it has expired, the server can record the creation time or set an expiration time stamp when caching the agent application object. Before reuse, the current time is obtained and compared with the recorded time, and if the preset validity period is exceeded, it is determined that it has expired. For the check of whether it has been damaged, the server can set verification information, such as a hash value, for the agent application object. The hash value is calculated and recorded when caching, the hash value of the current object is recalculated before reuse and compared with the recorded hash value, and if they are inconsistent, it is determined that it has been damaged.
[0108] If the state is available, necessary initialization or configuration operations are performed according to the specific circumstances of the second session. For example, the session identifier is updated to the session identifier of the second session, and the temporary data of the previous session that may be left over is cleared. Then, according to the service request initiated by the second user under the second session, the input information of the agent application object is determined (such as obtaining the user input information and the session state information). The input information is input into the target agent application object, the agent application object is run, and according to the component calling and processing flow described earlier, the response result of the service request initiated by the second user under the second session is determined. Temporary data is temporary information that may be left over by the agent application object during a session process and is related to the session. When the agent application object is reused in a new session, these temporary data need to be cleared to avoid affecting the processing of the new session.
[0109] For example, when the topology characterization graph comparison is performed, the server can calculate the hash value in the same way for the two topology characterization graphs that need to be compared, and if the calculated hash values are consistent, it can be determined that the two topology characterization graphs are consistent. Specifically, the two topology characterization graphs that need to be compared can be converted into text description information, for example, the node name and node connection relationship can be described in text, and then the same hash algorithm is used to calculate the hash value corresponding to the text description information.
[0110] The text description information is information that converts the topology characterization graph into text form, which records the node name and connection relationship in the order of topology node number by traversing the topology characterization graph, is connected by a specific separator, and is used to calculate the hash value of the topology characterization graph for comparison.
[0111] Specifically, when the topology characterization graph comparison is performed, the server can calculate the hash value in the same way for two topology characterization graphs that need to be compared, and if the calculated hash values are consistent, it can be determined that the two topology characterization graphs are consistent. Specifically, the method of converting the topology characterization graph into text description information is as follows: first, traverse the topology characterization graph, and record the node names in order according to the topology node number (if there is no number, sort). For each node, record all its connection relationships in the format of "node A->node B". Connect all node names and connection relationships with a specific delimiter (such as a newline character) to form text description information. For example, the topology characterization graph has nodes A, B, and C, and the connection relationship is A->B and B->C. The text description information is "A\nB\nC\nA->B\nB->C". Then, the SHA-256 hash algorithm is used to calculate the hash value corresponding to the text description information.
[0112] For example, the compilation parameter indicated by the service request can be a parameter specified by the user when initiating the service request, in which case the server can obtain the compilation parameter by parsing the service request; or the compilation parameter indicated by the service request can also be a default compilation parameter, and if the server does not parse the compilation parameter from the service request, the default parameter set for the topology characterization graph can be used as the compilation parameter indicated by the service request.
[0113] For example, by the target intelligent agent application object, the response result of the service request initiated by the second user under the second session can be specifically: determining the input information of the intelligent agent application object under the second session, inputting the input information into the intelligent agent application object, running the intelligent agent application object in a stateless mode to obtain output information, and the output information is the response result of the service request initiated by the second user under the second session.
[0114] In the above embodiments, when the server receives a new session request, if it is determined that the target topology characterization graph and the target compilation parameter indicated by the session request are consistent with the topology characterization graph and the compilation parameter associated with a certain intelligent application object in the cache, the intelligent application object can be directly reused, thereby greatly saving the computer resources required for the intelligent agent application object compilation process.
[0115] In some embodiments, the topology characterization graph is compiled to initialize components represented by compiled topology nodes, and the initialized components are obtained by: determining node dependency relationships between the topology nodes in the topology characterization graph according to connection relationships between the topology nodes in the topology characterization graph; determining a compilation order of the topology nodes based on the node dependency relationships between the topology nodes; and compiling the topology nodes in the topology characterization graph in the compilation order to initialize the components represented by the compiled topology nodes, and obtaining the initialized components.
[0116] The node dependency relationship is a relationship in which the function implementation of a component represented by a topology node in the topology characterization graph depends on a component represented by another topology node. The topology node pointed to by a connection line between topology nodes in the topology characterization graph needs to depend on an upstream topology node. The compilation order is a compilation sequence of the topology nodes determined according to the node dependency relationships between the topology nodes in the topology characterization graph when the topology characterization graph is compiled. Generally, nodes without upstream dependencies are compiled first. If there are multiple nodes without topology dependencies, the compilation order of the nodes can be determined randomly.
[0117] In this embodiment, the topology characterization graph can be a directed acyclic graph in which topology nodes are connected in a certain order, and the connection relationships between the topology nodes can represent the node dependency relationships between the topology nodes. The specific determination method is as follows: first, an empty dependency relationship dictionary is created to store the dependency information of each topology node. Then, all connection lines in the topology characterization graph are traversed. For each connection line, the starting node of the connection line is taken as a dependent node, and the pointed node of the connection line is taken as a dependency node. If the dependency node is not in the dependency relationship dictionary, an empty list is created for the dependency node in the dictionary, and the dependent node is added to the list. If the dependency node is already in the dictionary, the dependent node is directly added to the corresponding list. In this way, the dependency relationships between the topology nodes are recorded. If a complex circular dependency situation is found in the process of determining the dependency relationships, the server will prompt an error, stop the compilation process, and record the related information of the topology characterization graph. At the same time, the server can provide an analysis tool to help developers find the nodes and paths of the circular dependency, so as to adjust and optimize the topology characterization graph to meet the requirements of the directed acyclic graph.
[0118] The node dependency relationship is a relationship in which the function implementation of a component represented by a topology node needs to depend on a component represented by another topology node. In the topology characterization graph of this embodiment, the topology node pointed to by a connection line between topology nodes in the topology characterization graph is a node that needs to depend on an upstream topology node, for example, the agent node depends on the tool A node, the tool B node, and the tool C node in FIG. 3.
[0119] Specifically, in the process of compiling the topology characterization graph, the server can determine the node dependency relationship between the topology nodes in the topology characterization graph according to the connection relationship between the topology nodes in the topology characterization graph, and determine the compilation order of the topology nodes based on the node dependency relationship between the topology nodes. The specific determination method is as follows: first, an empty compilation order list and an in-degree dictionary are created, and the in-degree dictionary is used to record the number of upstream dependent nodes of each topology node. All connections in the topology characterization graph are traversed, and for each connection, the in-degree value of the pointing node in the in-degree dictionary is increased by 1. Then, all topology nodes with an in-degree of 0 (i.e., no upstream dependent nodes) are found and added to the compilation order list. Next, a node is taken out from the compilation order list, and the topology characterization graph is traversed to find all downstream nodes dependent on the node, and the in-degree value of these downstream nodes in the in-degree dictionary is reduced by 1. If the in-degree value of a downstream node becomes 0, the downstream node is added to the compilation order list. Repeat this process until the compilation order list contains all topology nodes.
[0120] For example, in the process of determining the compilation order, nodes without topology dependency relationship can be sorted first, if there are multiple topology nodes without topology dependency relationship, the compilation order of these topology nodes can be determined randomly, if a topology node has multiple upstream dependent nodes, the compilation order of the topology node is sorted after these dependent nodes, that is, for any two topology nodes, if there is no dependency relationship between the two topology nodes, the order between the two topology nodes can be determined arbitrarily, if there is a node dependency relationship between the two topology nodes, the compilation order of the topology node dependent on the other topology node is sorted before the other topology node. For example, taking the topology characterization graph in FIG. 4 as an example, assuming that the topology characterization graph only includes the topology nodes shown in FIG. 4, the determined compilation order can be tool A, tool B, tool C, large model, and agent.
[0121] Further, after determining the compilation order of each topology node, the server can compile each topology node in the topology characterization graph in order according to the compilation order. After compiling a topology node, the component represented by the topology node is initialized. In the initialization process, if the component represented by the topology node uses a public service in the public service pool, the function of the public service using component using the public service is compiled in the manner of referring to the public service used by the public service using component, and other functions are initialized according to the self-defined initialization of the component. After initialization, the initialized component of the topology node is obtained.
[0122] In the above embodiments, since the compiling order of each topology node can be determined based on the node dependency relationship between each topology node, each topology node in the topology representation graph is compiled in sequence according to the compiling order, the components represented by the compiled topology nodes are initialized, and the initialized components are obtained, which can improve the accuracy in the topology representation graph compiling process.
[0123] In some embodiments, in response to the service request, the topology representation graph indicated by the service request is determined, including: in response to the service request, the topology representation graph indicated by the service request and the compiling parameters are determined; the topology representation graph is compiled, and the components represented by the compiled topology nodes are initialized to obtain the initialized components, including: the topology representation graph is compiled according to the compiling parameters, so as to initialize the components represented by the compiled topology nodes to obtain the initialized components.
[0124] If the compiling parameter related information is missing in the request when the service request is parsed, the server will use the default compiling parameters preset for the topology representation graph. The default compiling parameters are general initialization parameters preset by the server for the topology representation graph when the compiling parameter related information is missing in the request when the service request is parsed, which are set for various components according to a large number of experiments and actual application situations. For example, for an Agent component, the default compiling parameters can include a default running mode, initial resource allocation, etc.; for an LLM component, the default compiling parameters can include a default model version, inference parameter, etc.; for a tool component, the default compiling parameters can include a default calling rule, data processing mode, etc.
[0125] The compiling parameters are parameters required for initialization of each component in the compiling process. The compiling parameters can be parsed from the service request, or can be default parameters corresponding to the topology representation graph indicated by the service request.
[0126] Specifically, after receiving the service request, the server can determine the topology representation graph and the compiling parameters indicated by the service request by parsing the service request. In the process of compiling the topology representation graph, each time a topology node is compiled, the server initializes the component represented by the topology node according to the compiling parameters corresponding to the component represented by the topology node, thereby obtaining the initialized component corresponding to the topology node.
[0127] In the above embodiments, by determining the compiling parameters, the topology representation graph is compiled according to the compiling parameters, the components represented by the compiled topology nodes are initialized, and the initialized components are obtained, which can improve the compiling accuracy.
[0128] In some specific embodiments, the application further provides an application scenario, in which the agent application of the application is an LLM-based agent application.
[0129] The popularity of ChatGPT and GPT4 has shown the potential of large language models (LLM) in various fields. In the education sector, LLMs are becoming powerful tools for students and teachers, providing personalized learning suggestions and instant answers. In the business sector, companies are using LLMs to automate customer service, improve efficiency, and analyze market trends and consumer sentiment through natural language processing. In the creative industry, writers and artists are using LLMs to inspire new works. In programming and technical services, LLMs are helping developers debug code and even write programs. The impact of LLMs goes beyond these areas. They are changing the way people access information and communicate, and even in professional fields such as law and medicine, LLMs are beginning to play a role in assisting analysis and decision-making. With the continuous advancement of technology, the application of LLMs will become more extensive, and its far-reaching impact will penetrate every corner of society, reshaping all aspects of work and life.
[0130] LLM can endow programs with intelligence, providing intelligent analysis, planning, execution, and reflection. By building Agent applications, combining LLM and tool libraries in the business field, LLM can autonomously analyze, think, and plan based on user input, accurately call tools to execute tasks at each stage, and reflect on the results of each stage to plan the next stage of action. The whole process is iteratively run by LLM without human intervention until the final result is output. Based on LLM, Agent can more efficiently solve various business scenarios, such as data search, analysis, and insight autonomous execution.
[0131] Agent applications greatly expand the application scenarios of LLMs. How to build efficient Agent application services is a pressing problem. The current mainstream method is to connect various components and LLMs in Agent applications through a DAG graph. Each node in the DAG graph represents various tools, LLMs, and other objects. Then compile the components in the DAG graph, initialize the LLM objects and component services, and finally get a runnable Agent object to provide Agent application services externally. Since the request parameters of each Agent service may differ, this method requires recompiling and building a new DAG graph for different users or even different sessions for the same user. Although the same DAG graph can be used to build corresponding Agent application services for different requests from different users, there is a large amount of DAG compilation and construction, which results in a waste of service resources and a long compilation and construction time when the number of request users is large.
[0132] Based on this, the embodiments of the present application aim at the problems of waste of computing, storage and network resources, and time-consuming of frequent DAG compilation and construction of the current LLM Agent service framework, and propose an agent application object generation method. In the method, an Agent application object, a common service, and a Session state separation Agent service framework are provided. The new Agent service framework can greatly solve the waste of computing, storage and network resources, greatly reduce the time-consuming of DAG compilation and construction, and significantly improve the Agent service response capability under a single node.
[0133] In a specific implementation, on the one hand, the application state and the common service are separated from the Agent application object, the Agent application object is minimized, the fast DAG compilation and construction of the Agent application object are implemented, the common service calling is more stable, and the user request is responded more quickly; on the other hand, the state data related to the Agent application Session is stored in a distributed cache, the common service adopts a unified resource pool for centralized management, the Agent application object that has been compiled and constructed is greatly reused between different users and different Sessions, the Agent application service realizes efficient and stable service in the case of resource minimization, a large amount of computing, storage and network resources are saved, and the single-node Agent service response and concurrency capability are improved.
[0134] The basic architecture of the Agent is shown in FIG. 3. The final node of the DAG graph in FIG. 3 is the Agent, which depends on the service nodes and various tool nodes provided by the LLM in the upstream, the various tools depend on the database nodes and search service nodes, and the whole is compiled and constructed into an Agent application object to provide external services. The Agent application object depends on the state in which it is located and external input, analyzes, plans and makes decisions based on the LLM, and calls the Tool that meets the demand to execute a specific task to obtain the task execution result as feedback.
[0135] Referring to FIG. 7, an Agent application service framework in the related art is shown. In the related art, the Agent provides services in the manner shown in FIG. 7. For different Session requests of each user, due to differences in Agent application construction parameters, such as differences in Agent parameters, LLM parameters, and application tool parameters, Agent application objects need to be independently compiled for each Session based on Agent DAG. Each Agent application object generates data under the current Session in the running process, including user conversation data, Agent state data, and other data, collectively referred to as Session state, which needs to be saved in the Agent application object. From the perspective of the service framework, it appears that each independent Agent application object provides services for different users and different Sessions. In many cases, even different Sessions of the same user need to recompile and build Agent application objects because of differences in Session state or Agent application construction parameters.
[0136] As can be seen from FIG. 7, in the related art, each Agent application object is bound with runtime data, which causes the compiled Agent to be unable to provide services for different Sessions. Differences in different Agent construction parameters cause Agent application objects to often need to be recompiled and built, and services such as LLM, Database, and Search Engine, which consume a large amount of network resources, cannot be shared. The above problems cause Agent application objects to often need to be recompiled and built from scratch for different Sessions of each user, resulting in waste of various resources, inability to share services, time-consuming DAG graph compilation and construction, and other problems.
[0137] The technical solution proposed in the present application can effectively solve the above problems of frequent compilation and construction of Agent application objects, and can solve the problems of Agent application object reuse, inability to share public service resources, and significantly reducing the time-consuming of Agent application object compilation and construction.
[0138] By separating the Session state from the Agent application object and storing it in a distributed cache, such as Redis, the Agent application object can be stateless, and the Agent application object can be reused in different users and different Sessions under the condition that the Agent construction parameters are the same.
[0139] By assembling the public services, such as LLM services, database connection services, search engine services, and the like, which are shareable public services into a public service Pool, the LLM and tools under each Agent application object use the unified public services to complete the corresponding request tasks, and it is not necessary to apply resources for each Agent application object to complete such request tasks, thereby saving a large amount of computing, storage, and network resources.
[0140] After the above two steps are completed, the compilation and construction of the Agent application object will become more easy and fast, and only the part of the Agent application object itself needs to be compiled and constructed, and the occupied computing and storage resources are small, thereby realizing the fast compilation and construction of the Agent DAG graph.
[0141] Referring to FIG. 8, under some embodiments, the Agent service framework provided by the agent application object generation method of the present application is shown. Compared with the Agent service framework in the related art, two main contents are added, which are as follows:
[0142] 1. Session data cache: using a distributed cache, such as Redis, to store the session state information generated by the Agent application object of different sessions into the distributed cache, thereby decoupling the Agent application object from the Session session data, realizing the stateless running of the Agent application object, and achieving the maximum reuse of the Agent application object between different users and sessions. For example, in FIG. 8, the agent application object can be reused between session A and session C.
[0143] 2. Common Services Pool: a public service resource pool, which separates the public services occupying a large amount of computing, storage, and network resources from the Agent application object to form an independent service. The advantages are mainly as follows: first, the DAG compilation and construction time is greatly saved, and in general, database connection, network resource initialization, and the like are time-consuming, and there are service stability guarantee problems in separate calling; second, a large amount of computing, storage, and network resources are saved, and different Agents only refer to the objects of the resource pool in the compilation and construction process, and are not independently occupied, and are applied and released when used, thereby realizing the sharing of public service resources, and greatly improving the resource utilization efficiency.
[0144] After the Agent service framework is adjusted, the compilation and construction of the DAG graph and the Agent application object service process will be changed accordingly, which will be described in detail.
[0145] Referring to FIG. 9, the specific compilation and construction of the Agent application object under the new Agent service framework is shown in FIG. 9, which includes the following steps:
[0146] 1. First determine whether the DAG graph of the Agent has been compiled and built under the current compilation parameters. If so, it can be reused, and jump to step 8. Otherwise, go to step 2 and recompile and build.
[0147] 2. Topologically sort the DAG graph of the Agent application to determine the order of compilation and construction of each node, generally starting with the node that has no upstream dependencies.
[0148] 3. Select nodes for initialization in order according to the topological sorting result. First determine whether the node has been initialized. If so, jump directly to step 7.
[0149] 4. Determine whether the initialization of the node depends on a common service, such as an LLM request, database connection, search engine service, etc. If so, jump to step 5. Otherwise, jump to step 6.
[0150] 5. During node initialization, the initialization of the common service object directly references the common service, without the need to reinitialize the corresponding service. The rest is performed according to the node's self-defined initialization.
[0151] 6. Complete the initialization according to the node's self-defined initialization.
[0152] 7. Determine whether all nodes have completed initialization. If so, jump to step 8. Otherwise, jump to step 3.
[0153] 8. End the Agent application compilation and construction process and output the Agent application object.
[0154] Under the new Agent service framework, the Agent application object caches the Session state information during the service process in the distributed cache. The Agent application object itself runs in a stateless manner and its output is entirely dependent on external input each time. This stateless working mode allows the Agent application object to be reused between different users and Sessions, and the application service can be easily horizontally expanded to run on multiple nodes to support more user access. Referring to FIG. 10, in some exemplary embodiments, the Agent application object service flowchart is shown in FIG. 10 and includes the following steps:
[0155] 1. Parse the Agent service request to obtain the current Session ID and user input information.
[0156] 2. Extract the session state information from the distributed cache according to the Session ID, including historical session information, Agent running state, etc.
[0157] 3. Based on the session state information and the current input information of the user, the final input information is obtained by combination and splicing, and the input information is taken as the input of the Agent application object.
[0158] 4. The Agent application object is run to process the final input information, and the execution result, i.e., the output information, is obtained.
[0159] 5. According to the output information, the cached session state information is updated.
[0160] 6. The output information is returned to the terminal of the user.
[0161] 7. End.
[0162] In summary, in the technical scheme proposed in the present application, the Agent application DAG graph is disassembled, and the public service is stripped out, which significantly reduces the consumption of initializing various services in the DAG graph compilation and construction process, realizes the reuse of public services, improves the service stability, and caches the compiled and constructed Agent application object for reuse, saving computing, storage and network resources. In the service process of the Agent application object, the execution result and the state information are cached, the stateless processing of the Agent application object is realized, the reusability and distributed horizontal scalability of the Agent application object are improved, and the overall service performance of the Agent service is improved.
[0163] The present application can be widely applied to Agent application scenarios based on LLM construction, such as medical, financial, transportation, consulting, big data and other application scenarios. Any scenario that needs LLM to build Agent and call various tools to solve problems independently can adopt the present scheme to obtain more efficient implementation.
[0164] For example, in a legal consultation scenario, when a user initiates legal consultation for a certain problem, a new session can be triggered to start in the server, the server generates an agent application object according to the topological representation graph corresponding to the consultation scenario, thereafter, the user can have multiple rounds of dialogue with the agent application object for the problem, in each dialogue, the server saves the input information of the user, the state information of the agent application object itself and the output information of the agent application object and other session state information in the distributed cache, when receiving the input information of the next dialogue of the user, the session state information of the user is obtained from the cache, the final input information of the agent application object is obtained by combining the input information of this time, the output information of the agent application object under this session is obtained by running the agent application object, and the output information is returned to the terminal of the user.
[0165] For another example, in a medical data analysis scenario, a user can input task description information, such as predicting the development trend of a certain disease, trigger a new session to be started in a server, and the server generates an agent application object according to a topology representation graph corresponding to the medical scenario, processes and analyzes massive medical data through the agent application object, and returns the development trend information of the disease to the terminal of the user.
[0166] It should be understood that, although each step in the flowchart involved in the above embodiments is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in the above embodiments can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.
[0167] Based on the same inventive concept, the embodiments of the present application also provide an agent application object generation apparatus for implementing the above-mentioned agent application object generation method. The implementation scheme for solving the problem provided by the apparatus is similar to the implementation scheme described in the above method, and therefore the specific limitations in one or more agent application object generation apparatus embodiments provided below can refer to the limitations of the agent application object generation method in the foregoing, which will not be described here again.
[0168] In some embodiments, as shown in FIG. 11, an agent application object generation apparatus 1100 is provided, which includes:
[0169] A topology representation graph determination module 1102 is configured to determine a topology representation graph indicated by a service request in response to the service request, the topology representation graph including a plurality of topology nodes, each topology node representing a component, and at least one component using a common service in a common service pool;
[0170] A topology representation graph compilation module 1104 is configured to compile the topology representation graph to initialize the components represented by the compiled topology nodes to obtain initialized components; and when the common service used by the common service used component uses a common service in the common service pool, the topology representation graph compilation module 1104 is configured to compile the function of the common service used by the common service used component by referencing the common service used by the common service used component in the common service pool.
[0171] The agent application object obtaining module 1106 is configured to, in response to completion of compiling the topology characterization graph, obtain an agent application object determined based on the initialized components corresponding to the topology nodes in the topology characterization graph.
[0172] The agent application object generation apparatus described above, in response to a service request, determines a topology characterization graph indicated by the service request, the topology characterization graph including a plurality of topology nodes, each topology node characterizing a component, at least one component using a common service in a common service pool, compiles the topology characterization graph to initialize the components characterized by the compiled topology nodes, obtains the initialized components, in response to completion of compiling the topology characterization graph, obtains an agent application object determined based on the initialized components corresponding to the topology nodes in the topology characterization graph, and when the compiled topology nodes characterize a common service using component using a common service in a common service pool, compiles the function of the common service using component using the common service in the common service pool in a manner of referencing the common service used by the common service using component in the common service pool, thus stripping the common service from the agent application object, so that only the part of the agent application object itself needs to be compiled and constructed during generation of the agent application object, and the occupied computing and storage resources are smaller, realizing fast compilation and construction of the agent application object, thereby improving the generation efficiency of the agent application object.
[0173] In some embodiments, the service request is a request initiated by a first user in a first session, and the agent application object generation apparatus further includes an agent application object running module configured to run the agent application object in a stateless mode to obtain output information of the agent application object for the service request, take the output information as a response result of the service request, and return the response result to a terminal of the first user.
[0174] In some embodiments, the agent application object running module is further configured to determine input information of the agent application object, input the input information to the agent application object to run the agent application object to obtain output information of the agent application object for the service request, and determine session state information in the first session according to the input information and state information in the running process of the agent application object, and save the session state information in association with a session identifier of the first session in the distributed cache.
[0175] In some embodiments, the agent application object running module is further configured to: input input information into the agent application object; invoke each initialized component in the agent application object to process the input information according to the topology characterization graph, and obtain output information of the agent application object for the service request; and when the invoked component is a component using a common service in the common service pool, send a service invocation request to the common service in the common service pool referenced by the invoked component, and continue to participate in processing of the input information based on a response result of the service invocation request.
[0176] In some embodiments, the agent application object generation apparatus further comprises an agent application object caching module configured to cache the agent application object, the topology characterization graph, and the compilation parameters corresponding to the topology characterization graph, and establish an association between the agent application object, the topology characterization graph, and the compilation parameters corresponding to the topology characterization graph.
[0177] In some embodiments, the service request is a request initiated by a first user in a first session, and the agent application object generation apparatus further comprises a service request response module configured to: in response to a newly initiated service request by the first user in the first session, obtain a session identifier of the first session and user input information from the newly initiated service request; obtain session state information of the first session according to the session identifier; determine input information of the agent application object according to the session state information and the user input information; and input the determined input information into the agent application object to obtain a response result of the newly initiated service request.
[0178] In some embodiments, the agent application object generation apparatus further comprises a response result caching module configured to save the response result of the newly initiated service request into a distributed cache to update session state information of the first session in the distributed cache, and the updated session state information is used to determine a response result of a subsequent service request initiated by the first user in the first session.
[0179] In some embodiments, the service request response module is further configured to: in response to a service request initiated by a second user in a second session, obtain a target topology characterization graph and target compilation parameters indicated by the service request initiated by the second user in the second session; when the target topology characterization graph is consistent with a topology characterization graph associated with a cached agent application object, and the target compilation parameters are consistent with compilation parameters associated with the cached agent application object, use the cached agent application object as a target agent application object for the second session; and determine a response result of the service request initiated by the second user in the second session through the target agent application object.
[0180] In some embodiments, the topology characterization graph compiling module is further configured to determine node dependency relationships between the topology nodes in the topology characterization graph according to the connection relationships between the topology nodes in the topology characterization graph; determine a compiling order of the topology nodes based on the node dependency relationships between the topology nodes; and compile the topology nodes in the topology characterization graph in sequence according to the compiling order, to initialize the components represented by the compiled topology nodes, and obtain the initialized components.
[0181] In some embodiments, the topology characterization graph determining module is further configured to determine the topology characterization graph and the compiling parameters indicated by the service request in response to the service request; and the topology characterization graph compiling module is further configured to compile the topology characterization graph according to the compiling parameters, to initialize the components represented by the compiled topology nodes, and obtain the initialized components.
[0182] In some embodiments, at least one of the components represented by the plurality of topology nodes is a pre-trained large language model.
[0183] The above various modules in the intelligent agent application object generation apparatus can be all or partially implemented by software, hardware, or a combination thereof. The above various modules can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the above various modules.
[0184] In some embodiments, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as shown in FIG. 12. The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement the intelligent agent application object generation method provided in the embodiments of the present application.
[0185] In some embodiments, a computer device is provided, which can be a terminal, and its internal structure diagram can be as shown in FIG. 13. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, mobile cellular network, NFC (near field communication) or other technologies. The computer program is executed by the processor to implement the agent application object generation method provided in the embodiments of the present application. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device. It can also be an external keyboard, touchpad or mouse, etc.
[0186] Those skilled in the art can understand that the structure shown in FIG. 12 and FIG. 13 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0187] In some embodiments, a computer device is provided, which includes a memory and a processor, and the memory stores a computer program. The processor executes the computer program to implement the steps of the above-mentioned agent application object generation method.
[0188] In some embodiments, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by the processor to implement the steps of the above-mentioned agent application object generation method.
[0189] In some embodiments, a computer program product is provided, which includes a computer program. The computer program is executed by the processor to implement the steps of the above-mentioned agent application object generation method.
[0190] In summary, the application provides an agent application object generation method and device, computer equipment, computer readable storage medium and computer program product. By responding to a service request, a topology representation graph indicated by the service request is determined, the topology representation graph includes a plurality of topology nodes, each node representing a component, and at least one component using a common service in a common service pool. In the prior art, the common service of each component needs to be initialized separately when the agent application object is generated, which not only consumes a large amount of computing resources, but also increases the time cost due to repeated initialization. The application separates the common service from the agent application object to form a common service pool. When compiling the topology representation graph, when compiling to the component using the common service, the function of using the common service of the component is compiled in the way of referencing the common service in the common service pool. For example, in an LLM-based agent application, the services such as databases and search engines that the components such as LLM and various tools depend on can form a common service pool, and these common services are directly referenced during compilation without reinitialization. In this way, the generation process only needs to compile the part related to the agent application object itself, reducing the occupation of computing and storage resources, avoiding the waste of resources caused by repeated initialization, realizing the rapid compilation and construction of the agent application object, and significantly improving the generation efficiency.
[0191] Further, when the service request is a request initiated by the first user in the first session, the agent application object is run in a stateless mode, the output information of the agent application object for the service request is obtained, and the output information is returned to the terminal of the first user as a response result. In the traditional mode, the agent application object saves state data, and the state data of different sessions are independent of each other, which makes it difficult to reuse the agent application object. In the stateless mode, no state data is saved, and the session state information can be saved to a local database or carried in the response result and returned to the user terminal. For example, the server saves the session state information in the session process to the local database, and when the next service request in the first session is initiated, the session state information can be queried from the local database. This way makes the agent application object reusable between different users and different sessions, avoids creating an agent application object for each session, improves the utilization of resources, and reduces the resource overhead of the system.
[0192] Further, when the agent application object running in the stateless mode obtains the output information, the input information is determined first, the input information is input into the agent application object to run and obtain the output information, and the session state information in the first session is determined according to the input information and the state information in the running process, and is associated with the session identifier of the first session and saved into the distributed cache. In the traditional mode, the session state information is bound to the agent application object, which is not conducive to data sharing and management. However, when the distributed cache such as Redis is used to store the session state information, centralized management and rapid access of data can be realized. When multiple users or sessions need to access the session state information, the session state information can be directly obtained from the distributed cache, avoiding repeated storage and management of data, realizing the separation and decoupling of the agent application object and the session state information, and maximizing the reuse of the agent application object between different users and different sessions. In addition, the distributed cache facilitates the sharing and management of state data, and improves the scalability and response speed of the system.
[0193] Further, when the input information is input into the agent application object to run and obtain the output information, each initialized component is called to process the input information according to the topological representation graph. When the called component uses the public service in the public service pool, a service call request is sent to the referenced public service, and the processing of the input information is continued based on the response result. In the traditional technology, each component needs to be requested and initialized separately when using the public service, which leads to resource waste and service response time increase. However, by referring to the public service in the public service pool, the application avoids repeated requests and initialization. For example, when the called component is a component using the public service in the public service pool, the server sends a service call request to the public service according to the access location of the public service used by the component recorded in the compilation process, and continues to process after integrating the response result with the input information. This way can realize efficient and stable service with minimal resources, save a lot of computing, storage and network resources, improve the single-node Agent service response and concurrency capability, make the system able to handle more requests, and improve the performance and efficiency of the system.
[0194] Further, the agent application object, the topology characterization map and the corresponding compilation parameters are cached, and the association relationship therebetween is established. In the traditional mode, each time a new session request is received, the topology characterization map needs to be recompiled to generate the agent application object, which causes a large amount of time and resource consumption. Through the caching mechanism, when a new session service request is received, whether there is a reusable agent application object can be queried in the cache according to the topology characterization map and the compilation parameters, and if so, the agent application object is directly reused. For example, when the topology characterization map and the compilation parameters of the new service request are consistent with those in the cache, the agent application object in the cache can be directly used, avoiding the time and resource consumption caused by multiple compilations, and improving the response speed and resource utilization of the system.
[0195] Further, when the first user newly initiates a service request under the first session, the session identifier and the user input information of the first session are obtained from the request, the session state information is obtained according to the session identifier, the input information of the agent application object is determined in combination with the user input information, and the input information is input into the agent application object to obtain a response result. In the traditional mode, for a newly initiated service request, the agent application object and the session state information may need to be reinitialized and obtained, resulting in a complex processing flow and a long response time. Through the caching of the agent application object and the session state information, when a new service request is initiated, the session state information can be quickly obtained and the input information can be determined, and the cached agent application object is directly used for processing. For example, the server quickly obtains the session state information from the distributed cache according to the session identifier, splices the user input information to serve as the input of the agent application object, determines the response result through the cached agent application object, improves the response efficiency for the newly initiated service request, and reduces the user waiting time.
[0196] Further, if the session state information under the first session is stored in the distributed cache, the response result of the newly initiated service request is saved to the distributed cache to update the session state information, and the updated information is used to determine the response result of the subsequent service request initiated by the first user under the first session. In the conventional technology, the update and management of the session state information is relatively complex, which may cause data inconsistency and management difficulty. However, the present application can realize real-time update and consistency management of data by storing and updating the session state information in the distributed cache. For example, the server adds the response result of the newly initiated service request as a new dialogue record to the user dialogue record, and updates the session state information in the distributed cache. When receiving the subsequent service request, the input information of the agent application object can be determined according to the latest session state information, and the response result can be determined through the cached agent application object. This way realizes the stateless processing of the Agent application object, improves its reusability and distributed horizontal scalability, and further improves the overall service performance of the Agent service, so that the system can better cope with high concurrency and large-scale service requests.
[0197] Further, when the second user initiates a service request under the second session, the target topology characterization map and the target compilation parameter are obtained, and if they are consistent with the topology characterization map and the compilation parameter associated with the cached agent application object, the cached agent application object is used as the target agent application object to determine the response result of the service request. In the conventional mode, for service requests of different users or sessions, it is usually necessary to recompile and generate an agent application object, which causes waste of resources and increase of processing time. However, the present application can directly reuse the cached agent application object when the target topology characterization map and the target compilation parameter are consistent with those in the cache through the caching and comparison mechanism. For example, the server compares the hash values of the topology characterization maps, and if they are consistent and the compilation parameters are also consistent, the cached agent application object is reused. In this way, the agent application object can be directly reused, which greatly saves the computer resources required in the compilation process of the agent application object, and improves the resource utilization and response speed of the system.
[0198] Further, when compiling the topology characterization graph to initialize the component, the node dependency relationship is determined according to the connection relationship between the topology nodes, and the compilation order is determined based on this. Each topology node is compiled in sequence to obtain the initialized component. In the traditional compilation method, the dependency relationship between the topology nodes may not be considered, resulting in chaotic compilation order, which may cause component initialization failure or repeated initialization problems. By determining the node dependency relationship and the compilation order, the correct initialization of the component is ensured. For example, nodes without upstream dependencies are compiled first. If there are multiple nodes without topology dependencies, the compilation order can be determined randomly. This can improve the accuracy of the topology characterization graph compilation process, avoid errors caused by improper compilation order, and improve the success rate and efficiency of compilation.
[0199] Further, in response to a service request, the topology characterization graph and the compilation parameters indicated by the service request are determined. When compiling the topology characterization graph, the component represented by the compiled topology node is initialized according to the compilation parameters to obtain the initialized component. In the traditional technology, there may be no explicit compilation parameters, which increases the uncertainty of component initialization. By determining the compilation parameters, the application provides clear guidance for component initialization. For example, different compilation parameters are set for different types of components, such as Agent components, LLM components, and tool components, to ensure that the components are initialized according to the correct parameters. This can improve the accuracy of compilation, reduce initialization failures caused by parameter errors, and improve the quality of the generated agent application object.
[0200] Further, among the components represented by the plurality of topology nodes, at least one component is a pre-trained large language model. The pre-trained large language model has strong language understanding and generation capabilities, and can endow the agent application object with intelligent analysis, planning, execution, and reflection capabilities. For example, in an LLM-based agent application, the LLM can analyze and plan according to user input problems, accurately call each tool to execute each stage task, and reflect on the execution result of each stage and plan the next stage action. The whole process is iteratively run autonomously without human intervention until the final result is output. The LLM-based Agent can more efficiently solve various business scenario problems, such as data search, analysis, and insight autonomous execution of the whole link, expanding the functionality and application range of the agent application and improving the intelligent level and processing capacity of the agent application.
[0201] In addition, under the new Agent service framework proposed in the present application, the DAG graph of the Agent application is topologically sorted, the dependency relationship between nodes is analyzed, the compilation and construction order of each node is determined, and the node without upstream dependency is initialized. This method ensures the order of component initialization, avoids initialization errors caused by chaotic dependency relationship, and improves the success rate and efficiency of compilation and construction. During node initialization, for nodes that depend on public services, the public services are directly referenced without the need to reinitialize the corresponding services, and the other parts are initialized according to the node self-defined initialization. This further reduces the repeated initialization of public services, saves resources and time, and improves the stability and response speed of the services. At the same time, under the new Agent service framework, the Agent application object caches the Session state information in the service process to the distributed cache, and itself runs in a stateless mode, and its output completely depends on each external input. This stateless working mode enables the Agent application object to be reused between different users and Sessions, and the application service can be easily horizontally expanded to run on multiple nodes to support more user access, thereby improving the scalability and concurrent processing capability of the system.
[0202] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0203] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0204] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0205] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of the present application. Therefore, the scope of protection of the patent of the present application should be subject to the appended claims.
Claims
1. An agent application object generation method, executed by a computer device, comprising: determining, in response to a service request, a topology characterization graph indicated by the service request, the topology characterization graph comprising a plurality of topology nodes, each topology node characterizing a component, at least one of the components using a common service in a common service pool; compiling the topology characterization graph to initialize the components characterized by the compiled topology nodes, to obtain initialized components; wherein, when a common service using component characterized by a compiled topology node uses a common service in the common service pool, the common service using component uses a function of the common service in the common service pool by referencing the common service used by the common service using component; and obtaining, in response to completion of the compilation of the topology characterization graph, an agent application object determined based on the initialized components corresponding to the topology nodes in the topology characterization graph. 2.The method of claim 1, wherein the service request is a request initiated by a first user in a first session, and the method further comprises: running the agent application object in a stateless mode to obtain output information of the agent application object for the service request; returning the output information as a response result of the service request to a terminal of the first user. 3.The method of claim 2, wherein the running the agent application object in a stateless mode to obtain output information of the agent application object for the service request comprises: determining input information of the agent application object; inputting the input information to the agent application object to run the agent application object to obtain output information of the agent application object for the service request; determining session state information of the first session according to the input information and state information during running of the agent application object, and saving the session state information in association with a session identifier of the first session in a distributed cache. 4.The method of claim 3, wherein the inputting the input information to the agent application object to run the agent application object to obtain output information of the agent application object for the service request comprises: inputting the input information to the agent application object; invoking each initialized component in the agent application object to process the input information according to the topology characterization graph to obtain output information of the agent application object for the service request; wherein, when the invoked component is a component using a common service in a common service pool, sending a service invocation request to the common service in the common service pool referenced by the invoked component, and continuing to participate in processing of the input information based on a response result of the service invocation request. 5.The method of any one of claims 1 to 4, further comprising: Cache the agent application object, the topology characterization graph and the compiling parameters corresponding to the topology characterization graph, and establish the association between the agent application object, the topology characterization graph and the compiling parameters corresponding to the topology characterization graph. 6.The method of claim 5, wherein the service request is a request initiated by a first user in a first session, and the method further comprises: in response to a newly initiated service request by the first user in the first session, obtaining a session identifier of the first session and user input information from the newly initiated service request; obtaining session state information of the first session according to the session identifier; determining input information of the agent application object according to the session state information and the user input information; inputting the determined input information into the agent application object to obtain a response result of the newly initiated service request. 7.The method of claim 6, wherein the session state information of the first session is stored in a distributed cache, and the method further comprises: saving the response result of the newly initiated service request into the distributed cache to update the session state information of the first session in the distributed cache; and the updated session state information is used to determine a response result of a subsequent service request initiated by the first user in the first session. 8.The method of any one of claims 5 to 7, wherein the method further comprises: in response to a service request initiated by a second user in a second session, obtaining a target topology characterization graph and target compiling parameters indicated by the service request initiated by the second user in the second session; when the target topology characterization graph is consistent with the topology characterization graph associated with the cached agent application object, and the target compiling parameters are consistent with the compiling parameters associated with the cached agent application object, using the cached agent application object as a target agent application object for the second session; and determining a response result of the service request initiated by the second user in the second session by using the target agent application object. 9.The method of any one of claims 1 to 8, wherein the compiling of the topology characterization graph is to initialize components represented by compiled topology nodes, and obtaining the initialized components comprises: determining node dependency relationships between topology nodes in the topology characterization graph according to connection relationships between the topology nodes in the topology characterization graph; determining a compiling order of the topology nodes based on the node dependency relationships between the topology nodes; and compiling the topology nodes in the topology characterization graph in the compiling order to initialize components represented by the compiled topology nodes, and obtaining the initialized components. 10.The method of any one of claims 1 to 9, wherein the determining of the topology characterization graph indicated by the service request comprises: determining the topology characterization graph and compiling parameters indicated by the service request in response to the service request. The topology characterization graph is compiled, and the components represented by the compiled topology nodes are initialized to obtain initialized components, including: The topology characterization graph is compiled according to the compilation parameters, so that the components represented by the compiled topology nodes are initialized to obtain initialized components.
11. The method of any one of claims 1-10, wherein at least one of the components represented by the plurality of topology nodes is a pre-trained large language model.
12. An agent application object generation apparatus, comprising: a topology characterization graph determination module configured to determine, in response to a service request, a topology characterization graph indicated by the service request, the topology characterization graph including a plurality of topology nodes, each topology node representing a component, and at least one of the components using a common service in a common service pool; a topology characterization graph compilation module configured to compile the topology characterization graph to initialize the components represented by the compiled topology nodes to obtain initialized components; wherein when a common service using component represented by a compiled topology node uses a common service in the common service pool, the common service using component uses the function of the common service in the common service pool by referencing the common service used by the common service using component; and an agent application object obtaining module configured to obtain, in response to completion of the compilation of the topology characterization graph, an agent application object determined based on the initialized components corresponding to the topology nodes in the topology characterization graph.
13. A computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps of the method of any one of claims 1-11 when executing the computer program.
14. A computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the method of any one of claims 1-11.
15. A computer program product comprising a computer program, the computer program being executed by a processor to implement the steps of the method of any one of claims 1-11.
Citation Information
Patent Citations
Service scheduling method and device, server and computer readable storage medium
CN113434283A
Intelligent agent construction and development method based on industry digitization
CN114356300A
Topological relation graph generation method and device, electronic equipment and storage medium
CN116405399A
Method for efficiently configuring and interacting public services in service bus
CN117033033A
Intelligent agent application object generation method and device, computer equipment and storage medium
CN118605884A