Intelligent agent deployment method and computer program product

Through the cellular architecture, the hardware resource cluster is split into general-purpose and GPU computing resource pools for the research and development of intelligent agents and the deployment of operation units. This solves the problems of resource management and user experience in intelligent agent deployment, and achieves efficient utilization and low-latency and stable user experience.

CN120704860APending Publication Date: 2025-09-26CHINA MOBILE GROUP ZHEJIANG +3
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510688201.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The current intelligent body deployment solution has difficulty in effectively managing and utilizing resources in To G or To B business scenarios, resulting in high user experience delays, poor stability, and the risk of single point failures and resource waste.

Method used

A cellular architecture is used to split the hardware resource cluster into a general computing resource pool and a GPU computing resource pool, which are used for the deployment of R&D units and intelligent units respectively. Combined with intelligent routing strategies and asynchronous data synchronization, efficient resource utilization and low-latency stability are achieved.

Benefits of technology

It improves resource utilization, meets low-latency and highly stable user experience, reduces the risk of single point failure, simplifies the expansion process, and improves system availability and business stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704860A_ABST
    Figure CN120704860A_ABST
Patent Text Reader

Abstract

The invention discloses an agent deployment method and a computer program product, belongs to the field of information technology support, and is used for reasonably utilizing resources and meeting low-delay and high-stability user experience. The method comprises the steps that a virtual resource pool of a hardware resource cluster is split into a first computing power resource pool and a second computing power resource pool, and the quota of general computing power resources of the first computing power resource pool is larger than the quota of general computing power resources of the second computing power resource pool; the quota of the graphics processor resources of the second computing power resource pool is greater than the quota of the graphics processor resources of the first computing power resource pool; deploying a research and development unit of the intelligent agent through the first computing power resource pool, wherein the research and development unit has an intelligent agent research and development function; and deploying an agent unit of the agent through the second computing power resource pool, wherein the agent unit comprises an agent operation function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of information technology support, and specifically relates to a deployment method and computer program product of an intelligent body. Background Art

[0002] An AI agent is an intelligent entity capable of perceiving its environment, making decisions, and executing actions. It uses a large language model (LLM) as its core computing engine, enabling it to engage in conversation, execute tasks, reason, and exhibit a degree of autonomy. Currently, the most popular solution for AI agents is the Retrieval-Augmented Generation (RAG) framework, which leverages an external knowledge base to significantly improve the quality and accuracy of the output from large models.

[0003] Intelligent agents typically have large model sizes and require substantial computing resources to support training and inference.

[0004] Current intelligent agent solutions in the industry tend to address the challenges of large-scale model deployment by focusing on improving resource utilization while ensuring a basic user experience. In these scenarios, users prefer low-latency, highly stable services over extreme resource utilization. Therefore, effective resource management and utilization while ensuring a low-latency, highly stable user experience remains a challenge in current intelligent agent deployments. Summary of the Invention

[0005] The embodiments of the present application provide a method for deploying an intelligent agent and a computer program product, which can solve the problem that the current intelligent agent deployment cannot effectively manage and utilize resources and meet the user's low-latency and highly stable user experience.

[0006] In a first aspect, an embodiment of the present application provides a method for deploying an intelligent agent, the method comprising: splitting the virtual resource pool of a hardware resource cluster into a first computing power resource pool and a second computing power resource pool, the quota of general computing power resources of the first computing power resource pool being greater than the quota of general computing power resources of the second computing power resource pool, and the quota of graphics processor resources of the second computing power resource pool being greater than the quota of graphics processor resources of the first computing power resource pool; deploying the R&D unit of the intelligent agent through the first computing power resource pool, the R&D unit including the intelligent agent R&D function; deploying the intelligent agent unit of the intelligent agent through the second computing power resource pool, the intelligent agent unit including the intelligent agent operation function.

[0007] In the second aspect, an embodiment of the present application provides a deployment device for an intelligent agent, which includes: a splitting module for splitting the virtual resource pool of a hardware resource cluster into a first computing power resource pool and a second computing power resource pool, the quota of general computing power resources of the first computing power resource pool is greater than the quota of general computing power resources of the second computing power resource pool, and the quota of graphics processor resources of the second computing power resource pool is greater than the quota of graphics processor resources of the first computing power resource pool; a first deployment module for deploying the R&D unit of the intelligent agent through the first computing power resource pool, the R&D unit including the intelligent agent R&D function; a second deployment module for deploying the intelligent agent unit of the intelligent agent through the second computing power resource pool, the intelligent agent unit including the intelligent agent operation function.

[0008] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.

[0009] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0010] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer performs the steps of the method described in the first aspect.

[0011] In an embodiment of the present application, the virtual resource pool of the hardware resource cluster is split into a first computing resource pool and a second computing resource pool, the quota of general computing resources of the first computing resource pool is greater than the quota of general computing resources of the second computing resource pool, and the quota of graphics processor resources of the second computing resource pool is greater than the quota of graphics processor resources of the first computing resource pool; the R&D unit of the intelligent body is deployed through the first computing resource pool, and the R&D unit includes the intelligent body R&D function; the intelligent body unit of the intelligent body is deployed through the second computing resource pool, and the intelligent body unit includes the intelligent body operation function, thereby improving the resource utilization efficiency and being able to meet the low-latency and highly stable user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is a flow chart of a method for deploying an intelligent agent provided in an embodiment of the present application; Figure 2 This is a schematic diagram of resource pool splitting provided by an embodiment of the present application; Figure 3 This is a schematic diagram of an intelligent agent architecture provided by an embodiment of the present application; Figure 4 This is a schematic diagram of an intelligent agent deployment architecture provided by an embodiment of the present application; Figure 5 This is a schematic diagram of the separation of a research and development unit and an intelligent body unit provided in an embodiment of the present application; Figure 6 This is a schematic diagram of the structure of a deployment device for an intelligent agent provided in an embodiment of the present application; Figure 7 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0013] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0014] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0015] When it comes to solving the challenges of "large-scale" model deployment, the current intelligent agent solutions in the industry tend to "improve resource utilization while only ensuring basic user experience." The business characteristics and challenges of To G or To B intelligent agent business scenarios are completely opposite to the mainstream "large-scale, single model" business scenarios. On the one hand, for reasons of security and confidentiality, users tend to build their own clusters and use the same Maas platform to implement all aspects of the entire process of intelligent agent research and development, release, and inference deployment. On the other hand, the intelligent agents released by users based on the Maas platform have the characteristics of "one many and one small", that is, the number of intelligent agents is large and the number of users of a single intelligent agent is small.

[0016] In this type of business scenario, users prefer low-latency, highly stable services rather than extreme resource utilization. While the aforementioned technologies have greatly improved resource utilization and met users' requirements for low-latency, highly stable user experience, their benefits will decline or even fail in ToG and ToB business scenarios. These shortcomings include: Traffic throttling can impair user experience for some users: While traffic throttling can increase the number of users enjoying the service, it limits the experience of individual users, which is inversely proportional to the frequency and duration of use. In B2B or B2G business scenarios, users cannot tolerate service interruptions caused by traffic throttling.

[0017] Intelligent routing is limited to scheduling the same inference instance: A key prerequisite for intelligent routing scheduling strategies is the predictability of request resource consumption. Based on the assumption that "the resources consumed by a single request are similar," the algorithm can estimate which instance can handle the corresponding request without causing performance degradation. Therefore, a key requirement for using the intelligent routing mode is that "multiple inference instances of the same model are deployed in the same service group." However, in To B and To G intelligent agent business scenarios, there are many intelligent agent models, and the amount of GPU resources required for a single user request under different models varies greatly, causing the effectiveness of the intelligent routing algorithm to decline or even fail.

[0018] Resource reuse leads to increased latency: Using intelligent routing or the Mooncake architecture for resource reuse can significantly reduce server deployment costs, but it also increases latency for some users. This is because scheduling triggers a cold start of model instances in the cluster, and user requests experience significant latency during a cold start. Furthermore, for cost considerations, scheduling strategies maximize resource reuse, but resource reuse can lead to GPU contention, increasing average request latency.

[0019] Microservices architecture introduces single point of risk: The product-as-a-service deployments offered by the aforementioned platforms and vendors all utilize a microservices architecture, which incorporates common middleware to achieve unified data storage, data communication, and resource scheduling capabilities. However, due to the existence of this middleware, a failure in any particular piece of middleware can impact the entire system. Furthermore, due to the resource-intensive nature of this model, recovery costs are high.

[0020] In response to the above problems, this application focuses on To B or To G business scenarios, and proposes an intelligent platform architecture and deployment method based on cellular architecture to solve the problems of managing and utilizing computing resources and meeting low-latency and highly stable user experience.

[0021] Compared with the microservice architecture, the cellular architecture introduces more detailed isolation and autonomous "units", controls the risk range of single point failures in the microservice architecture, provides higher isolation, reduces performance degradation caused by resource competition, and provides business stability and dynamic expansion capabilities far higher than the microservice architecture.

[0022] Inference servers deployed based on GPUs are usually equipped with high-performance CPUs, memory, and hard disks. However, the GPU is the primary consumer of model inference, resulting in waste of these CPU, memory, and disk resources. This application draws on Mooncake's split architecture to split the MaaS platform's training and deployment functions and the deployment of intelligent products, corresponding to CPU-intensive and memory-intensive clusters and GPU-intensive clusters, respectively. It also introduces a cellular architecture combined with intelligent routing strategies to implement resource scheduling, thereby achieving more reasonable resource utilization, a better user experience, and greater stability.

[0023] The following, in conjunction with the accompanying drawings, describes in detail the deployment method and computer program product of the intelligent agent provided in the embodiments of the present application through specific embodiments and their application scenarios.

[0024] Figure 1 An embodiment of the present application provides a method for deploying an agent. The method can be performed by an electronic device, which may include a server and / or a terminal device. In other words, the method can be performed by software or hardware installed on the electronic device. The method includes the following steps: Step S102: Split the virtual resource pool of the hardware resource cluster into a first computing resource pool and a second computing resource pool.

[0025] like Figure 2 As shown, the embodiment of the present application adopts a separate resource allocation scheme at the resource layer, splitting the virtual resource pool of a homogeneous or heterogeneous hardware resource cluster into a first computing resource pool and a second computing resource pool, wherein the general computing resource quota of the first computing resource pool is greater than the general computing resource quota of the second computing resource pool, and the graphics processing unit (GPU) resource quota of the second computing resource pool is greater than the graphics processing unit (GPU) resource quota of the first computing resource pool. In the embodiment of the present application, the first computing resource pool can also be referred to as the general computing resource pool, and the second computing resource pool can also be referred to as the GPU computing resource pool.

[0026] In one implementation, the general computing resources include: central processing unit resources, memory resources and hard disk resources.

[0027] Specifically, in this embodiment of the present application, the general computing resources of the first computing resource pool (general computing resource pool) are mainly central processing unit (CPU) resources, memory resources, and hard disk resources, serving computing or IO-intensive applications. The second computing resource pool (GPU computing resource pool) is mainly GPU resources, serving GPU-intensive applications such as large model training or inference.

[0028] In one implementation, splitting the virtual resource pool of the hardware resource cluster into a first computing resource pool and a second computing resource pool includes: performing a virtualization operation on the hardware resource cluster to obtain the virtual resource pool; splitting the virtual resource pool into a first part of resources and a second part of resources, allocating the first part of resources to a first namespace to obtain the first computing resource pool; allocating the second part of resources to a second namespace to obtain the second computing resource pool.

[0029] Specifically, splitting the virtual resource pool of the hardware resource cluster into the first computing power resource pool and the second computing power resource pool includes the following steps: the first step is to deploy the K8S cluster on the hardware resource cluster to realize the virtualization operation of the resources and obtain the virtual resource pool. Through this operation, the resource quotas of all CPU resources, GPU resources, memory resources, and hard disk resources are obtained. The second step is to define two namespaces (namespaces) based on the namespace mechanism of K8S, namely "first namespace (general)" and "second namespace (agent)", and then split the above-mentioned virtual resource pool into the first part of resources and the second part of resources. The first part of resources is allocated to the first namespace to obtain the first resource pool, and the second part of resources is allocated to the second namespace to obtain the second resource pool. Due to the resource isolation characteristics of the namespace, the present application obtains the advantages of hardware resource sharing and virtual resource isolation at the same time.

[0030] This application virtualizes and then breaks up homogeneous or heterogeneous hardware clusters, splitting the resource pool into two parts: a general computing resource pool that focuses on CPU resources, memory resources, and hard disk resources, and a GPU computing resource pool that focuses on GPUs. These two parts are used to serve general computing-intensive businesses and GPU-intensive inference businesses, respectively. The CPU and memory resources that were previously wasted in GPU clusters are also utilized to achieve more reasonable use of cluster resources.

[0031] Step S104: Deploy the R&D unit of the intelligent agent through the first computing resource pool.

[0032] The embodiment of the present application proposes that the intelligent agent has the following Figure 3 The system architecture shown in the figure is divided into two cellular units from the user's perspective: the R&D unit and the intelligent unit. Figure 2As shown, in an embodiment of the present application, the first computing resource pool is used to deploy the R&D unit.

[0033] In the embodiment of the present application, the main object of deployment is the "unit". The unit contains the logical functions, databases, models and other software entities required to implement the unit function body, and the unit gateway is used to receive external requests and request external capabilities. All the involved logical functions are unitized and then organically organized to form a honeycomb form, that is, a honeycomb architecture. The deployment architecture is as follows: Figure 4 As shown, this deployment architecture includes an R&D unit entity. This unit serves R&D personnel and has a low level of concurrent business operations. It is typically deployed using a single instance in the first computing resource pool. It contains large models, databases, and microservices, with the unit gateway receiving user front-end requests.

[0034] like Figure 3 As shown, the R&D unit integrates all functional modules involved in the R&D process and fully manages every aspect of the R&D process. Specifically, the R&D unit includes agent development functionality. This unit provides R&D users with: operations for adding, deleting, modifying, and querying agents; operations for creating, modifying, and integrating models with agent knowledge bases; configuration and integration of tool plug-ins that agents rely on; and the ability to assemble models, knowledge bases, tools, and other tools through process orchestration. Furthermore, this unit maintains the foundational capabilities necessary to implement these business functions, including model access management, database management, knowledge acquisition and processing, and user permission management. Inputs to this unit include, but are not limited to, knowledge text, tool plug-in APIs, training data, and R&D personnel's operations on the platform. The output is a complete definition of an agent application (including the agent frontend or interface, the agent's internal operational processes, the knowledge base and tools that the agent relies on, and the underlying model that the agent relies on). Based on the different inputs in the R&D process, developers can use this unit to output a variety of agent application definitions with varying capabilities.

[0035] Step S106: deploying the agent units of the agent through the second computing resource pool.

[0036] like Figure 2 As shown, in the embodiment of the present application, the second computing resource pool is used to deploy intelligent units. Figure 4 As shown, the deployment architecture includes an intelligent unit entity. In this embodiment of the application, the intelligent unit of the intelligent agent can be deployed through the second computing resource pool. There are usually multiple intelligent units, and each intelligent unit can provide completely different services to users. If a certain intelligent agent business requires a high service concurrency, multiple identical replica units can be deployed to improve overall performance through horizontal expansion.

[0037] like Figure 3As shown, the agent unit includes agent operation functions. An agent unit is a specific instance of an agent application and a concrete implementation of the complete definition of the agent application output by the R&D unit. It integrates all the functions and data required for the operation of the agent application and can independently provide complete agent capabilities and services to users. Specifically, the agent unit includes an agent front-end or interface gateway, the agent's internal operation process, the agent-dependent knowledge base and tools, and the agent-dependent foundation model. The agent front-end or interface gateway serves as the entry point for the agent, allowing users to access the agent through the front-end interface or through RESTful API calls. The internal operation process of the agent unit is for the agent to drive each atomic capability node to implement business logic based on the workflow definition. This includes configuration data for the workflow definition, the docking configuration of the atomic capabilities, and the process engine responsible for execution. The knowledge base that the agent relies on refers to knowledge data that can be directly used by the agent, including database entities that support the use of the knowledge base, vector search capability interfaces, etc.

[0038] The tools that the agent relies on refer to specific business capabilities obtained by calling external interfaces through RESTful APIs or code snippets and plug-ins injected by R&D into the platform. These include the aforementioned configurations and the environment in which they run. The base model that the agent relies on refers to the large language model called by the agent core, which also includes all other components required for model runtime. The input to this unit includes user requests, knowledge base information synchronized by the R&D unit, and information passed in by external interface tools. The output is the agent's response.

[0039] The deployment method of the intelligent body provided in the embodiment of the present application splits the virtual resource pool of the hardware resource cluster into a first computing power resource pool and a second computing power resource pool, the quota of general computing power resources of the first computing power resource pool is greater than the quota of general computing power resources of the second computing power resource pool, and the quota of graphics processor resources of the second computing power resource pool is greater than the quota of graphics processor resources of the first computing power resource pool; the R&D unit of the intelligent body is deployed through the first computing power resource pool, and the R&D unit includes the intelligent body R&D function; the intelligent body unit of the intelligent body is deployed through the second computing power resource pool, and the intelligent body unit includes the intelligent body operation function, which can solve the problem of improving resource utilization efficiency and meet the low-latency and highly stable user experience.

[0040] In the embodiment of the present application, all functional modules required for R&D can be integrated into R&D units according to business divisions, and the functional modules required by intelligent agents can be integrated into intelligent agent units. Although there are certain duplicate functions in the two types of units, since the functions within the unit are complete and there is no need to call services across units, all user requests can be completed within the unit, which greatly reduces the data transmission cost and request call delay, while isolating the intelligent agent unit and the R&D unit, and the faults between different intelligent agent units. In addition, since the infrastructure between units has clear boundaries, this fundamentally eliminates the resource competition and security issues that exist under the microservice architecture. If the units are physically isolated (deployed in different computer rooms), the system-level reliability can be further improved.

[0041] like Figure 5 The separation of the R&D unit and the intelligent body unit shown in the figure, this application is different from the unified system of the current intelligent body platform in the industry. This application separates the R&D unit and the intelligent body unit logically and physically. After separation, the logic of the platform on the user side is clearer and more focused, and it is easier to control the user's permissions. On the resource side, the independent reuse of CPU, memory, and GPU can be achieved by using the resource pool separation solution. Although in the architecture of this application, multiple functional entities such as models, knowledge bases, and databases exist in the R&D unit and the intelligent body unit at the same time, the infrastructure, software, and data of the intelligent body unit are logical subsets of the R&D unit and are physically independent. This application introduces a new cellular architecture, which breaks up the organization of the intelligent body platform and reconstructs it into a cellular deployment mode. Through the implementation of the cellular architecture, better availability capabilities are obtained.

[0042] In one implementation, after splitting the virtual resource pool of the hardware resource cluster into a first computing power resource pool and a second computing power resource pool, it also includes: a control unit for deploying the intelligent body through the first computing power resource pool, and the control unit includes an intelligent routing unit, a resource management unit, an operation and maintenance management unit, and a data synchronization unit.

[0043] In the embodiments of this application, Figure 4 As shown, the deployment architecture also includes a control unit, which includes an intelligent routing unit, a resource management unit, an operations and maintenance management unit, and a data synchronization unit. All control unit types in this architecture can be deployed using multiple replica units. Due to the sidecar model, business units have a weak dependency on the control unit. Multiple replica deployment can greatly improve availability and service capacity.

[0044] like Figure 3As shown, the operation and maintenance management unit aggregates the capabilities related to operation and maintenance, integration, and release, and is positioned to connect the R&D unit and the intelligent agent unit. It is responsible for instantiating the output of the R&D unit and deploying it into the intelligent agent unit. Specifically, the unit obtains the complete definition of the intelligent agent application from the R&D unit, schedules the unit resources required for the intelligent agent definition through the collaborative resource management unit, and then reads the relevant configuration through scripts or IaC tools to complete the creation and deployment operations of the intelligent agent unit's dependent infrastructure. Finally, the collaborative data synchronization unit synchronizes the operating data information that the intelligent agent depends on from the R&D unit to the intelligent agent unit. The input of this unit includes the complete definition of the intelligent agent application and the basic resource information provided by the resource management unit. It can be triggered manually by the R&D or by the resource management module, and the output is an intelligent agent unit instance.

[0045] like Figure 3 As shown, the data synchronization unit is the data synchronization pipeline, specifically implemented as asynchronous data synchronization middleware. This unit's data synchronization consists of two links: asynchronous synchronization of data received from the R&D unit to the agent unit, and data synchronization between units of the same type (synchronization between R&D units and synchronization between agent units). This unit adopts a passive, asynchronous design and is loosely coupled with the main business unit, eliminating the possibility of fault correlation.

[0046] like Figure 3 As shown in the figure, the resource management unit integrates resource control and traffic scheduling capabilities. Specifically, it implements resource capacity and performance awareness, unit fault detection, intelligent routing, and user permission management. This unit is similar to the sidecar framework in microservices. It physically decouples resource and traffic control from the business application plane, isolating the business application's dependence on the control plane and thus achieving fault isolation.

[0047] The embodiment of the present application realizes the separation of the data layer control plane and the business plane: in the microservice architecture, the data layer adopts the PaaS model for unified management, realizes sharing from the bottom layer data, and the data synchronization of the business layer is strongly integrated with PaaS. The present application goes a step further and separates the data synchronization link to the control unit (control plane). Since this architecture splits the data layer into various units, the sharing and coordination of data between different units should not be handled by the business unit itself (synchronous calls will be formed). Introducing a data synchronization unit as a third party for processing decouples the strong coupling relationship between units due to data. In addition, since the data synchronization unit adopts an asynchronous synchronization mode, it provides retry, error tolerance, fault isolation and other mechanisms to improve the stability of the system.

[0048] In one implementation, after the intelligent body unit of the intelligent body is deployed through the second computing power resource pool, it also includes: when a user's access request is received, obtaining the uniform resource locator of the access request; when the uniform resource locator matches a preset target functional path, determining the target gateway according to the target gateway address corresponding to the target functional path; forwarding the access request to the target gateway, the target gateway being the gateway of the R&D unit or the intelligent body unit.

[0049] like Figure 4 As shown in the figure, users request to access services using a unified domain name. After all traffic passes through load balancing, it first passes through the intelligent routing unit. This unit determines which business unit the request should be directed to based on routing rules such as tenant and business type, and then forwards the request to the corresponding unit gateway.

[0050] In an embodiment of the present application, after the R&D unit and the intelligent body unit complete the deployment, they register the customized function path and gateway IP with the resource management module. For example, the R&D unit is defined as ( / devops / xxx; 192.168.1.1) and the intelligent body unit is located as ( / agent / xxx, 192.168.2.1). The intelligent routing unit stores the above binary data in a high-performance cache. The user requests to access the service using a unified domain name, and all traffic must first pass through the intelligent routing unit after load balancing. When the user's access request is received, the uniform resource locator URL in the access request is extracted, and the uniform resource locator is matched with the target function path in the cache. When the uniform resource locator matches the preset target function path, the target gateway is determined according to the target gateway address corresponding to the target function path; then the access request is forwarded to the target gateway, which is the gateway of the R&D unit or the intelligent body unit.

[0051] This application introduces an intelligent routing strategy at the request scheduling layer, which enables dynamic distribution of user requests, that is, it can route requests to the correct function provider and achieve balanced resource utilization by adjusting the traffic to each unit.

[0052] In one implementation, after the intelligent body unit of the intelligent body is deployed through the second computing power resource pool, it also includes: when receiving an access request from a user, obtaining the tenant identifier in the access request; when the tenant identifier matches a preset target tenant identifier, determining the target gateway according to the target gateway address corresponding to the target tenant identifier; forwarding the access request to the target gateway, which is the gateway of the intelligent body unit.

[0053] In an embodiment of the present application, a tenant identifier (ID) can also be used as a basis for traffic distribution. The intelligent routing unit stores the tuple information of the intelligent unit (target tenant identifier ID, target gateway address IP) in a high-performance cache. The premise of using this strategy is that the tenant ID needs to be reported as additional information on the user client. After receiving the user's access request, the tenant identifier ID provided in the request can be extracted. The tenant identifier is matched with the target tenant identifier in the tuple in the cache. If the match is successful, the corresponding target gateway address IP is obtained, and then the access request is routed to the target gateway corresponding to the target gateway address. Since the routing rules of this strategy are manually assigned and static, it can ensure that the traffic of a certain unit is fixed, which is suitable for scenarios with high requirements for customer experience.

[0054] In one implementation, a capacity threshold of the intelligent unit is obtained; when the traffic of access requests distributed to the intelligent unit is greater than the capacity threshold, the intelligent unit is expanded by deployment to carry the traffic of the access requests.

[0055] Specifically, this strategy can be used for groups formed by multiple instance units (multiple copies) of the same business. The resource management unit regularly reports resource load information, and the intelligent routing unit can grasp the unit performance indicators and the level of incoming traffic requests. It predicts traffic patterns and demand for unit resources through reinforcement learning algorithms, and continuously adjusts traffic allocation strategies based on service performance and resource usage to achieve optimal resource utilization. In addition, a capacity threshold can be set for the intelligent unit. When the incoming traffic exceeds the traffic threshold that all instances of the current intelligent unit can carry, the elastic expansion operation is immediately triggered. Since the performance of the intelligent unit is fixed and predictable, the number of intelligent units required for expansion can be linearly calculated without complex evaluation. The expansion operation can be completed in minutes, thereby achieving lossless traffic transfer.

[0056] In one implementation, after the intelligent body units of the intelligent body are deployed through the second computing power resource pool, it also includes: when a change in the cluster deployment structure is detected, new intelligent body units are added, intelligent body units are expanded, intelligent body units are reduced, or intelligent body units are replaced according to the change in the cluster deployment structure.

[0057] In the implementation of this application, if Figure 4As shown, cluster resource information is stored in the resource management unit. When the resource management unit detects changes in the cluster deployment structure, it pushes the corresponding change information to the intelligent routing unit. The intelligent routing unit includes a high-performance, highly available routing policy storage service, which updates static routing rules immediately after receiving cluster resource changes. There are three situations that trigger intelligent routing rule adjustments: adding a new intelligent unit, expanding an intelligent unit, reducing an intelligent unit, or replacing an intelligent unit.

[0058] The process of creating, scaling, and replacing intelligent body units: When new intelligent body services are added and intelligent body units are replaced, it is usually triggered by the R&D unit. The expansion and reduction of intelligent body units are triggered by the resource management unit. When the event is triggered, the operation and maintenance management unit will be activated first. The operation and maintenance management unit will first obtain the original definition of the unit to be deployed from the R&D unit, and obtain the available computing power resource information from the resource management module. Then, it will prepare the intelligent body instance based on the internal integration process of the operation and maintenance management as input, and complete the deployment action on the computing power resources based on the IaC tool. After the basic deployment is completed, the data synchronization operation will be triggered. When the unit is ready, the operation and maintenance management unit will notify the resource management unit to update the resource information.

[0059] In one implementation, after the agent unit of the agent is deployed through the second computing resource pool, the method further includes: synchronizing the data of the R&D unit to the agent unit through the data synchronization unit.

[0060] In the application examples, Figure 4 As shown, data synchronization primarily occurs when an agent unit is created or updated, with data flowing from the R&D unit to the newly created agent unit. Data synchronization occurs asynchronously, with the R&D unit pushing data to the data synchronization unit, and the agent unit consuming data from the data synchronization unit using a pull method.

[0061] The embodiments of the present application improve resource utilization without compromising user experience: Through a separate resource pool solution, CPU and memory resources previously wasted in GPU clusters are freed up for general-purpose computing-intensive business units, improving resource utilization. Intelligent routing solutions flexibly adjust traffic to each unit to achieve optimal resource utilization. Due to the isolation and rapid expansion of cellular units, user requests can be ensured to always be processed within a reasonable range of unit performance, avoiding degradation of user experience caused by competition for underlying resources.

[0062] The embodiments of this application provide higher system availability than microservice architectures: this application confines business functions and their dependencies to the unit, and the control plane adopts a weakly coupled sidecar model, thereby eliminating the strong communication dependencies between different units and avoiding the fault correlation effects caused by strong dependencies. Compared to the impact of underlying failures on a wide range of businesses in a microservice architecture, the architecture of this application can ensure that the scope of failure impact is always limited to the unit, greatly reducing the scope of the explosion of single-point risks.

[0063] The embodiments of the present application provide a simpler and more flexible capacity expansion mechanism: the present application unitizes business functions, and units of the same type naturally form strong consistency in resource requirements, performance, and business capacity, thereby enabling rapid replication (deployment) of units through standard automated deployment mechanisms. Due to the strong cohesion of the units, the business traffic that the units can carry is fixed, so the evaluation of the business traffic after the unit expansion can be calculated simply by simple linear addition. Compared to the microservice architecture, which requires evaluation through full-link stress testing and other means after expansion, this architecture simplifies the process, thereby providing faster elastic capacity expansion capabilities.

[0064] It should be noted that the agent deployment method provided in the embodiments of the present application can be executed by the agent deployment device, or by a control module in the agent deployment device for executing the method xxx. In the embodiments of the present application, the agent deployment method executed by the agent deployment device is used as an example to illustrate the agent deployment device provided in the embodiments of the present application.

[0065] Figure 6 Schematic diagram of the structure of the deployment device of the intelligent body according to the embodiment of the present application. Figure 6 As shown, the deployment device 600 of the intelligent agent includes: a splitting module 610, a first deployment module 620 and a second deployment module 630.

[0066] The splitting module 610 is used to split the virtual resource pool of the hardware resource cluster into a first computing power resource pool and a second computing power resource pool, wherein the quota of general computing power resources of the first computing power resource pool is greater than the quota of general computing power resources of the second computing power resource pool, and the quota of graphics processor resources of the second computing power resource pool is greater than the quota of graphics processor resources of the first computing power resource pool; the first deployment module 620 is used to deploy the R&D unit of the intelligent agent through the first computing power resource pool, and the R&D unit includes the intelligent agent R&D function; the second deployment module 630 is used to deploy the intelligent agent unit of the intelligent agent through the second computing power resource pool, and the intelligent agent unit includes the intelligent agent operation function.

[0067] In one implementation, the splitting module 610 is used to perform virtualization operations on the hardware resource cluster to obtain the virtual resource pool; split the virtual resource pool into a first part of resources and a second part of resources, allocate the first part of resources to a first namespace to obtain the first computing power resource pool; allocate the second part of resources to a second namespace to obtain the second computing power resource pool.

[0068] In one implementation, the first deployment module 620 is also used to deploy the control unit of the intelligent entity through the first computing power resource pool. The control unit includes an intelligent routing unit, a resource management unit, an operation and maintenance management unit, and a data synchronization unit.

[0069] In one implementation, the first deployment module 620 is further used to obtain the uniform resource locator of the access request when receiving an access request from the user; determine the target gateway according to the target gateway address corresponding to the target functional path when the uniform resource locator matches the preset target functional path; and forward the access request to the target gateway, which is the gateway of the R&D unit or the intelligent unit.

[0070] In one implementation, the first deployment module 620 is further used to obtain the tenant identifier in the access request when receiving the user's access request; determine the target gateway based on the target gateway address corresponding to the target tenant identifier when the tenant identifier matches the preset target tenant identifier; and forward the access request to the target gateway, which is the gateway of the intelligent unit.

[0071] In one implementation, the first deployment module 620 is further configured to obtain a capacity threshold of the intelligent unit; when the traffic of access requests distributed to the intelligent unit is greater than the capacity threshold, the intelligent unit is expanded by deployment to carry the traffic of the access requests.

[0072] In one implementation, the first deployment module 620 is further configured to, upon detecting a change in the cluster deployment structure, add new intelligent units, expand the capacity of intelligent units, reduce the capacity of intelligent units, or replace intelligent units according to the change in the cluster deployment structure.

[0073] In one implementation, the first deployment module 620 is further configured to synchronize the data of the R&D unit to the intelligent unit via the data synchronization unit.

[0074] In one implementation, the general computing resources include: central processing unit resources, memory resources and hard disk resources.

[0075] The deployment device of the intelligent agent in the embodiments of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc., which are not specifically limited in the embodiments of the present application.

[0076] The deployment device of the intelligent agent in the embodiment of the present application can be a device having an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0077] The deployment device of the intelligent body provided in the embodiment of the present application can achieve Figures 1 to 5 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0078] like Figure 7 As shown, an embodiment of the present application further provides an electronic device 700, including a processor 701 and a memory 702, wherein the memory 702 stores a program or instruction that can be run on the processor 701, and when the program or instruction is executed by the processor 701, it is implemented as follows: the virtual resource pool of the hardware resource cluster is split into a first computing power resource pool and a second computing power resource pool, the quota of general computing power resources of the first computing power resource pool is greater than the quota of general computing power resources of the second computing power resource pool, and the quota of graphics processor resources of the second computing power resource pool is greater than the quota of graphics processor resources of the first computing power resource pool; the R&D unit of the intelligent body is deployed through the first computing power resource pool, and the R&D unit includes an intelligent body R&D function; the intelligent body unit of the intelligent body is deployed through the second computing power resource pool, and the intelligent body unit includes an intelligent body operation function.

[0079] In one implementation, the hardware resource cluster is virtualized to obtain the virtual resource pool; the virtual resource pool is split into a first part of resources and a second part of resources, the first part of resources is allocated to a first namespace to obtain the first computing power resource pool; the second part of resources is allocated to a second namespace to obtain the second computing power resource pool.

[0080] In one implementation, after the virtual resource pool of the hardware resource cluster is split into a first computing resource pool and a second computing resource pool, a control unit of the intelligent body is deployed through the first computing resource pool, and the control unit includes an intelligent routing unit, a resource management unit, an operation and maintenance management unit, and a data synchronization unit.

[0081] In one implementation, after the intelligent body unit of the intelligent body is deployed through the second computing power resource pool, when a user's access request is received, the uniform resource locator of the access request is obtained; when the uniform resource locator matches the preset target functional path, the target gateway is determined according to the target gateway address corresponding to the target functional path; the access request is forwarded to the target gateway, and the target gateway is the gateway of the R&D unit or the intelligent body unit.

[0082] In one implementation, after the intelligent body unit of the intelligent body is deployed through the second computing power resource pool, when receiving a user's access request, the tenant identifier in the access request is obtained; when the tenant identifier matches the preset target tenant identifier, the target gateway is determined according to the target gateway address corresponding to the target tenant identifier; the access request is forwarded to the target gateway, and the target gateway is the gateway of the intelligent body unit.

[0083] In one implementation, after the intelligent body unit of the intelligent body is deployed through the second computing power resource pool, the capacity threshold of the intelligent body unit is obtained; when the traffic of the access request distributed to the intelligent body unit is greater than the capacity threshold, the intelligent body unit is expanded by deployment to carry the traffic of the access request.

[0084] In one implementation, after the intelligent body units of the intelligent body are deployed through the second computing power resource pool, when a change in the cluster deployment structure is detected, new intelligent body units are added, intelligent body units are expanded, intelligent body units are reduced, or intelligent body units are replaced according to the change in the cluster deployment structure.

[0085] In one implementation, after the agent unit of the agent is deployed through the second computing resource pool, the data of the R&D unit is synchronized to the agent unit through the data synchronization unit.

[0086] In one implementation, the general computing resources include: central processing unit resources, memory resources and hard disk resources.

[0087] The specific execution steps can refer to the various steps of the above-mentioned intelligent agent deployment method embodiment, and can achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0088] It should be noted that the electronic devices in the embodiments of the present application include: servers, terminals, or other devices other than terminals.

[0089] The above electronic device structure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. For example, the input unit may include a graphics processing unit (GPU) and a microphone, and the display unit may be configured as a display panel in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit includes at least one of a touch panel and other input devices. A touch panel is also called a touch screen. Other input devices may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be detailed here.

[0090] The memory can be used to store software programs and various data. The memory may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory may include volatile memory or non-volatile memory, or the memory may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct rambus random access memory (DRRAM).

[0091] The processor may include one or more processing units; optionally, the processor may integrate an application processor and a modem processor, wherein the application processor primarily handles operations related to the operating system, user interface, and application programs, and the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into the processor.

[0092] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned intelligent body deployment method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0093] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as ROM, RAM, magnetic disk or optical disk.

[0094] An embodiment of the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the various processes of the embodiment of the above-mentioned intelligent agent deployment method and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0095] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0096] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of this application.

[0097] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A method for deploying an intelligent agent, characterized in that: include: Splitting the virtual resource pool of the hardware resource cluster into a first computing resource pool and a second computing resource pool, wherein the general computing resource quota of the first computing resource pool is greater than the general computing resource quota of the second computing resource pool, and the graphics processing unit resource quota of the second computing resource pool is greater than the graphics processing unit resource quota of the first computing resource pool; Deploy an agent R&D unit through the first computing resource pool, wherein the R&D unit includes an agent R&D function; The agent unit of the agent is deployed through the second computing power resource pool, and the agent unit includes an agent operation function.

2. The deployment method according to claim 1, characterized in that: Splitting the virtual resource pool of the hardware resource cluster into a first computing resource pool and a second computing resource pool includes: Performing a virtualization operation on the hardware resource cluster to obtain the virtual resource pool; Splitting the virtual resource pool into a first portion of resources and a second portion of resources, Allocate the first portion of resources to the first namespace to obtain the first computing resource pool; Allocate the second part of resources to the second namespace to obtain the second computing power resource pool.

3. The deployment method according to claim 1, wherein: After splitting the virtual resource pool of the hardware resource cluster into the first computing resource pool and the second computing resource pool, the method further includes: The control unit of the intelligent body is deployed through the first computing power resource pool, and the control unit includes an intelligent routing unit, a resource management unit, an operation and maintenance management unit, and a data synchronization unit.

4. The deployment method according to claim 1, wherein: After deploying the agent unit of the agent through the second computing resource pool, the method further includes: Upon receiving an access request from a user, obtaining a uniform resource locator (URL) of the access request; In the case where the uniform resource locator matches the preset target function path, determining the target gateway according to the target gateway address corresponding to the target function path; The access request is forwarded to the target gateway, where the target gateway is the gateway of the R&D unit or the intelligent unit.

5. The deployment method according to claim 1, wherein: After deploying the agent unit of the agent through the second computing resource pool, the method further includes: When receiving an access request from a user, obtaining a tenant identifier in the access request; When the tenant identifier matches the preset target tenant identifier, determining the target gateway according to the target gateway address corresponding to the target tenant identifier; The access request is forwarded to the target gateway, which is the gateway of the intelligent unit.

6. The deployment method according to claim 1, wherein: After deploying the agent unit of the agent through the second computing resource pool, the method further includes: Obtaining a capacity threshold of the intelligent unit; In a case where the traffic of access requests distributed to the intelligent unit is greater than the capacity threshold, the intelligent unit is expanded by deployment to carry the traffic of the access requests.

7. The deployment method according to claim 1, characterized in that: After deploying the agent unit of the agent through the second computing resource pool, the method further includes: When a change in the cluster deployment structure is detected, new intelligent units are added, intelligent units are expanded, intelligent units are reduced, or intelligent units are replaced according to the change in the cluster deployment structure.

8. The deployment method according to claim 3, characterized in that: After deploying the agent unit of the agent through the second computing resource pool, the method further includes: The data of the R&D unit is synchronized to the intelligent unit through the data synchronization unit.

9. The deployment method according to claim 1, wherein: The general computing resources include: CPU resources, memory resources, and hard disk resources.

10. A computer program product, characterized in that The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to perform the steps of the method for deploying an intelligent agent as described in any one of claims 1 to 9.

Citation Information

Cited By

  • Deployment method, device and equipment of large model agent and medium

    CN121116649A