Intelligent agent platform resource management method and equipment based on cloud native architecture, and medium

Through containerization technology and multi-objective optimization algorithm, the resource management of the intelligent platform realizes dynamic scheduling and collaborative decision-making of resources, solves the problems of inflexible resource allocation and low efficiency of collaborative decision-making, and improves the system's task allocation efficiency and stability.

CN120407184APending Publication Date: 2025-08-01SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 17 Cited by

Patent Information

Application Number
CN202510552309.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the resource management of existing intelligent platform, there are inflexible resource allocation, delayed tasks or wasted resources, low efficiency of collaborative decision-making of intelligent agents, lack of dynamic optimization mechanisms, and unable to meet the needs of high concurrency and low latency business.

Method used

The containerization technology is used to encapsulate the application of the agent as an independent container example, and a dynamic scheduling scheme is generated in combination with the multi-objective optimization algorithm. The degree of fit is calculated by matching the agent's capability vector and the task demand vector, and resource utilization is monitored in real time and elastic scaling mechanisms are triggered to achieve collaborative decision-making and resource optimization.

Benefits of technology

It realizes flexible deployment and scheduling of resources, improves task allocation efficiency, reduces delays, enhances the scientificity and system stability of collaborative decision-making, avoids waste of resources, and meets the adaptive capabilities in high concurrency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407184A_ABST
    Figure CN120407184A_ABST
Patent Text Reader

Abstract

The invention discloses an agent platform resource management method and device based on a cloud native architecture and a medium, and the method comprises the steps: packaging an agent application into an independent container instance based on a containerization technology, and deploying the container instance to a target node; acquiring task demand information of the intelligent agent in real time, and generating a dynamic scheduling scheme by combining the resource state data and through a multi-target optimization algorithm so as to allocate the task to a target container instance; according to a matching function of the capability vector of the intelligent agent and the task demand vector, calculating the integrating degree of the intelligent agent and the task so as to generate a collaborative decision-making result and issue the collaborative decision-making result to the target intelligent agent; the resource utilization rate and the task execution state of the intelligent agent are monitored, an elastic telescoping mechanism or task rescheduling is triggered according to feedback data monitored in real time, and a resource allocation strategy is dynamically adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of resource management, and particularly to a resource management method, device, and medium for an agent platform based on a cloud-native architecture. Background Art

[0002] In the context of the deep integration of cloud computing and artificial intelligence technologies, as the core carrier of distributed task execution and collaborative decision-making, the resource management efficiency of the agent platform directly affects the system performance and response capabilities in complex scenarios. Especially in scenarios where cloud-native architectures are widely used, the agent platform needs to handle dynamically changing computing loads and diverse task requirements, and the technical bottlenecks of traditional resource scheduling and collaborative decision-making mechanisms are becoming increasingly prominent.

[0003] Most existing resource management of agent platforms adopts static allocation strategies, and the allocation and recycling of resource instances lack flexible adjustment capabilities. During peak task periods, fixed resource pools are difficult to cope with sudden loads, resulting in task delays or execution failures. During low-load periods, the resource idle rate rises significantly, causing waste of computing resources. In addition, traditional scheduling methods rely on manually preset rules or simple priority sorting and cannot be dynamically optimized according to real-time task requirements and resource status, resulting in low task allocation efficiency and difficulty in meeting high-concurrency and low-latency business requirements.

[0004] At the level of agent collaborative decision-making, existing technologies usually adopt centralized control or simple communication protocols, making it difficult to achieve efficient cooperation among multiple agents. The task allocation and resource matching processes lack a quantitative evaluation mechanism, resulting in insufficient matching between agent capabilities and task requirements. At the same time, when multiple agents compete for the same resource or there are dependency conflicts in tasks, there is a lack of effective conflict resolution strategies, easily leading to inconsistent decision results or resource preemption problems. Although existing solutions have tried to introduce cloud-native technologies to improve resource management efficiency, there are still deficiencies in the deep integration of dynamic scheduling algorithms and agent collaborative mechanisms, and closed-loop optimization of resource allocation and task execution cannot be achieved. Summary of the Invention

[0005] Embodiments of this application provide a resource management method, device, and medium for an agent platform based on a cloud-native architecture to solve the above technical problems.

[0006] On the one hand, embodiments of this application provide a resource management method for an agent platform based on a cloud-native architecture, including: Encapsulating agent applications into independent container instances based on containerization technology and deploying the container instances to target nodes; Real-time collecting task requirement information of agents, combining resource status data, and generating a dynamic scheduling plan through a multi-objective optimization algorithm to allocate tasks to target container instances; Calculate the fitness between the agent and the task according to the matching function of the agent's ability vector and the task requirement vector, so as to generate a collaborative decision result and send it to the target agent; Monitor the resource utilization rate and task execution status of the agent, and trigger the elastic scaling mechanism or task rescheduling according to the feedback data of real-time monitoring, and dynamically adjust the resource allocation strategy.

[0007] In an implementation manner of the present application, combine the resource status data and generate a dynamic scheduling plan through a multi-objective optimization algorithm to allocate tasks to the target container instance, specifically including: Predict the resource requirements of the agent tasks in the future time period through a pre-trained time series analysis model, and filter the container instances in combination with the policy constraint conditions at the time of task release to eliminate the target nodes that do not match the policy constraint conditions; Calculate the comprehensive scores of the remaining candidate nodes through a multi-objective optimization algorithm, and combine the task priority, delay sensitivity and resource matching degree to generate a dynamic scheduling instruction; Push the dynamic scheduling instruction to the target container instance through an asynchronous message bus, and monitor the task startup status and resource occupancy in real time.

[0008] In an implementation manner of the present application, calculate the fitness between the agent and the task according to the matching function of the agent's ability vector and the task requirement vector, so as to generate a collaborative decision result and send it to the target agent, specifically including: Quantify the agent's perception ability, decision success rate and parallel execution ability into a multi-dimensional ability vector, and normalize the task requirements into a requirement vector of resource type, priority and deadline; Calculate the fitness between the ability vector and the task requirement vector through vector dot product, and superimpose the resource cost to generate a comprehensive matching score; Sort the candidate agents according to the comprehensive matching score, and select the agent with the highest comprehensive matching score among multiple candidate agents to execute the task.

[0009] In an implementation manner of the present application, generate a collaborative decision result, specifically including: Establish a secure communication channel between agents, and call a multi-agent cooperation algorithm to select a decision-making strategy according to the task type; When detecting resource conflicts or task dependency conflicts between agents, start a weighted voting mechanism according to the corresponding decision-making strategy to generate a final collaborative decision result.

[0010] In an implementation manner of the present application, trigger the elastic scaling mechanism or task rescheduling according to the feedback data of real-time monitoring, and dynamically adjust the resource allocation strategy, specifically including: When it is detected that the resource utilization rate of the agent exceeds the preset threshold or the task execution status is delayed, trigger the elastic scaling mechanism to expand or reduce resource instances; Based on the preset priority rules, reserve resource instances for high-priority tasks among multiple tasks, and suspend the resource occupancy of low-priority tasks among multiple tasks when the task requirements are greater than the resource supply; Migrate the agent instance with the largest load among multiple tasks to an idle node through container migration technology and update the resource allocation status.

[0011] In one implementation manner of the present application, dynamically adjust the resource allocation policy, specifically including: Input the resource utilization rate and the task execution status into a pre-trained machine learning model to predict the system load change and resource demand trend in the future time period; Generate a resource pre-allocation suggestion according to the prediction result and dynamically adjust the elastic scaling mechanism of the container instance; When it is detected that the deviation between the resource allocation and the actual demand exceeds the preset threshold, trigger the adaptive optimization algorithm to recalculate the resource allocation weight and task scheduling priority.

[0012] In one implementation manner of the present application, monitor the resource utilization rate and task execution status of the agent, specifically including: Monitor the running status of each agent and record the decision result and execution path of the agent; the running status includes the execution progress, time used, and remaining time of the current task; Real-time collect the resource consumption data of each agent and calculate the resource utilization rate of the agent; the resource consumption data includes the consumption data of CPU, memory, GPU, and network bandwidth; Track the actual execution result of the task and determine the task execution status according to the actual execution result; the actual execution result includes the success rate, delay rate, and error rate of the task.

[0013] In one implementation manner of the present application, it further includes: Continuously track the communication link status of the secure communication channel between agents during the task execution process. When communication delay or data packet loss is detected, dynamically switch to a redundant network path; According to the real-time feedback of the task execution status, perform a secondary prediction on the remaining resource requirements of the unfinished tasks and update the dynamic scheduling instruction; If it is detected that the resource allocation of the container instance is continuously lower than the preset utilization threshold within a preset time period, trigger the resource recovery mechanism to release the idle resources to the global resource pool for other tasks to call.

[0014] On the other hand, the embodiments of the present application also provide an agent platform resource management device based on a cloud-native architecture. The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the agent platform resource management method based on the cloud-native architecture as described above.

[0015] On the other hand, the embodiments of the present application also provide a non-volatile computer storage medium storing computer-executable instructions, which when executed, implement the agent platform resource management method based on the cloud-native architecture as described above.

[0016] The embodiments of the present application provide an agent platform resource management method, device and medium based on a cloud-native architecture, including at least the following beneficial effects: By encapsulating the agent application as an independent instance through containerization technology, lightweight deployment and rapid migration of resources are achieved, and resource instances can be dynamically scaled up or down according to load changes, avoiding resource shortage during peak task periods and resource waste during low-load periods; a dynamic scheduling scheme is generated based on a multi-objective optimization algorithm, combined with real-time resource status and task priorities, to accurately match task requirements with available resource instances, optimize the task allocation path, reduce task execution latency and improve system throughput; through the quantization matching mechanism of the agent capability vector and the task requirement vector, accurate evaluation of the fit between tasks and agents is realized, enhancing the scientificity and reliability of collaborative decision-making, and reducing execution failures or efficiency losses caused by mismatched capabilities; the linkage design of real-time monitoring feedback and elastic scaling mechanism can dynamically adjust resource policies according to task execution status, forming a closed-loop optimization process to ensure the adaptive ability and stability of the system in high-concurrency and multi-task scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings: Figure 1 It is a schematic flowchart of the agent platform resource management method based on the cloud-native architecture provided by the embodiments of the present application; Figure 2 It is an internal structure schematic diagram of the agent platform resource management device based on the cloud-native architecture provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.

[0019] The following will, with reference to the drawings, elaborate on the technical solutions provided by each embodiment of this application.

[0020] Figure 1 It is a schematic flowchart of the method for managing resources of an agent platform based on a cloud-native architecture provided by an embodiment of this application.

[0021] The implementation of the analysis method involved in the embodiments of this application can be a terminal device or a server, and this application does not make special restrictions on this. For the convenience of understanding and description, the following embodiments will be described in detail by taking the server as an example.

[0022] It should be noted that this server can be a single device or a system composed of multiple devices, that is, a distributed server, and this application does not make specific limitations on this.

[0023] As Figure 1 shown, the method for managing resources of an agent platform based on a cloud-native architecture provided by the embodiments of this application includes: Step 101: Package the agent application into an independent container instance based on containerization technology and deploy the container instance to a target node.

[0024] In this embodiment, the agent application and its dependent running environment are packaged into an independent container instance through containerization technology. It should be noted that the dependent running environment in the embodiments of this application includes third-party libraries, configuration files, etc. Exemplarily, the Docker tool is used to encapsulate the code logic, dependencies, and runtime environment of the agent into an image to form a portable containerized unit. It should be noted that containerization technology ensures the independence of the agent running environment through isolation, avoiding resource competition or dependency conflicts between different agents. It can be understood that the lightweight feature of the container instance enables it to be quickly deployed to a target node, such as a physical machine or a virtual machine in a Kubernetes cluster.

[0025] Specifically, during the deployment process, the container orchestration engine (such as Kubernetes) matches the resource requirements of the agent according to the labels of the target node, such as GPU model, memory capacity, and network bandwidth. For example, in an autonomous driving scenario, the perception algorithm agent with high computing requirements will be preferentially scheduled to a node equipped with a GPU accelerator.

[0026] In addition, when the target node fails to run due to high load or hardware failure, the container instances can be dynamically migrated to other idle nodes to ensure the continuity of task execution. Exemplarily, in the industrial robot control scenario, if a certain node has a CPU overload due to task accumulation, the system automatically migrates some container instances to the low-load nodes to maintain the balance of resource allocation.

[0027] Step 102: Collect the task requirement information of the agent in real time, combine it with the resource status data, and generate a dynamic scheduling plan through a multi-objective optimization algorithm to allocate tasks to the target container instances.

[0028] In this embodiment, the task requirement information is actively submitted by the agent or automatically parsed and generated by the platform, including task type (such as compute-intensive, real-time control type), priority (such as emergency task, background task), required resource type (CPU / GPU / memory), and deadline. It can be understood that the resource status data is collected in real time through the monitoring component, covering indicators such as CPU utilization rate, remaining memory, and network bandwidth of each node.

[0029] Exemplarily, a time series analysis model (such as LSTM) is used to predict the resource demand trend in the future time period. For example, in the e-commerce big promotion scenario, the system predicts that the computing resource demand for order processing tasks will surge during a specific period, and thus allocates elastic resource instances in advance. Specifically, combined with the policy constraint conditions at the time of task release, such as node label matching and resource quota limitation, the candidate nodes are filtered to eliminate the nodes that do not meet the conditions. Subsequently, a multi-objective optimization algorithm, such as the combination of genetic algorithm and ant colony algorithm, is used to calculate the comprehensive scores of the remaining nodes. The genetic algorithm is used for the evolution of the global resource allocation scheme, while the ant colony algorithm is used for the optimization search for local resource bottlenecks (such as nodes with too high network latency). The finally generated dynamic scheduling instruction is pushed to the target container instance through the Kafka asynchronous message bus, and the task startup status and resource occupancy are monitored in real time to ensure the efficient execution of the scheduling instruction.

[0030] Step 103: Calculate the fitness between the agent and the task according to the matching function of the agent's ability vector and the task requirement vector, so as to generate a collaborative decision result and send it to the target agent.

[0031] In this embodiment, the matching mechanism of the ability vector and the demand vector and the collaborative decision generation process are defined. It should be noted that the agent's ability is quantified by a multi-dimensional vector, such as indicators such as perception ability, decision success rate, and parallel execution ability. The perception ability is, for example, the image recognition accuracy rate, the decision success rate is, for example, the proportion of the number of effective path planning times, and the parallel execution ability is, for example, the number of tasks processed simultaneously. The task requirements are normalized into a demand vector, including resource type, priority, deadline, and task dependencies.

[0032] Specifically, the matching function is defined as the dot product similarity calculation between the capability vector and the demand vector, and the resource cost is superimposed to generate a comprehensive matching score, such as cross-node communication latency or resource rental cost. Exemplarily, in a logistics and warehousing scenario, if a handling robot has high-load handling capabilities and is currently idle, its matching score with a cargo transportation task will be significantly higher than that of other candidate agents. After sorting according to the scores, the system selects the agent with the highest comprehensive matching score to execute the task.

[0033] It can be understood that when multiple agents compete for the same resource, the weighted voting mechanism will be triggered. Different voting weights are assigned according to the historical task success rate of the agents, and finally an agreement is reached by majority vote or weighted results. For example, in a collaborative inspection task of a drone swarm, if two drones simultaneously request the right to inspect the same area, the system assigns weights according to their historical task success rates, and the drone with the higher weight obtains the execution right first. The collaborative decision result is sent to the target agent through a secure communication channel, such as a TLS encrypted link, to ensure the security of instruction transmission.

[0034] Step 104: Monitor the resource utilization rate and task execution status of the agent, and trigger the elastic scaling mechanism or task rescheduling according to the feedback data of real-time monitoring, and dynamically adjust the resource allocation strategy.

[0035] In this embodiment, the specific implementation of resource monitoring and dynamic adjustment is defined. It should be noted that the resource utilization rate data and task execution status are continuously collected through the real-time monitoring module. The resource utilization rate data includes CPU, GPU, memory, and network bandwidth consumption, and the task execution status includes progress, delay rate, and error rate. Exemplarily, in a video stream processing scenario, if the GPU utilization rate of a container instance continuously exceeds the preset threshold, the system determines that it is in an overloaded state.

[0036] Specifically, the elastic scaling mechanism is triggered. Kubernetes automatically scales the number of container instances to share the load. For example, new GPU instances are added to process the backlogged video rendering tasks. At the same time, based on the priority rules, high-priority tasks (such as real-time payment verification) will preferentially occupy the reserved resource instances, while low-priority tasks (such as log analysis) may be suspended to release resources.

[0037] In addition, the resource recycling mechanism takes effect when long-term idle resources are detected. For example, if the CPU utilization rate of a container instance continuously falls below the preset threshold, the system automatically reclaims its resources and releases them to the global resource pool for other tasks to call.

[0038] It is understandable that the task rescheduling process combines the prediction results of the machine learning model, predicts the future system load changes based on historical data, and dynamically adjusts the resource allocation weights. For example, in an intelligent factory, if it is predicted that the production line is about to enter the maintenance period, the system reduces the number of resource instances in advance to save costs. Finally, the resource allocation strategy is continuously optimized through a closed-loop feedback mechanism to ensure the stability and efficiency of the platform in high-concurrency and multi-task scenarios.

[0039] The above is the method embodiment proposed in this application. Based on the same inventive concept, the embodiment of this application also provides an intelligent agent platform resource management device based on the cloud-native architecture, and its structure is as Figure 2 shown.

[0040] Figure 2 It is the internal structure schematic diagram of the intelligent agent platform resource management device provided by the embodiment of this application. As Figure 2 shown, the device includes: At least one processor; And a memory communicatively connected to at least one processor; Wherein, the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor to enable at least one processor to: Package the intelligent agent application as an independent container instance based on containerization technology and deploy the container instance to the target node; Real-time collect the task requirement information of the intelligent agent, combine the resource status data and generate a dynamic scheduling plan through a multi-objective optimization algorithm to allocate tasks to the target container instance; Calculate the fitness between the intelligent agent and the task according to the matching function of the ability vector and the task requirement vector of the intelligent agent to generate a collaborative decision result and send it to the target intelligent agent; Monitor the resource utilization rate and task execution status of the intelligent agent, and trigger an elastic scaling mechanism or task rescheduling according to the real-time monitored feedback data to dynamically adjust the resource allocation strategy.

[0041] The embodiment of this application also provides a non-volatile computer storage medium, storing computer-executable instructions, and when the computer-executable instructions are executed, they can: Package the intelligent agent application as an independent container instance based on containerization technology and deploy the container instance to the target node; Real-time collect the task requirement information of the intelligent agent, combine the resource status data and generate a dynamic scheduling plan through a multi-objective optimization algorithm to allocate tasks to the target container instance; Calculate the fitness between the intelligent agent and the task according to the matching function of the ability vector and the task requirement vector of the intelligent agent to generate a collaborative decision result and send it to the target intelligent agent; Monitor the resource utilization rate and task execution status of the intelligent agent, and trigger the elastic scaling mechanism or task rescheduling according to the feedback data of real-time monitoring, and dynamically adjust the resource allocation strategy.

[0042] Each embodiment in this application is described in a progressive manner. For the same or similar parts between each embodiment, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0043] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0044] The devices and media provided in the embodiments of the present application correspond one-to-one with the methods. Therefore, the devices and media also have beneficial technical effects similar to those of the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.

[0045] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0046] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0047] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0048] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0049] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0050] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0051] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technologies, compact disc read only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0052] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.

[0053] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for managing resources of an agent platform based on a cloud-native architecture, characterized in that The method includes: Encapsulating the agent application as an independent container instance based on containerization technology and deploying the container instance to a target node; Collecting the task requirement information of the agent in real time, combining with the resource status data and through a multi-objective optimization algorithm, generating a dynamic scheduling scheme to allocate tasks to the target container instance; Calculating the fitness between the agent and the task according to the matching function of the agent's ability vector and the task requirement vector to generate a collaborative decision result and sending it to the target agent; Monitoring the resource utilization rate and task execution status of the agent, and triggering an elastic scaling mechanism or task rescheduling according to the real-time monitored feedback data to dynamically adjust the resource allocation strategy.

2. The method for managing agent platform resources based on a cloud-native architecture according to claim 1, wherein Combining with the resource status data and through a multi-objective optimization algorithm, generating a dynamic scheduling scheme to allocate tasks to the target container instance, specifically including: Predicting the resource requirements of the agent tasks in the future time period through a pre-trained time series analysis model, and filtering the container instances in combination with the policy constraint conditions at the time of task release to eliminate the target nodes that do not match the policy constraint conditions; Calculating the comprehensive scores of the remaining candidate nodes through a multi-objective optimization algorithm and combining task priority, latency sensitivity and resource matching degree to generate a dynamic scheduling instruction; Pushing the dynamic scheduling instruction to the target container instance through an asynchronous message bus and monitoring the task start status and resource occupancy in real time.

3. The method for managing resources of an agent platform based on a cloud-native architecture according to claim 1, wherein, Calculating the fitness between the agent and the task according to the matching function of the agent's ability vector and the task requirement vector to generate a collaborative decision result and sending it to the target agent, specifically including: Quantifying the agent's perception ability, decision success rate and parallel execution ability into a multi-dimensional ability vector, and normalizing the task requirements into a requirement vector of resource type, priority and deadline; Calculating the fitness between the ability vector and the task requirement vector through vector dot product and superimposing the resource cost to generate a comprehensive matching score; Sorting the candidate agents according to the comprehensive matching score and selecting the agent with the highest comprehensive matching score among multiple candidate agents to execute the task.

4. The method for managing agent platform resources based on a cloud-native architecture according to claim 3, wherein, Generating a collaborative decision result, specifically including: Establishing a secure communication channel between agents, calling a multi-agent cooperation algorithm, and selecting a decision-making strategy according to the task type; When detecting resource conflicts or task dependency conflicts between agents, starting a weighted voting mechanism according to the corresponding decision-making strategy to generate a final collaborative decision result.

5. The method for managing agent platform resources based on a cloud-native architecture according to claim 1, wherein Triggering an elastic scaling mechanism or task rescheduling according to the real-time monitored feedback data to dynamically adjust the resource allocation strategy, specifically including: When it is monitored that the resource utilization rate of the agent exceeds the preset threshold or the task execution status is delayed, triggering an elastic scaling mechanism to expand or reduce resource instances; Based on the preset priority rules, reserving resource instances for high-priority tasks among multiple tasks, and suspending the resource occupancy of low-priority tasks among multiple tasks when the task requirements are greater than the resource supply; Migrating the agent instance with the largest load among multiple tasks to an idle node through container migration technology and updating the resource allocation status.

6. The method for managing agent platform resources based on a cloud-native architecture according to claim 5, wherein Dynamically adjusting the resource allocation strategy, specifically including: Input the resource utilization rate and the task execution status into a pre-trained machine learning model to predict the system load change and resource demand trend in a future time period; Generate resource pre-allocation suggestions based on the prediction results and dynamically adjust the elastic scaling mechanism of container instances; When it is detected that the deviation between the resource allocation and the actual demand exceeds a preset threshold, trigger an adaptive optimization algorithm to recalculate the resource allocation weights and task scheduling priorities.

7. The method for managing agent platform resources based on a cloud-native architecture according to claim 1, wherein Monitor the resource utilization rate and task execution status of the intelligent agent, specifically including: Monitor the running status of each intelligent agent and record the decision results and execution paths of the intelligent agent; the running status includes the execution progress, the time used, and the remaining time of the current task; Real-time collect the resource consumption data of each intelligent agent and calculate the resource utilization rate of the intelligent agent; the resource consumption data includes the consumption data of CPU, memory, GPU, and network bandwidth; Track the actual execution results of the task and determine the task execution status according to the actual execution results; the actual execution results include the success rate, delay rate, and error rate of the task.

8. The method for managing resources of an agent platform based on a cloud-native architecture according to claim 1, characterized in that, The method further includes: Continuously track the communication link status of the secure communication channel between intelligent agents during the task execution process. When communication delay or data packet loss is detected, dynamically switch to a redundant network path; According to the real-time feedback of the task execution status, perform a secondary prediction on the remaining resource requirements of the unfinished tasks and update the dynamic scheduling instructions; If it is monitored that the resource allocation of the container instance is continuously lower than the preset utilization rate threshold within a preset time period, trigger a resource recovery mechanism to release the idle resources to the global resource pool for other tasks to call.

9. An intelligent agent platform resource management device based on a cloud-native architecture, characterized in that, The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the resource management method of the intelligent agent platform based on the cloud native architecture according to any one of claims 1-8.

10. A non-volatile computer storage medium stores computer-executable instructions, characterized in that, When the computer-executable instructions are executed, the resource management method of the intelligent agent platform based on the cloud native architecture according to any one of claims 1-8 is implemented.

Citation Information

Cited By

  • Container security detection method, device and system and storage medium

    CN116707890A

  • Model management resource pool dynamic allocation method and device, server and medium

    CN120631602A

  • Dynamic containerization deployment device and method based on runtime feature perception

    CN120743441A

  • GPU resource dynamic allocation method and system based on load awareness

    CN120832243A

  • Data processing method and device, equipment and medium

    CN120872553A