CPU (Central Processing Unit) resource allocation method for DPU (Distributed Processing Unit) centralized service grid
By obtaining the indicator information of the monitoring platform and preset scaling rules in the DPU centralized service grid, dynamically adjusting CPU resources, the problems of resource waste and performance degradation under static allocation methods are solved, and efficient resource utilization and operation and maintenance management are achieved.
Patent Information
- Application Number
- CN202510389662.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the CPU resource allocation method of the DPU centralized service mesh is static allocation, which cannot adapt to changes in dynamic business load, resulting in waste of resources, degraded service performance and increased operation and maintenance complexity.
By obtaining the indicator information of the monitoring platform, the CPU resources of the DPU centralized service mesh are dynamically adjusted using preset expansion and scaling rules, including receiving rule parameters input by users, generating scaling rules, and storing and monitoring their changes in the database, executing scaling instructions based on the indicator information, and adjusting the CPU resources of the service mesh data surface agent.
It realizes dynamic adjustment of CPU resources according to actual load needs, avoid resource waste, ensure service performance, improve resource utilization and operation and maintenance efficiency, and reduce manual intervention.
Smart Images

Figure CN120295789A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of centralized service meshes, and in particular to a method for allocating CPU resources for a DPU centralized service mesh, as well as a device and an electronic device for allocating CPU resources for a DPU centralized service mesh, a computer storage medium and a computer program product for implementing the method for allocating CPU resources for a DPU centralized service mesh. Background Art
[0002] In a Data Processing Unit (DPU) centralized service mesh, the service mesh data plane (Envoy) is deployed on the System on a Chip (SOC) side of the DPU, and containers are isolated using an application container engine (Docker), separated from the host-side container orchestration platform (Kubernetes) environment. The DPU does not occupy host resources, increasing the number of microservices deployed on the host side; in addition, the DPU can directly process traffic without sending it to the host, significantly improving network latency and forwarding efficiency.
[0003] Traditional DPU centralized service mesh solutions usually deploy the Envoy service using a static resource allocation method. That is, a fixed amount of CPU resources is pre-allocated to the Envoy proxy during the initialization phase, for example, 2 CPU cores are allocated. Thereafter, regardless of how the actual business traffic and load change, the CPU resources of the Envoy proxy remain unchanged.
[0004] This passive resource adjustment method has the following problems: 1. Resource waste: When the business traffic is low, the pre-allocated CPU resources cannot be fully utilized, resulting in resource waste. 2. Service performance degradation: When the peak business traffic arrives, if the pre-allocated CPU resources are insufficient, the Envoy proxy may experience performance bottlenecks, leading to problems such as increased service latency and request failures, affecting the overall service quality. 3. Lack of flexibility: The static resource allocation method cannot adapt to dynamically changing business loads, and manual intervention is required to adjust the CPU resources, increasing the operation and maintenance costs and complexity. Summary of the Invention
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, this application provides a method, device and equipment for allocating CPU resources for a DPU centralized service mesh, which can dynamically allocate CPU resources according to changes in business loads, avoid resource waste, improve service performance, and eliminate the need for manual intervention.
[0006] To achieve the above object, the technical solutions provided in the embodiments of this application are as follows:
[0007] In a first aspect, the present application provides a method for allocating CPU resources in a DPU centralized service mesh, the method including:
[0008] Obtain the metric information of the service mesh data plane proxy collected by the monitoring platform, where the metric information includes CPU utilization rate;
[0009] Read a preset scaling rule from the database, and determine whether to scale the CPU resources according to the metric information; wherein, the preset scaling rule is pre-configured and stored in the database;
[0010] If so, execute the scaling instruction to adjust the CPU resources of the service mesh data plane proxy.
[0011] As an optional implementation manner of the embodiment of the present application, before obtaining the metric information of the service mesh data plane proxy collected by the monitoring platform, the method further includes: receiving rule parameters input by the user according to business requirements, where the rule parameters include: target object of action, trigger condition, and polling interval, and the trigger condition includes metric type and trigger threshold; generating a scaling rule according to the rule parameters; storing the scaling rule in the database.
[0012] As an optional implementation manner of the embodiment of the present application, after storing the scaling rule in the database, the method further includes: monitoring whether the scaling rule in the database changes through the monitoring platform; if so, updating the scaling policy corresponding to the scaling rule.
[0013] As an optional implementation manner of the embodiment of the present application, if so, execute the scaling instruction to adjust the CPU resources of the service mesh data plane proxy, including: in the case of determining to scale the CPU resources, determining the scaling policy corresponding to the scaling rule; generating a scaling instruction according to the scaling policy; executing the scaling instruction through the scaling component to adjust the CPU resources of the service mesh data plane proxy.
[0014] As an optional implementation manner of the embodiment of the present application, after if so, execute the scaling instruction to adjust the CPU resources of the service mesh data plane proxy, the method further includes: synchronizing the CPU resource adjustment event to the communication tool to notify the user of the CPU resource status.
[0015] In a second aspect, the present application provides a device for allocating CPU resources in a DPU centralized service mesh, the device including:
[0016] A metric adapter, configured to obtain the metric information of the service mesh data plane proxy collected by the monitoring platform, where the metric information includes CPU utilization rate;
[0017] A controller for reading preset scaling rules from a database and determining whether to scale the CPU resources based on metric information; wherein the preset scaling rules are pre-configured and stored in the database;
[0018] An expander for, if so, executing a scaling instruction to adjust the CPU resources of the service mesh data plane proxy.
[0019] As an optional implementation manner of an embodiment of the present application, the device further includes a rule configurator for: receiving rule parameters input by a user according to service requirements, where the rule parameters include: a target object of action, a trigger condition, and a polling interval, and the trigger condition includes a metric type and a trigger threshold; generating a scaling rule according to the rule parameters; storing the scaling rule in the database.
[0020] As an optional implementation manner of an embodiment of the present application, the rule configurator is further used for: monitoring whether the scaling rule in the database changes through a monitoring platform; if so, updating the scaling policy corresponding to the scaling rule.
[0021] As an optional implementation manner of an embodiment of the present application, the expander is specifically used for: in the case of determining to scale the CPU resources, determining the scaling policy corresponding to the scaling rule; generating a scaling instruction according to the scaling policy; executing the scaling instruction through a scaling component to adjust the CPU resources of the service mesh data plane proxy.
[0022] As an optional implementation manner of an embodiment of the present application, the device further includes a communicator for: synchronizing the CPU resource adjustment event to a communication tool to notify the user of the CPU resource status.
[0023] In a third aspect, the present application provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, it implements the DPU centralized service mesh CPU resource allocation method as described in the first aspect or any one of its optional implementation manners.
[0024] In a fourth aspect, the present application provides a computer-readable storage medium, including: a computer program stored on the computer-readable storage medium, where when the computer program is executed by a processor, it implements the DPU centralized service mesh CPU resource allocation method as described in the first aspect or any one of its optional implementation manners.
[0025] In a fifth aspect, the present application provides a computer program product, including: the computer program product includes a computer program, and when the computer program runs on a computer, it causes the computer to implement the DPU centralized service mesh CPU resource allocation method as described in the first aspect or any one of its optional implementation manners.
[0026] The technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:
[0027] The embodiment of the present application provides a CPU resource allocation method, device, equipment, storage medium and program product for a DPU centralized service mesh. The method includes: first, obtaining the metric information of the service mesh data plane proxy collected by the monitoring platform, where the metric information includes CPU utilization, and reading a preset scaling rule from the database, and determining whether to perform CPU resource scaling according to the preset scaling rule and the metric information. If so, execute the scaling instruction to adjust the CPU resources of the service mesh data plane proxy.
[0028] The embodiment of the present application determines whether to scale the CPU resources of the service mesh data plane proxy according to the monitored metric information and the preset scaling rule, can automatically reduce CPU resources when the business traffic is low to avoid waste, and automatically increase CPU resources when the business traffic is at the peak to ensure service performance. Thus, through the dynamic CPU resource scaling mechanism, the service mesh data plane proxy can always utilize sufficient CPU resources to process the process, ensuring low latency and high throughput of the service and improving service quality. It realizes the automatic management of CPU resources without manual intervention, improving resource utilization and operation and maintenance efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present application and used together with the description to explain the principles of the present application.
[0030] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0031] Figure 1 It is a flowchart of a CPU resource allocation method for a DPU centralized service mesh provided by the embodiment of the present application;
[0032] Figure 2 It is a schematic diagram of a system architecture provided by the embodiment of the present application;
[0033] Figure 3 It provides a CPU resource allocation device for a DPU centralized service mesh according to the embodiment of the present application;
[0034] Figure 4 It is a schematic diagram of the structure of an electronic device described in the embodiment of the present application. Detailed implementation manners
[0035] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the technical terms required in the description of the embodiments or the prior art:
[0036] Data Processing Unit (DPU for short), which is a major category of newly developed dedicated processors. It is the third important computing chip in the data center scenario after the CPU and GPU, providing a computing engine for high-bandwidth, low-latency, data-intensive computing scenarios. It is a new generation of computing chip centered on data, I / O-intensive, adopting a software-defined technology route to support the virtualization of the infrastructure resource layer, capable of improving the efficiency of the computing system, reducing the total cost of ownership of the overall system, and enhancing data processing efficiency while reducing the performance loss of other computing chips. The DPU network card is installed on the cloud server node in the data center, providing a high-bandwidth, low-latency heterogeneous network computing acceleration engine for the cloud server node. After correctly installing the DPU network card on the node, the DPU network card resources appear in the kernel network space of the node in the form of VF network interfaces and PF network interfaces.
[0037] The DPU system on chip (SOC) is an operating system deployed on the DPU network card.
[0038] Docker is a container running platform that enables the rapid deployment and operation of applications by packaging software and its dependencies into containers. Docker container technology is based on operating system-level virtualization, packaging applications and all their dependencies into an independent container image to ensure that applications can run consistently in any environment. This technology is applicable not only to Linux applications but also supports Windows applications, thus achieving cross-platform compatibility. By providing a standardized, lightweight, isolated, and highly flexible solution, Docker greatly simplifies the application deployment and management process, improving development efficiency and system resource utilization.
[0039] The container orchestration platform Kubernetes, also known as K8s, is an open-source system for automatically deploying, scaling, and managing containerized applications. It combines the containers that make up an application into logical units for easy management and service discovery. K8s is typically used to build a multi-container network interface (Container Network Interface, CNI) network. Linux containers provide a lightweight virtualization method that allows multiple virtual environments (containers) to run simultaneously on a single host. Containers provide virtualization at the operating system level, where the kernel controls the isolated containers.
[0040] Istio is an open-source service mesh platform designed to connect, protect, control, and observe services. Istio reduces the complexity of deployment and management by providing a unified microservices architecture solution. It achieves functions such as traffic control, security assurance, and metric collection by introducing a transparent proxy layer in the microservices architecture to automatically manage network communication between services.
[0041] As a component of the data plane, all requests are sent to Envoy and then forwarded by Envoy to the backend server. It is a core part of the Istio architecture, responsible for handling functions such as routing of network requests, load balancing, service discovery, and health checks. Envoy supports HTTP / 1.1 and HTTP / 2 protocols and can act as a two-way transparent proxy to bridge clients and servers of these two protocols. It is recommended to use the HTTP / 2 protocol to create a persistent connection grid for request and response multiplexing. In addition, Envoy also supports routing and load balancing of gRPC requests and responses, as well as L7 sniffing, statistics, and logging of MongoDB connections, providing extensive support for key components in modern web applications.
[0042] With the rapid development of cloud-native technologies, the microservices architecture has become a popular choice for building modern applications. However, the microservices architecture also brings new challenges, such as complex communication between services, network management, and observability issues. Service Mesh has emerged to provide unified traffic management, security, and observability functions for microservices.
[0043] Traditional service meshes usually adopt the Sidecar Proxy mode, where each service instance is accompanied by a proxy to handle traffic in and out of the service. Although the Sidecar Proxy mode simplifies communication between services, it introduces additional resource overhead. For example, each service instance requires additional CPU and memory to run the proxy. In addition, in the Sidecar Proxy mode, traffic needs to be processed by the host network protocol stack twice, increasing latency.
[0044] To overcome the disadvantages of the Sidecar Proxy mode, the Centralized Service Mesh has emerged. The Centralized Service Mesh centralizes all proxy functions onto one or more dedicated nodes, thus reducing resource overhead and providing higher performance.
[0045] The emergence of DPU provides new possibilities for building high-performance and low-latency centralized service meshes. A DPU is a specialized processor dedicated to handling network, storage, and security tasks, which can offload these tasks from the CPU and provide higher throughput and lower latency.
[0046] The DPU centralized service mesh deploys the service mesh data plane (Envoy) to the SOC side of the DPU, isolates it using Docker containers, and separates it from the host-side Kubernetes environment. The DPU does not consume host resources, increasing the number of microservices deployed on the host side. In addition, the DPU can directly process traffic without sending it to the host, significantly improving network latency and forwarding efficiency.
[0047] However, deploying the service mesh to the DPU also faces new challenges: 1. Resource limitations: The resources of the DPU (such as CPU and memory) are usually more limited than those of the host server. 2. Dynamic load: The load of microservices usually changes dynamically, and the service mesh needs to be able to automatically adjust resource usage according to the load situation. 3. Efficient management: A mechanism is needed to effectively manage and monitor the service mesh components deployed on the DPU.
[0048] In the prior art, when deploying the Envoy service on the DPU, a static resource allocation method is usually adopted. That is, a fixed amount of CPU resources is pre-allocated to the Envoy container during the initialization phase, for example, 2 CPU cores are allocated. After that, regardless of how the actual business traffic and load change, the CPU resources of Envoy remain unchanged. Specifically, the common practices in the prior art for handling insufficient Envoy CPU resources on the DPU are as follows: ① Monitor service performance: By monitoring metrics such as CPU utilization and request latency, determine whether there are performance problems with the Envoy service. ② Manual intervention: When it is found that the service performance degrades due to insufficient Envoy CPU resources, the operation and maintenance personnel need to intervene manually and perform the following operations: analyze the cause of the problem, determine the amount of CPU resources that need to be increased; stop the existing Envoy container; modify the configuration of the Envoy container to increase CPU resources; redeploy the Envoy container.
[0049] This passive resource adjustment method has the following problems: ① Resource waste: When the business traffic is low, the pre-allocated CPU resources cannot be fully utilized, resulting in resource waste. ② Service performance degradation: When the peak business traffic arrives, if the pre-allocated CPU resources are insufficient, the Envoy container may experience performance bottlenecks, leading to problems such as increased service latency and request failures, affecting the overall service quality. ③ Lack of flexibility: The static resource allocation method cannot adapt to the dynamically changing business load, requires manual intervention to adjust CPU resources, and cannot respond to sudden traffic in a timely manner, increasing the operation and maintenance cost and complexity.
[0050] To solve some or all of the technical problems existing in the related art, an embodiment of the present application provides a method for allocating CPU resources of a DPU centralized service mesh. The method first obtains the metric information of the service mesh data plane proxy collected by the monitoring platform. The metric information includes CPU utilization rate, and reads the preset scaling rules from the database. It determines whether to scale the CPU resources according to the preset scaling rules and the metric information. If so, it executes the scaling instruction to adjust the CPU resources of the service mesh data plane proxy.
[0051] The present application solves the problems in the existing DPU-based centralized service mesh, such as the lack of elasticity in Envoy service resource allocation, inability to adapt to dynamic business loads, resulting in low resource utilization, degraded service performance, and complex operation and maintenance management. It has the following effects: improving resource utilization, dynamically adjusting the CPU resources of the DPU service according to actual load requirements, avoiding resource waste, and reducing costs. Enhancing service availability, ensuring that the DPU service has sufficient CPU resources to handle the load and improving service availability. Automatically managing the number of workloads, reducing manual intervention, and improving development efficiency.
[0052] In order to more clearly understand the above-mentioned objects, features, and advantages of the present application, the solution of the present application will be further described below. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0053] Many specific details are set forth in the following description to facilitate a thorough understanding of the present application, but the present application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present application, rather than all of the embodiments.
[0054] A method for allocating CPU resources of a DPU centralized service mesh provided in an embodiment of the present application can be implemented by a CPU resource allocation device or an electronic device of the DPU centralized service mesh. The electronic device includes, but is not limited to, a server, a personal computer, a laptop computer, a tablet computer, a smart phone, etc. The operating system of the electronic device may include Android, iOS developed by Apple Inc., Windows developed by Microsoft Corporation of the United States, etc., and the embodiments of the present application do not limit this. The electronic device can run alone to implement the present application, or can be connected to the network and implement the present application through interaction with other computer devices in the network. Among them, the network where the electronic device is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN) network, etc.
[0055] It should be noted that the protection scope of the CPU resource allocation method for a DPU centralized service mesh described in the embodiments of this application is not limited to the execution order of the steps listed in this embodiment. Any solutions achieved by adding or reducing steps of the prior art and replacing steps according to the principle of this application are included in the protection scope of this application.
[0056] As Figure 1 shown, Figure 1 FIG. is a schematic flowchart of a CPU resource allocation method for a DPU centralized service mesh provided by an embodiment of this application. This method can be executed by a CPU resource allocation device of the DPU centralized service mesh. This method mainly includes the following steps S101 to S103:
[0057] S101. Obtain the metric information of the service mesh data plane proxy collected by the monitoring platform.
[0058] Among them, the metric information includes CPU utilization rate, etc. The monitoring platform is deployed on the DPU SOC side of the centralized service mesh. The service mesh data plane (Envoy) proxy is deployed in a Docker container on the DPU SOC side.
[0059] As Figure 2 shown, Figure 2 FIG. is a schematic system architecture diagram provided by an embodiment of this application. This system includes a control node (Master NODE), a worker node (WoekerNode1), and a DPU. The worker node is the host side of the DPU and runs a container orchestration platform (Kubernetes); the worker node communicates with the DPU based on the peripheral component interconnect express (PCIE) standard.
[0060] Figure 2 The DPU in is used to accelerate tasks such as network, storage, and security; the DPU SOC includes Docker containers, a monitoring platform, a database, and a scaling component. The service mesh data plane (Envoy) and the CPU resource service run in the Docker container. Envoy is responsible for handling the traffic between services. The CPU resource service is used to read the scaling rules stored in the database, determine whether CPU resource scaling is required according to the metric information of Envoy feedback by the monitoring platform, and execute the scaling instruction if necessary to adjust the CPU resources. The monitoring platform is used to collect and display the metric information of Envoy for operation and maintenance personnel to monitor and analyze. The database is used to store the scaling rules of Envoy.
[0061] When executing step S101, the CPU resource service calls the monitoring platform to obtain the real-time metric information of the Envoy proxy collected by it. The object of the CPU resource service is the Envoy. It realizes decoupling from the monitoring platform, facilitating users to select a suitable monitoring platform according to the actual situation.
[0062] S102. Read the preset scaling rules from the database and determine whether to scale the CPU resources according to the metric information.
[0063] Among them, the preset scaling rules are the scaling rules pre-configured by the operation and maintenance personnel and stored in the database. The preset scaling rules include but are not limited to the target object (scaleTargetRef), trigger conditions (triggers), and polling interval (pollingInterval). The trigger conditions include the metric type (metricType) and the trigger threshold (threshold).
[0064] Exemplarily, the pseudo-code of the preset scaling rules is as follows:
[0065]
[0066] The above preset scaling rules mean that the metric information (CPU utilization rate) is judged as follows every 30 seconds: judge whether the metric information exceeds the trigger threshold (70%). If so, execute step S103; if not, return to step S101 to continue monitoring the metric information for re-judgment.
[0067] S103. If so, execute the scaling instruction to adjust the CPU resources of the service mesh data plane proxy.
[0068] The scaling instruction is used to instruct the DPU to adjust the CPU resources of the Envoy proxy. Exemplarily, the scaling instruction instructs to increase the number of CPU cores of the Envoy proxy, or the scaling instruction instructs to decrease the number of CPU cores of the Envoy proxy.
[0069] In the above steps, the CPU resource service judges whether it is necessary to scale the CPU resources of the Envoy proxy according to the monitored metric information of the Envoy proxy and the read preset scaling rules; if necessary, it issues a scaling instruction to the DPU to adjust the CPU resources of the Envoy proxy. It can realize automatically reducing the CPU resources when the business traffic is low to avoid waste; when the peak period of business traffic arrives, it automatically expands the CPU resources to ensure service performance, realizing elastic scaling of resources and improving resource utilization.
[0070] In some embodiments, before the above steps S101 to S103, the method further includes a scaling rule configuration process, which specifically includes the following steps S201 to S203:
[0071] S201. Receive rule parameters input by the user according to business requirements.
[0072] The rule parameters include the target object of action, trigger conditions, and polling interval. The trigger conditions include the metric type and trigger threshold.
[0073] S202. Generate a scaling rule according to the rule parameters.
[0074] S203. Store the scaling rule in the database.
[0075] Multiple scaling rules are stored in the database, and the scaling rules support modification.
[0076] In the above embodiments, the configuration of the scaling rules for the service mesh data plane proxy is stored in the database and managed by a dedicated CPU resource service, no longer relying on the Kubernetes API, thereby improving the flexibility and scalability of the system. By managing the rules through the database, the configurability and maintainability of the system are improved.
[0077] In some embodiments, after configuring the scaling rules, the following steps S301 to S302 are further included:
[0078] S301. Monitor whether the scaling rules in the database change through the monitoring platform.
[0079] Exemplarily, monitor whether any scaling rule (such as scaling rule 1) in the database changes. The pseudocode of scaling rule 1:
[0080]
[0081] S302. If so, update the scaling policy corresponding to the scaling rule.
[0082] The scaling policy corresponding to scaling rule 1 indicates that when the CPU utilization rate exceeds 50%, increase the number of CPU cores of the Envoy container.
[0083] Continuing with the above example, if scaling rule 1 changes to:
[0084]
[0085] If the above trigger threshold changes, update the corresponding scaling policy. The changed scaling policy indicates that when the CPU utilization rate exceeds 70%, increase the number of CPU cores of the Envoy container.
[0086] Based on the above embodiments, step S103 (if so, send a scaling instruction to the data processing unit to adjust the CPU resources of the sidecar proxy container) specifically includes: if so, determine the scaling policy corresponding to the scaling rule; generate a scaling instruction according to the scaling policy; execute the scaling instruction through the scaling component to adjust the CPU resources of the Envoy proxy.
[0087] The above embodiments achieve the adjustment of CPU resources by executing the scaling instruction, without restarting the Docker container, ensuring the continuity of the service mesh data plane proxy.
[0088] In some embodiments, after step S103, the resource adjustment event is sent to the communication tool to notify the user of the resource status.
[0089] Exemplarily, by configuring Webhook notifications to send resource adjustment events to communication tools such as Slack and Email, the DPU operations and maintenance team can timely understand the auto-scaling status of the application. Among them, Webhook is a mechanism that automatically executes custom scripts or notifies external systems when a specific event occurs. Webhook allows some operations to be automatically executed when a specific event occurs, such as code submission, review, deployment, etc., such as sending notifications, automatic building, and deploying to servers. It triggers an automated workflow by providing a special uniform resource locator (URL) and sending data to the URL when the specified event occurs. Slack is an enterprise chat tool designed to improve team collaboration efficiency. It integrates various tools, utilizes the powerful capabilities of generative AI, automates routine tasks, and simplifies the workflow through preference applications that are always available in Slack.
[0090] In summary, this application preset the scaling rule, dynamically adjusts the CPU resources of Envoy in the DPU SOC side Docker container according to real-time metrics, realizes the elastic scaling of CPU resources, improves resource utilization, and thus ensures the stability and continuity of the service mesh data plane proxy. Through the dynamic CPU resource scaling mechanism, the service mesh data plane proxy can always utilize sufficient CPU resources to process the process, ensuring low latency and high throughput of the service, and improving the service quality. It realizes the automated management of CPU resources without manual intervention. It improves resource utilization and operation and maintenance efficiency, thereby reducing the operation cost of the DPU centralized service mesh.
[0091] In some embodiments, the CPU resource allocation method for the service mesh data plane proxy provided by the embodiments of this application includes a scaling rule configuration stage and a scaling stage:
[0092] In the scaling rule configuration stage, the operation and maintenance personnel configure the CPU resource scaling rules for the Envoy proxy according to the business requirements, and then store the configured scaling rules in the database. And monitor the database in real time to detect whether the currently configured scaling rules have changed. If so, change the corresponding scaling policy for the scaling rules.
[0093] In the scaling stage, call the monitoring platform to obtain the metric information (such as CPU utilization rate) of the Envoy proxy in real time; the CPU resource service determines whether to scale the CPU resources of the Envoy proxy according to the scaling rules read from the database and the metric information. If necessary, the CPU resource service sends a scaling instruction to the DPU to be received and executed by the DPU to adjust the CPU resources of the Envoy proxy; further, send the CPU resource scaling event to the communication tool (such as Slack, Email) to prompt the operation and maintenance team to understand the CPU resource status. If it is not necessary to scale the CPU resources, the CPU resource service does not perform any operation.
[0094] The above embodiment scales the CPU resources of the Envoy proxy deployed in the Docker container on the DPU SOC side, and no longer depends on the Kubernetes API for scaling. Instead, the CPU resource scaling rules of the Envoy proxy are stored in the database and managed by a dedicated CPU resource service, which improves the flexibility and scalability of the system. The CPU resource service obtains the metric information of the Envoy proxy by calling the monitoring platform, realizing the decoupling from the monitoring platform, and facilitating users to select a suitable monitoring platform according to the actual situation. The adjustment of the CPU resources is completed through the scaling instruction, and there is no need to restart the Docker container, ensuring the continuity of the Envoy service. It provides an efficient, flexible and easy-to-manage solution for the CPU resources of the DPU centralized service mesh, without manual intervention, and can realize the reasonable utilization of the CPU resources on the DPU and the stability of Envoy through preset scaling rules.
[0095] As Figure 3 shown, Figure 3 This application embodiment provides a CPU resource allocation device for a DPU centralized service mesh. The device includes:
[0096] A metric adapter 301 for obtaining the metric information of the service mesh data plane proxy collected by the monitoring platform, where the metric information includes CPU utilization rate;
[0097] A controller 302 for reading the preset scaling rules from the database and determining whether to scale the CPU resources according to the metric information; where the preset scaling rules are pre-configured and stored in the database;
[0098] An expander 303, which is used to execute a scaling instruction if so, so as to adjust the CPU resources of the service mesh data plane proxy.
[0099] As an alternative implementation manner of the embodiment of the present application, the device further includes a rule configurator 304, which is used to: receive rule parameters input by the user according to service requirements, where the rule parameters include: a target object of action, a trigger condition, and a polling interval, and the trigger condition includes an index type and a trigger threshold; generate a scaling rule according to the rule parameters; store the scaling rule in a database.
[0100] As an alternative implementation manner of the embodiment of the present application, the rule configurator 304 is further used to: monitor whether the scaling rule in the database changes through a monitoring platform; if so, update the scaling policy corresponding to the scaling rule.
[0101] As an alternative implementation manner of the embodiment of the present application, the expander 303 is specifically used to: determine a scaling policy corresponding to the scaling rule in the case of judging to scale the CPU resources; generate a scaling instruction according to the scaling policy; execute the scaling instruction through a scaling component to adjust the CPU resources of the service mesh data plane proxy.
[0102] As an alternative implementation manner of the embodiment of the present application, the device further includes a communicator 305, which is used to: synchronize the CPU resource adjustment event to a communication tool to notify the user of the CPU resource status.
[0103] For the specific limitation of the CPU resource allocation device of the DPU centralized service mesh, reference can be made to the limitation of the CPU resource allocation method of the DPU centralized service mesh in the above text, which will not be elaborated here. Each component in the above CPU resource allocation device of the DPU centralized service mesh can be implemented in whole or in part by software, hardware and their combination. The above components can be embedded in the processor in the computer device in hardware form or independent of it, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above components.
[0104] In one embodiment, the present application provides an electronic device, which can be a terminal, and its internal structure diagram can be as Figure 4As shown in the figure. The electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for monitoring lags. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0105] Those skilled in the art can understand that Figure 4 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0106] The embodiment of the present application provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements each process of the continuous prefix fine-tuning method of the large language model in the above method embodiment and can achieve the same technical effect.
[0107] Among them, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0108] The embodiment of the present application provides a computer program product. The computer program product stores a computer program. When the computer program is executed by a processor, it implements each process of the structural safety intelligent monitoring method for the multi-stage construction process of a building in the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0109] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media that contain computer-usable program code.
[0110] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0111] In the present application, the processor can be a central processing unit (CPU), or can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor, etc.
[0112] In the present application, the memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0113] In this application, computer-readable media include both permanent and non-permanent, removable and non-removable storage media. The storage media can implement information storage by any method or technology, and the information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory media such as modulated data signals and carrier waves.
[0114] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0115] The above are only specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A CPU resource allocation method for a DPU centralized service grid, characterized in that Including: Obtain the metric information of the service mesh data plane proxy collected by the monitoring platform, where the metric information includes CPU utilization; Read the preset scaling rules from the database, and determine whether to scale the CPU resources according to the metric information; where the preset scaling rules are pre-configured and stored in the database; If so, execute the scaling instruction to adjust the CPU resources of the service mesh data plane proxy.
2. The method according to claim 1, wherein Before obtaining the metric information of the service mesh data plane proxy collected by the monitoring platform, the method further includes: Receive rule parameters input by the user according to business requirements, where the rule parameters include: target object, trigger condition, and polling interval, and the trigger condition includes metric type and trigger threshold; Generate scaling rules according to the rule parameters; Store the scaling rules in the database.
3. The method according to claim 2, wherein, After storing the scaling rules in the database, the method further includes: Monitor whether the scaling rules in the database change through the monitoring platform; If so, update the scaling policy corresponding to the scaling rules.
4. The method according to claim 3, characterized in that, The "if so, execute the scaling instruction to adjust the CPU resources of the service mesh data plane proxy" includes: When it is determined to scale the CPU resources, determine the scaling policy corresponding to the scaling rules; Generate a scaling instruction according to the scaling policy; Execute the scaling instruction through the scaling component to adjust the CPU resources of the service mesh data plane proxy.
5. The method according to claim 1, characterized in that, After "if so, execute the scaling instruction to adjust the CPU resources of the service mesh data plane proxy", the method further includes: Synchronize the CPU resource adjustment event to the communication tool to notify the user of the CPU resource status.
6. A CPU resource allocation device for a DPU centralized service grid, characterized in that, Including: A metric adapter for obtaining the metric information of the service mesh data plane proxy collected by the monitoring platform, where the metric information includes CPU utilization; A controller for reading the preset scaling rules from the database and determining whether to scale the CPU resources according to the metric information; where the preset scaling rules are pre-configured and stored in the database; An expander for, if so, executing the scaling instruction to adjust the CPU resources of the service mesh data plane proxy.
7. The device according to claim 6, characterized in that The device further includes a rule configurator for: Receive rule parameters input by the user according to business requirements, where the rule parameters include: target object, trigger condition, and polling interval, and the trigger condition includes metric type and trigger threshold; Generate scaling rules according to the rule parameters; Store the scaling rules in the database.
8. An electronic device, characterized in that, Including: A processor, a memory, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, it implements the DPU centralized service mesh CPU resource allocation method according to any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that, Including: A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the DPU centralized service mesh CPU resource allocation method according to any one of claims 1 to 5.
10. A computer program product, characterized in that, Including: The computer program product includes a computer program which, when running on a computer, causes the computer to implement the CPU resource allocation method for the DPU centralized service grid according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method, device and equipment for thermally updating CPU cores based on kubernetes cluster containers and readable medium
CN113806075A
Deployment method and device of service grid unit, equipment and storage medium
CN115733746A
Service grid monitoring index acquisition method, service deployment method and device
CN116962238A
Method and device for dynamically adjusting network resources of DPU (Data Processing Unit) application
CN117240923A
Container scheduling method and device of distributed system, storage medium and electronic equipment
CN118733253A