CNF network element adaptive large page memory scheduling method and device

By adding large-page memory adaptive annotations to the POD template and introducing large-page adaptation mechanisms in Kube Scheduler and Kubelet, the problem of resource scheduling is solved in the traditional CNF network element deployment method, and flexible resource scheduling and high system reliability are achieved.

CN120029776APending Publication Date: 2025-05-23CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510139638.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the Xinchuang cloud environment, the traditional CNF network element multi-copy application deployment method has different support for large page memory by different CPU architectures, resulting in resource scheduling being limited, unable to effectively utilize resources, affecting the reliability of the system.

Method used

By adding large-page memory adaptive annotations to the POD template, the Kube Scheduler scheduler performs large-page adaptation of nodes of different architectures during the scheduling stage to obtain alternative nodes. Kubelet adapts the requested large-page resources based on the large-page memory adaptive annotation information and creates a container.

Benefits of technology

It realizes flexible resource scheduling and reuse of CNF network elements in a multi-architecture CPU environment, enhances the reliability and fault tolerance of the system, improves resource utilization, and avoids waste caused by resource mismatch.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029776A_ABST
    Figure CN120029776A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large-page memories, and provides a CNF network element adaptive large-page memory scheduling method and device, and the method comprises the steps: creating a network element multi-copy instance, and adding a large-page memory adaptive annotation in a POD template; the method comprises the following steps of: in a scheduling stage of a Kube Scheduler, carrying out large page adaptation on nodes of different architectures to obtain alternative nodes; in the stage that the Kubelet where the node is located runs the container on the node, the requested large-page resource information is adapted to the large-page size supported by the node according to the large-page memory self-adaptive annotation information, and the container is created; and the Kubelet updates the POD state according to the container creation result to complete the creation of the POD. According to the invention, flexible scheduling and multiplexing of heterogeneous resources of the CNF network element can be realized, efficient utilization of resources and high reliability of a system in a one-cloud multi-core environment are realized, a stable operation environment is provided for the CNF network element, the efficient forwarding performance of the CNF network element is guaranteed, the user experience is greatly improved, and the learning and use cost of a user is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of large page memory technology, and in particular to a CNF network element adaptive large page memory scheduling method and device. Background Art

[0002] With the rapid development of information technology application innovation (ICT), ICT cloud environment has become an important development direction in the current cloud computing field. In ICT cloud environment, in order to achieve efficient utilization and flexible deployment of resources, one cloud multi-core technology is usually adopted, that is, deploying hosts with multiple CPU architectures (such as X86, ARM, etc.) in the same cloud environment. This diversity of architectures brings higher flexibility and compatibility to the cloud environment, but also brings new technical challenges.

[0003] For high-performance network elements (such as container-based NFV elements, CNF), the Data Plane Development Kit (DPDK) is usually used to improve network forwarding performance. DPDK significantly improves data forwarding performance by using large page memory to reduce page fault exceptions. However, there are differences in the support for large page memory between CPUs of different architectures. For example, X86-based CPUs usually support 1G large page memory, while ARM-based CPUs generally support 512M large page memory. This difference leads to scheduling limitations when using traditional deployment methods (such as Kubernetes Deployment) when building multi-copy applications of CNF.

[0004] In traditional Deployment deployment, Template can specify a type of hugepage memory size. However, this method can only be scheduled to CPU nodes that support this type of hugepage memory. For example, when the hugepage memory is specified as hugepage-1Gi, only X86 nodes can meet the requirement, while ARM nodes cannot be scheduled even if there is enough free memory because their hugepage memory is hugepages-512Mi. This scheduling restriction leads to a waste of resources, especially in failure scenarios, even if there are free resources, effective scheduling cannot be performed, so that the resource utilization cannot be fully utilized and high reliability cannot be provided.

[0005] Therefore, in the existing trusted cloud environment, the CNF multi-copy application deployment method based on large page memory has obvious limitations. This limitation not only limits the full utilization of resources, but also affects the reliability of the system. In order to solve this problem, a new technical solution is needed to support different types of large page memory in the trusted cloud environment where multi-architecture CPUs coexist, thereby improving resource utilization and system reliability. Summary of the invention

[0006] In view of this, in order to overcome the deficiencies of the prior art, the present application aims to provide a CNF network element adaptive large page memory scheduling method and device.

[0007] According to a first aspect of the present application, a CNF network element adaptive large page memory scheduling method is provided, the method comprising: Create multiple replica instances of network elements and add large page memory adaptation annotations in the POD template; During the scheduling phase, the Kube Scheduler obtains candidate nodes by adapting large pages to nodes of different architectures. During the container running phase on the node, the Kubelet of the node adapts the requested large page resource information to the large page size supported by the node according to the large page memory adaptive annotation information, and creates a container. Kubelet updates the POD status based on the container creation result and completes the POD creation.

[0008] Optionally, in the CNF network element adaptive large page memory scheduling method of the present application, when creating multiple copy instances of a network element, deployment is performed through a Deployment controller, and a large page memory adaptive annotation is added to the POD template of the Deployment controller.

[0009] Optionally, in the CNF network element adaptive large page memory scheduling method of the present application, when the Kube ControllerManager component monitors the creation of the Deployment controller, it triggers the POD creation operation in sequence according to the number of replica instances, and the declarative definition of the desired state of the POD is copied from the declarative definition of the desired state of the Deployment controller.

[0010] Optionally, in the CNF network element adaptive large page memory scheduling method of the present application, the creation of unbound PODs is monitored by the Kube Scheduler scheduler. When the creation of unbound PODs is monitored, the POD binding node selection is entered.

[0011] Optionally, in the CNF network element adaptive large page memory scheduling method of the present application, the Kube Scheduler obtains candidate nodes by performing large page adaptation on nodes of different architectures during the scheduling stage, including: during the scheduling stage, the Kube Scheduler performs large page adaptation on nodes of different architectures in the filter according to the CPU architecture of the node and the available large page memory resources, and uses heterogeneous nodes that meet the conditions as candidate nodes.

[0012] Optionally, in the CNF network element adaptive large page memory scheduling method of the present application, the Kube Scheduler obtains candidate nodes by performing large page adaptation on nodes of different architectures during the scheduling phase, including: When the node is an X86 node, determine whether the huge page memory reported by the node is greater than the requested huge page memory amount. If the huge page memory reported by the node is greater than the requested huge page memory amount, use the node as a candidate node. When the node is an ARM node, determine whether the available large page memory resources reported by the node are greater than the requested large page memory quantity. If the available large page memory resources reported by the node are greater than the requested large page memory quantity, the node is used as a candidate node.

[0013] Optionally, in the CNF network element adaptive large page memory scheduling method of the present application, Kubelet runs the container stage on the node, reads the system default large page information, matches the requested large page information with the system default large page information, and creates a container based on the matching result.

[0014] Optionally, in the CNF network element adaptive large page memory scheduling method of the present application, when the requested large page information matches the system default large page information, a container is created through the container runtime interface.

[0015] Optionally, in the CNF network element adaptive large page memory scheduling method of the present application, when the requested large page information does not match the large page information missing from the system, the requested large page resource information is converted into a large page size supported by the node, and then a container is created through the container runtime interface.

[0016] According to a second aspect of the present application, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in the first aspect of the present application is implemented.

[0017] The CNF network element adaptive large page memory scheduling method and device of the present application have the following beneficial technical effects: 1. By adding adaptive large page annotations to Pods and adapting large pages to systems with different architectures during the scheduling phase, heterogeneous nodes can be used as candidate nodes, thus breaking the resource limitations caused by differences in node architecture and large pages in traditional scheduling, and realizing flexible scheduling and reuse of heterogeneous resources of CNF network elements.

[0018] 2. The improved container scheduling mechanism introduces a large page memory adaptation mechanism in the Pod scheduling stage and the node container running stage, so that CNF network elements are no longer restricted by the CPU architecture and large page differences of the nodes during resource scheduling. Since the differences in node architectures are shielded, CNF network elements can be flexibly migrated after a failure without being restricted by the architecture, thereby enhancing the reliability and fault tolerance of the system, truly realizing the efficient utilization of resources and high reliability of the system in a one-cloud, multi-core environment, and avoiding waste caused by resource mismatch.

[0019] 3. After selecting the node, in the stage of running the container on the node, the container can be adapted and started according to the large page size supported by the node to ensure the normal operation of the container, thereby providing a stable operating environment for the CNF network element and ensuring its efficient forwarding performance.

[0020] 4. While achieving the above functions, users are completely unaware of it. Users do not need to worry about the complexity of the underlying node architecture and large page configuration, and can use large page resources seamlessly, which greatly improves the user experience and reduces the user's learning and usage costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1 A flowchart of a method for adaptive large page memory scheduling of CNF network elements according to an embodiment of the present application; Figure 2 This is an example diagram of an execution timing of the CNF network element adaptive large page memory scheduling method according to an embodiment of the present application; Figure 3 A schematic diagram of the structure of the device provided in this application. DETAILED DESCRIPTION

[0023] The embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0024] It should be noted that the following embodiments and features in the embodiments may be combined with each other in the absence of conflict; and, based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making any creative work are within the scope of protection of the present disclosure.

[0025] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.

[0026] Figure 1 FIG. 1 is a flowchart of a method for adaptively scheduling large page memory for a CNF network element according to an embodiment of the present application. Figure 1 As shown, in this embodiment, the CNF network element adaptive large page memory scheduling method includes the following steps: Step S101: Create multiple replica instances of network elements and add large page memory adaptive annotations in the POD template.

[0027] As an optional example, when creating a multiple-copy instance of a network element, deployment is performed through a Deployment controller, and a large page memory adaptation annotation is added to a POD template of the Deployment controller.

[0028] When the Kube Controller Manager component listens to the creation of the Deployment controller, it triggers the creation of the POD in sequence according to the number of replica instances, and the declarative definition of the desired state of the POD is copied from the declarative definition of the desired state of the Deployment controller.

[0029] Step S102: During the scheduling phase, the Kube Scheduler obtains candidate nodes by performing large page adaptation on nodes of different architectures.

[0030] In this embodiment, the Kube Scheduler monitors the creation of unbound PODs. When the creation of unbound PODs is monitored, the POD binding node selection is performed.

[0031] As an optional example, during the scheduling phase, the Kube Scheduler performs large page adaptation on nodes of different architectures in the filter according to the CPU architecture of the node and the available large page memory resources, and selects heterogeneous nodes that meet the conditions as candidate nodes.

[0032] When the node is an X86 node, determine whether the large page memory reported by the node is greater than the requested large page memory quantity. If so, the node is selected as a candidate node. When the node is an ARM node, determine whether the large page memory available resources reported by the node are greater than the requested large page memory quantity. If so, the node is selected as a candidate node.

[0033] Step S103: During the container running phase on the node, the Kubelet of the node adapts the requested large page resource information to the large page size supported by the node according to the large page memory adaptive annotation information, and creates a container.

[0034] As an optional example, in this embodiment, Kubelet runs the container phase on the node, reads the system default large page information, matches the requested large page information with the system default large page information, and creates a container based on the matching result. When the requested large page information matches the system default large page information, a container is created through the container runtime interface. When the requested large page information does not match the system missing large page information, the requested large page resource information is converted to the large page size supported by the node, and then the container is created through the container runtime interface.

[0035] Step S104: Kubelet updates the POD status according to the container creation result and completes the creation of the POD.

[0036] The following further describes in detail the CNF network element adaptive large page memory scheduling method of this embodiment in a specific scenario. Figure 2 FIG. 1 is an example diagram of an execution timing of the CNF network element adaptive large page memory scheduling method according to an embodiment of the present application, such as Figure 2 As shown in the figure, in this scenario, the required large page memory amount is 2Gi, and the user's request to create a Deployment is as follows:

[0037] Because adaptive adaptation of huge pages is required, the annotation hugepage.kuberetes.io: "adptive" is added to the POD template at this stage.

[0038] Kube Controller Manager listens to the creation of Deployment and triggers the creation of PODs in sequence according to the number of replicas. The Spec of POD is copied from the Spec of Deloyment.

[0039] Kube Scheduler monitors the creation of an unbound pod and enters the POD scheduling process. The huge page resource is used as the user's resource request. The default logic in the prior art is to judge according to the huge page type and resource request reported by the node. For this requirement, if it is an ARM node, the huge page memory resource reported by the node is hugepage-512Mi. Since the huge page in the POD resource request is hugepages-1Gi, it will be returned that there is no such resource, which will cause the ARM node to not be used as an alternative scheduling node. In this application, the node filtering node is introduced Figure 2 Processing (b) in the above process is to determine whether the pod needs to perform adaptive adaptation of huge pages. If the node is an X86 node, the reported hugepages-1Gi is sufficient. If the node is an ARM node, because the reported hugepage memory is hugepage-512Mi, there is no need to determine the hugepage type. Instead, it is necessary to determine whether the available resources of hugepage-512Mi are greater than 2Gi. If they are greater than 2Gi, the hugepage memory resource check passes and the node is added to the candidate node. Subsequent scoring operations are performed. Therefore, only the screening strategy is changed, and the request resource information of the POD is not changed. At this time, from the perspective of the scheduler, it is considered that the node uses hugepages-1Gi of hugepage memory. The scheduler notifies the API server of the node bound to the last selected pod.

[0040] The Kubelet of the node listens to the pod's binding node and notifies the runtime to create a container based on the pod's request information. Figure 2 The processing of (b) in the above figure may be scheduled to nodes with different large page sizes, so the Figure 2 In (c), according to the hugepage.kuberetes.io: "adptive" annotation information, the requested hugepage information and the default hugepage information in the system (cat / proc / meminfo | grep Hugepagesize) are used for judgment. If the requested hugepage matches, the creation is initiated directly. If it does not match, the requested hugepage resource information is converted from hugepages-1Gi:2Gi in the request to hugepages-512Mi:2Gi of the ARM system, and the runtime is notified to create the container. After adaptation, the hugepage memory can match the system support for hugepages, ensuring the successful creation of the container. The container operation notifies the Kubelet of the creation result, and the Kubelet updates the POD status to complete the creation of the POD. Multi-copy scheduling is successful, and the Deployment is successfully created.

[0041] The CNF network element adaptive large page memory scheduling method of the embodiment of the present application has the following beneficial technical effects: 1. By adding adaptive large page annotations to Pods and adapting large pages to systems with different architectures during the scheduling phase, heterogeneous nodes can be used as candidate nodes, thus breaking the resource limitations caused by differences in node architecture and large pages in traditional scheduling, and realizing flexible scheduling and reuse of heterogeneous resources of CNF network elements.

[0042] 2. The improved container scheduling mechanism introduces a large page memory adaptation mechanism in the Pod scheduling stage and the node container running stage, so that CNF network elements are no longer restricted by the CPU architecture and large page differences of the nodes during resource scheduling. Since the differences in node architectures are shielded, CNF network elements can be flexibly migrated after a failure without being restricted by the architecture, thereby enhancing the reliability and fault tolerance of the system, truly realizing the efficient utilization of resources and high reliability of the system in a one-cloud, multi-core environment, and avoiding waste caused by resource mismatch.

[0043] 3. After selecting the node, in the stage of running the container on the node, the container can be adapted and started according to the large page size supported by the node to ensure the normal operation of the container, thereby providing a stable operating environment for the CNF network element and ensuring its efficient forwarding performance.

[0044] 4. While achieving the above functions, users are completely unaware of it. Users do not need to worry about the complexity of the underlying node architecture and large page configuration, and can use large page resources seamlessly, which greatly improves the user experience and reduces the user's learning and usage costs.

[0045] like Figure 3 As shown, the present application also provides a device, including a processor 210, a communication interface 220, a memory 230 for storing a processor executable computer program, and a communication bus 240. The processor 210, the communication interface 220, and the memory 230 communicate with each other through the communication bus 240. The processor 210 implements the above-mentioned CNF network element adaptive large page memory scheduling method by running the executable computer program.

[0046] Among them, the computer program in the memory 230 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0047] The system embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected based on actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art may understand and implement it without creative effort.

[0048] Through the description of the above implementation modes, those skilled in the art can clearly understand that each implementation mode can be implemented by means of software plus a necessary general hardware platform, or of course by hardware. Based on such an understanding, the above technical solution can essentially or in other words be embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiment.

[0049] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A CNF network element adaptive large page memory scheduling method, characterized in that: The method comprises: Create multiple replica instances of network elements and add large page memory adaptation annotations in the POD template; During the scheduling phase, the Kube Scheduler obtains candidate nodes by adapting large pages to nodes of different architectures. During the container running phase on the node, the Kubelet of the node adapts the requested large page resource information to the large page size supported by the node according to the large page memory adaptive annotation information, and creates a container. Kubelet updates the POD status based on the container creation result and completes the POD creation.

2. The CNF network element adaptive large page memory scheduling method according to claim 1, characterized in that: When creating multiple replica instances of network elements, deploy them through the Deployment controller and add the large page memory adaptation annotation to the POD template of the Deployment controller.

3. The CNF network element adaptive large page memory scheduling method according to claim 1, characterized in that: When the KubeController Manager component listens to the creation of the Deployment controller, it triggers the creation of PODs in sequence according to the number of replica instances, and the declarative definition of the desired state of the POD is copied from the declarative definition of the desired state of the Deployment controller.

4. The CNF network element adaptive large page memory scheduling method according to claim 1, characterized in that: The KubeScheduler scheduler monitors the creation of unbound PODs. When the creation of unbound PODs is monitored, the POD binding node selection is performed.

5. The CNF network element adaptive large page memory scheduling method according to claim 1, characterized in that: During the scheduling phase, the KubeScheduler scheduler obtains candidate nodes by adapting large pages to nodes of different architectures, including: During the scheduling phase, the Kube Scheduler scheduler adapts large pages to nodes of different architectures in the filter according to the CPU architecture of the node and the available large page memory resources, and selects heterogeneous nodes that meet the conditions as candidate nodes.

6. The CNF network element adaptive large page memory scheduling method according to claim 1, characterized in that: During the scheduling phase, the KubeScheduler scheduler obtains candidate nodes by adapting large pages to nodes of different architectures, including: When the node is an X86 node, determine whether the huge page memory reported by the node is greater than the requested huge page memory amount. If the huge page memory reported by the node is greater than the requested huge page memory amount, use the node as a candidate node. When the node is an ARM node, determine whether the available large page memory resources reported by the node are greater than the requested large page memory quantity. If the available large page memory resources reported by the node are greater than the requested large page memory quantity, the node is used as a candidate node.

7. The CNF network element adaptive large page memory scheduling method according to claim 1, characterized in that: Kubelet runs the container phase on the node, reads the system default huge page information, matches the requested huge page information with the system default huge page information, and creates a container based on the matching result.

8. The CNF network element adaptive large page memory scheduling method according to claim 7, characterized in that: When the requested huge page information matches the system default huge page information, a container is created through the container runtime interface.

9. The CNF network element adaptive large page memory scheduling method according to claim 8, characterized in that: When the requested huge page information does not match the huge page information missing from the system, the requested huge page resource information is converted to the huge page size supported by the node, and then the container is created through the container runtime interface.

10. A computer device, characterized in that: The computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the method according to any one of claims 1 to 9 when executing the program.