A CPU Core Binding Method and System for Cloud Native Applications

By reading processor information and device information, generating a processor core object pool, updating and accurately selecting the CPU core in real time, solving the problem that Kubernetes cannot perceive CPU core attributes in a cloud-native environment, and improving processor resource utilization efficiency and application performance.

CN119938341BActive Publication Date: 2025-07-18北京志凌海纳科技股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510413519.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

In the prior art, Kubernetes cannot perceive the attribute characteristics of different CPU cores in the cloud-native environment, resulting in cross-NUMA access problems, affecting the performance of performance-sensitive applications.

Method used

By reading the memory architecture area identification and core type attributes from the processor information files and the system device subsystem, a processor core object pool is generated, and based on real-time updates of hardware changes, the appropriate processor core collection is accurately selected, and an independent resource group configuration is created for the container group to realize thread-level core binding.

Benefits of technology

It improves the efficiency of processor resources utilization, reduces memory access latency and I/O operation latency, ensures that the processor resources of critical threads are not disturbed, and improves the performance and stability of the application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938341B_ABST
    Figure CN119938341B_ABST
Patent Text Reader

Abstract

A CPU core binding method and system for cloud-native applications, which relates to the field of electrical digital data processing. It reads the memory architecture region identifier, generates processor core objects, stores the processor core objects in a container cluster to obtain a processor core object pool, updates the affected processor core objects in the processor core object pool according to hardware changes, selects a set of processor cores that meet the processor demand configuration from the processor core object pool and uses the set of processor cores as a container group, removes the selected set of processor cores from the processor core object pool to obtain a new container group, creates independent resource group configurations for each thread in the new container group, and binds each thread to the processor core with the corresponding number in the new container group according to the correspondence between the threads and the processor core numbers in the processor demand configuration. In this method, the present application improves the performance and stability of cloud-native applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic digital data processing, and particularly relates to a CPU core binding method and system for cloud-native applications. Background Art

[0002] With the continuous development of computer technology, multi-core modern processors (CPUs) have been widely popularized in personal computers and data center servers. In the traditional Linux software running environment, processes are generally managed through the systemd method or directly run in the background of the operating system through the command line. In this way, processes will frequently migrate between different cores, resulting in a large number of context switches, increasing system overhead, and at the same time increasing memory access latency due to cache invalidation, seriously affecting system performance.

[0003] In the related art, Kubernetes provides a Control CPU Management function to implement CPU core binding for processes in a cloud-native environment. This function divides all CPU cores into a shared core pool and an exclusive core pool. When a user creates a Pod and declares the required number of cores in the configuration file, Kubernetes will allocate the corresponding number of cores from the exclusive core pool for the Pod to use independently, thus avoiding the frequent migration of processes between different cores.

[0004] However, since this function treats all CPU cores equally and it is difficult to perceive the attribute characteristics of different CPU cores (such as numa regions, performance cores, or energy-efficient cores, etc.), problems such as cross-numa access are often caused during core allocation. Especially in performance-sensitive applications, this undifferentiated allocation method will reduce the performance of the application. Summary of the Invention

[0005] This application provides a CPU core binding method and system for cloud-native applications, which is used to improve the performance and stability of cloud-native applications.

[0006] In the first aspect, this application provides a CPU core binding method for cloud-native applications, which reads the memory architecture region identifier and core type attribute of the allocable processor cores from the processor information file;

[0007] Reads the memory architecture region identifiers of network cards and disk devices from the system device subsystem;

[0008] Generates a processor core object according to the memory architecture region identifier of the processor core, the core type attribute, and the memory architecture region identifiers of the network card and the disk device;

[0009] Stores the processor core object in the container cluster in the form of a custom resource to obtain a processor core object pool;

[0010] Detect hardware changes by subscribing to device events of the system kernel;

[0011] Update the processor core object pool according to the hardware changes;

[0012] Receive a container group creation request and obtain the processor requirement configuration, where the processor requirement configuration includes the number of cores, core type, memory architecture area requirements, and thread binding relationship;

[0013] Select a set of processor cores that meet the processor requirement configuration from the updated processor core object pool, and use the set of processor cores as the container group;

[0014] Remove the selected set of processor cores from the updated processor core object pool to obtain a new container group;

[0015] Create an independent resource group configuration for each thread in the new container group;

[0016] According to the correspondence between the threads and the processor core numbers in the processor requirement configuration, bind each thread to the processor core with the corresponding number in the new container group through the independent resource group configuration.

[0017] By adopting the above technical solution, the core attributes are read from the processor information file and the device information is read from the system device subsystem, a processor core object pool is established, and the object pool information is updated in real time based on hardware changes. When creating a container group, a suitable set of cores is selected from the object pool according to the processor requirement configuration, and an independent resource group configuration is created for each thread. This technical solution enables the container to more precisely allocate computing tasks to the most suitable processor cores for execution, reducing the memory access latency caused by cross-node access. At the same time, by managing the processor core object pool, the situation where multiple containers compete for the same processor core is avoided, improving the utilization efficiency of processor resources. Considering the memory architecture area identifiers of network card and disk devices, the system can schedule network-intensive or storage-intensive containers to the processor cores close to the corresponding devices, reducing the path length of data transmission and lowering the latency of I / O operations. The fine-grained thread-level core binding further ensures that the processor resources of critical threads are not interfered by other tasks, improving the performance stability of the application program.

[0018] Combined with some embodiments of the first aspect, in some embodiments, reading the memory architecture area identifier and core type attributes of the allocable processor cores from the processor information file specifically includes:

[0019] Obtain the physical identification numbers and physical location information of all processor cores in the system;

[0020] Read the cache hierarchy information of each processor core according to the physical identification number, where the cache hierarchy information includes the size of the first-level cache, the size of the second-level cache, and the three-level cache sharing relationship;

[0021] Construct a physical topology relationship graph of the processor core based on the physical location information and the cache hierarchy information;

[0022] According to the physical topology relationship graph, divide the processor cores with the same three-level cache sharing relationship into the same memory architecture area to obtain the memory architecture area identifier;

[0023] Read the static operating parameters of each processor core, where the static operating parameters include the maximum operating frequency, the minimum operating frequency, and the peak power consumption;

[0024] According to the static operating parameters, determine the processor cores with a maximum operating frequency greater than the preset frequency threshold and a peak power consumption greater than the preset power consumption threshold as performance cores, and determine the remaining processor cores as energy-efficient cores to obtain the core type attribute.

[0025] By adopting the above technical solutions, the memory architecture area is divided based on the physical topology relationship graph of the processor core, improving the accuracy of identifying the processor core group sharing the three-level cache. Combining the static operating parameters to classify the processor cores in terms of performance and energy efficiency, a fine-grained core profile is established, enabling the container scheduling system to fully understand the physical architecture and performance characteristics of the processor, so that the cache sharing advantage of the processor can be maximally utilized when allocating processor cores. By identifying the performance cores and energy-efficient cores, the system can preferentially schedule computationally intensive tasks to the performance cores for execution, and schedule background tasks to the energy-efficient cores for execution, achieving a reasonable allocation of processor resources, reducing energy consumption while ensuring performance. Threads with strong correlation can be scheduled to the cores sharing the cache, reducing the cache synchronization overhead and improving the cache hit rate.

[0026] Combined with some embodiments of the first aspect, in some embodiments, by subscribing to the device events of the system kernel, hardware changes are detected, specifically including:

[0027] Create an event listening thread for the system kernel and establish a communication channel with the kernel device management subsystem;

[0028] Only subscribe to the processor core hot plug event, the memory topology change event, and the device status change event to obtain the event filtering rule;

[0029] Register an event listening callback function with the kernel according to the event filtering rule;

[0030] Receive the device event message sent by the system kernel, where the device event message includes the event type, the device identifier, and the status flag;

[0031] Extract the hardware components and status changes involved in the device event message;

[0032] Obtain the hardware change based on the event type and status change.

[0033] By adopting the above technical solution, the event listening mechanism is used to capture the hardware change information of the system kernel in real time. Through precise event filtering rules, only key events such as processor core hot plugging, memory topology change, and device status change are concerned. The established communication channel ensures that the hardware change information can be transmitted to the container scheduling system in time, enabling the system to obtain immediate notification when a processor core fails, is newly added, or degrades in performance, and accordingly update the information in the processor core object pool. By adjusting the available processor core resources in real time, the system avoids scheduling containers to failed or degraded processor cores, improving the reliability of container operation. The event filtering mechanism reduces the overhead of the system for processing hardware events, reduces resource waste caused by irrelevant events, and improves the real-time accuracy of the processor resource view.

[0034] Combined with some embodiments of the first aspect, in some embodiments, after binding each thread to the corresponding numbered processor core in the new container group through independent resource group configuration, the method further includes:

[0035] Obtain the instruction execution sequences of the threads in the new container group, and detect the encryption operation instructions in the instruction execution sequences;

[0036] Statistically calculate the time proportion of the encryption operation instructions, and mark the threads with a time proportion greater than the preset threshold as encryption operation threads;

[0037] Before the encryption operation thread executes, adjust the working frequency of the processor core to which the encryption operation thread is bound to the maximum frequency threshold supported by the processor;

[0038] Collect the execution completion time of the encryption operation thread at the maximum frequency threshold, and use the execution completion time as the reference execution time;

[0039] Determine the expected execution interval of the next encryption operation thread according to the reference execution time, and adjust the working frequency of the processor core to which the encryption operation thread is bound to the maximum frequency threshold in advance before the expected execution interval.

[0040] By adopting the above technical solutions, the execution efficiency of the encryption operation thread is identified and optimized by analyzing the instruction execution sequence of the thread. The system accurately identifies the encryption operation threads that require high-performance support by counting the time proportion of the encryption operation instructions. After identifying the encryption operation threads, an accurate baseline execution time is established by adjusting the working frequency of the cores to which they are bound to the maximum frequency threshold. Based on the execution time prediction mechanism, the system can adjust the processor core frequency in advance to ensure that the encryption operation threads obtain the best processor performance support during execution, reducing the energy waste caused by the processor continuously working at a high frequency state, while ensuring that sufficient computing power can be obtained when encryption operations are required. Through the precise identification and targeted optimization of the encryption operation threads, the execution efficiency of the encryption operation is improved, and the impact of the encryption and decryption operations on the application performance is reduced.

[0041] Combined with some embodiments of the first aspect, in some embodiments, determining the predicted execution interval of the next encryption operation thread according to the baseline execution time specifically includes:

[0042] Recording the execution time intervals of the encryption operation threads for N consecutive times, where N is an integer greater than 1;

[0043] Calculating the periodic regularity index of the encryption operation based on the execution time intervals;

[0044] When the periodic regularity index is greater than the preset period threshold, determining the next predicted execution interval as one period after the completion time of the previous execution;

[0045] When the periodic regularity index is less than the preset period threshold, determining the next predicted execution interval as twice the baseline execution time after the completion time of the previous execution.

[0046] By adopting the above technical solutions, when the periodic regularity index is greater than the preset period threshold, it indicates that the encryption operation presents a stable periodic characteristic, and setting the predicted execution interval to one period can accurately predict the next execution time. When the periodic regularity index is less than the preset period threshold, it indicates that the execution of the encryption operation is uncertain, and setting the predicted execution interval to twice the baseline execution time can reserve sufficient execution time for the irregular encryption operation. Dynamically adjusting the predicted execution interval according to the actual execution characteristics of the encryption operation not only reduces the waste of resources caused by reserving too much execution time for periodic encryption operations, but also ensures that there is sufficient execution time for non-periodic encryption operations, thereby improving the utilization efficiency of the processor resources.

[0047] Combined with some embodiments of the first aspect, in some embodiments, before adjusting the working frequency of the processor core to which the encryption operation thread is bound to the maximum frequency threshold before the predicted execution interval, the method further includes:

[0048] Obtain the expected execution intervals of the encryption operation threads of all container groups;

[0049] Calculate the overlapping time periods of the expected execution intervals;

[0050] When it is detected that the number of encryption operation threads in the overlapping time period exceeds the preset concurrency threshold, allocate the encryption operation threads to different time slices according to their priorities;

[0051] According to the allocation results of the time slices, adjust the operating frequencies of the processor cores bound to each encryption operation thread in sequence.

[0052] By adopting the above technical solution, obtaining the expected execution intervals of the encryption operation threads of all container groups and calculating the overlapping time periods, the system can discover the situation where multiple encryption operation threads execute concurrently in the same time period. When the number of encryption operation threads in the overlapping time period exceeds the preset concurrency threshold, allocating the threads to different time slices according to their priorities can avoid too many encryption operation threads competing for processor resources simultaneously. Making a staggered adjustment to the operating frequencies of the processor cores enables each encryption operation thread to execute in different high-frequency time windows, reducing the performance loss caused by resource competition, ensuring the execution efficiency of the encryption operation, reducing the overall power consumption of the processor, and at the same time avoiding the problem of the processor temperature rising sharply due to intensive concurrent encryption operations.

[0053] Combined with some embodiments of the first aspect, in some embodiments, allocating the encryption operation threads to different time slices according to their priorities specifically includes:

[0054] Obtain the service level identifiers of the container groups to which each encryption operation thread belongs;

[0055] Perform a priority sorting on the encryption operation threads according to the service level identifiers to obtain a sorting result;

[0056] Allocate the top preset number of encryption operation threads according to the sorting result to the original expected execution intervals;

[0057] Allocate the encryption operation threads other than the top preset number to the nearest idle time slice.

[0058] By adopting the above technical solution, arranging the encryption operation threads with higher rankings in the original expected execution intervals ensures the timely response of high-priority services. And scheduling the remaining encryption operation threads to the nearest idle time slices can make full use of the idle computing resources of the processor, while ensuring the performance of critical services, also realizing the reasonable reuse of processor resources, improving the system's processing ability for encryption operation tasks with different priorities, enabling the limited high-performance computing resources to give priority to important services, and taking into account the execution requirements of other services at the same time.

[0059] In a second aspect, an embodiment of the present application provides a CPU core binding system for cloud-native applications. The CPU core binding system for cloud-native applications includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code. The computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the system to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0060] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions, which when running on the system, enable the system to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0061] In a fourth aspect, an embodiment of the present application provides a computer program product, which when running on the system, enables the system to execute the method described in any possible implementation manner in the first aspect.

[0062] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0063] 1. The present application provides a CPU core binding method for cloud-native applications, which reads core attributes from a processor information file and device information from a system device subsystem, establishes a processor core object pool, and updates the object pool information in real time based on hardware changes. When creating a container group, a suitable core set is selected from the object pool according to the processor requirements configuration, and an independent resource group configuration is created for each thread. This technical solution enables the container to more precisely allocate computing tasks to the most suitable processor cores for execution, reducing the memory access latency caused by cross-node access. At the same time, through the pooled management of processor cores, the situation where multiple containers compete for the same processor core is avoided, improving the utilization efficiency of processor resources. Considering the memory architecture region identifiers of network cards and disk devices, the system can schedule network-intensive or storage-intensive containers to the processor cores close to the corresponding devices, reducing the path length of data transmission and lowering the latency of I / O operations. The fine-grained core binding at the thread level further ensures that the processor resources of critical threads are not interfered with by other tasks, improving the performance stability of the application program.

[0064] 2. This application provides a CPU core binding method for cloud-native applications. By analyzing the instruction execution sequence of threads, the execution efficiency of encryption operation threads is identified and optimized. The system accurately identifies the encryption operation threads that require high-performance support by statistically analyzing the time proportion of encryption operation instructions. After identifying the encryption operation threads, an accurate baseline execution time is established by adjusting the working frequency of their bound cores to the maximum frequency threshold. Based on the execution time prediction mechanism, the system can adjust the processor core frequency in advance to ensure that the encryption operation threads obtain the best processor performance support during execution, reducing the energy waste caused by the processor continuously working at a high frequency state, and at the same time ensuring that sufficient computing power can be obtained when encryption operations are required. Through the precise identification and targeted optimization of encryption operation threads, the execution efficiency of encryption operations is improved, and the impact of encryption and decryption operations on the performance of the application program is reduced.

[0065] 3. This application provides a CPU core binding method for cloud-native applications. By obtaining the predicted execution intervals of all container group encryption operation threads and calculating the overlapping time periods, the system can discover the situation where multiple encryption operation threads execute concurrently in the same time period. When the number of encryption operation threads in the overlapping time period exceeds the preset concurrency threshold, allocating the threads to different time slices according to the priority can avoid too many encryption operation threads competing for processor resources at the same time. The working frequency of the processor core is adjusted in a staggered manner, so that each encryption operation thread executes in different high-frequency time windows, reducing the performance loss caused by resource competition, ensuring the execution efficiency of encryption operations, reducing the overall power consumption of the processor, and at the same time avoiding the problem of the processor temperature rising sharply due to intensive concurrent encryption operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 is a flowchart of a CPU core binding method for a cloud-native application in an embodiment of the present application.

[0067] Figure 2 is a flowchart of a processor frequency optimization method for a specific computing task in an embodiment of the present application.

[0068] Figure 3 is a schematic structural diagram of an entity device of a CPU core binding system for a cloud-native application provided in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any and all possible combinations including one or more of the listed items.

[0070] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and should not be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0071] Next, an embodiment is used and combined with Figure 1 to describe a method for binding CPU cores of a cloud-native application in the embodiments of the present application:

[0072] Please refer to Figure 1 , which is a schematic flowchart of a method for binding CPU cores of a cloud-native application in the embodiments of the present application.

[0073] S101. Read the memory architecture region identifier and core type attribute of the allocable processor cores from the processor information file;

[0074] The system reads the memory architecture region identifier and core type attribute of the allocable processor cores from the processor information file, specifically: obtain the physical identification numbers and physical location information of all processor cores in the system; read the cache hierarchy structure information of each processor core according to the physical identification number, and the cache hierarchy structure information includes the size of the first-level cache, the size of the second-level cache, and the three-level cache sharing relationship; construct a physical topology relationship graph of the processor cores based on the physical location information and the cache hierarchy structure information; divide the processor cores with the same three-level cache sharing relationship into the same memory architecture region according to the physical topology relationship graph to obtain the memory architecture region identifier; read the static operating parameters of each processor core, and the static operating parameters include the maximum operating frequency, the minimum operating frequency, and the peak power consumption; determine the processor cores with the maximum operating frequency greater than the preset frequency threshold and the peak power consumption greater than the preset power consumption threshold as performance cores according to the static operating parameters, and determine the remaining processor cores as energy-efficient cores to obtain the core type attribute.

[0075] In this step, the system reads the memory architecture region identifiers and core type attributes of the allocable processor cores from the processor information file. This step aims to obtain the detailed information of the processor cores in the system, providing a basis for subsequent resource allocation and management. The system can obtain processor information through various methods, such as directly accessing hardware registers, calling API interfaces provided by the operating system, etc.

[0076] Specifically, the system first obtains the physical identification numbers and physical location information of all processor cores in the system. The physical identification number can be the unique identifier of the processor core, such as the APIC ID, etc.; the physical location information can include the processor Socket number, the location of the core in the physical package, etc. Then, the system reads the cache hierarchy structure information of each processor core according to the physical identification number, including the size of the first-level cache, the size of the second-level cache, and the three-level cache sharing relationship. This information can be obtained from the return value of the CPUID instruction of the processor. Next, the system constructs a physical topology relationship graph of the processor cores based on the physical location information and the cache hierarchy structure information to reflect the physical connection and cache sharing relationship between the cores. On this basis, the system divides the processor cores with the same three-level cache sharing relationship into the same memory architecture region to obtain the memory architecture region identifier. Finally, the system reads the static operating parameters of each processor core, including the maximum operating frequency, the minimum operating frequency, and the peak power consumption, etc., and divides the processor cores into performance cores and energy-efficient cores according to the preset frequency and power consumption thresholds to obtain the core type attributes.

[0077] S102. Read the memory architecture region identifiers of the network card and disk device from the system device subsystem;

[0078] In this step, the system reads the memory architecture region identifiers of the network card and disk device from the device subsystem. The purpose of this step is to understand the physical location relationship between the I / O devices and the processor cores, providing support for implementing NUMA-aware resource scheduling.

[0079] In specific implementation, the system can call the device management interface provided by the operating system kernel, such as the sysfs file system of Linux, to obtain the NUMA Node information to which each I / O device is connected. For PCIe devices, the system can also obtain its NUMA Node ID by parsing the configuration space register of the device. These information can be mapped to the memory architecture region identifiers of the processor cores to establish the corresponding relationship between the I / O devices and the processor cores.

[0080] In some scenarios, I / O devices may be connected to the processor through components such as PCIe switches, making the physical location relationship between the devices and the processor complex. In response to this situation, the system can further analyze the PCIe topology to identify the processor Socket and memory controller to which the device is actually connected, so as to correctly determine its NUMA affiliation. In addition, for old devices that do not support NUMA awareness, the system can divide them into the default memory architecture area according to the bus number and device number to which they are connected.

[0081] S103. Generate a processor core object according to the memory architecture area identifier of the processor core, the core type attribute, and the memory architecture area identifiers of the network card and disk devices;

[0082] In this step, the system generates a processor core object for management and allocation according to the previously obtained attribute information of the processor core and I / O devices. This step aims to aggregate and abstract the scattered hardware resource information to form a unified resource view for subsequent resource control operations.

[0083] In specific implementation, the system can define a data structure of the processor core object, including fields such as the memory architecture area identifier, the core type attribute, and the list of associated local I / O devices. Then, the system traverses each processor core and fills in the corresponding fields according to its memory architecture area identifier and core type attribute. At the same time, the system also needs to traverse each network card and disk device, associate its memory architecture area identifier with the processor core object, and construct a complete resource topology.

[0084] In the process of generating the processor core object, the system may need to handle some special situations. For example, in some processors with asymmetric architectures, different cores may have different instruction set extensions or functional units, resulting in significant differences in their performance and power consumption characteristics. In this regard, when generating the core object, the system can introduce more fine-grained attribute descriptions, such as the instruction set architecture version, special function flags, etc., to fully characterize the heterogeneous characteristics of the core. In addition, considering the diversity of node hardware configurations in the cloud environment, the system can also support the dynamic extension and customization of the processor core object, allowing users to add custom attribute fields according to needs to adapt to different application scenarios.

[0085] S104. Store the processor core object in the container cluster in the form of a custom resource to obtain a processor core object pool;

[0086] In this step, the system stores the generated processor core objects in the container cluster in the form of custom resources, constructing a global processor core object pool. The purpose of this is to unify the hardware resource information with the resource model of the container cluster, enabling the container orchestration system to perceive and schedule the underlying CPU resources.

[0087] In specific implementation, the system can utilize the Custom Resource Definition (CRD) mechanism provided by Kubernetes to define a new resource type to describe the processor core objects. For example, a CRD named "ProcessorCore" can be defined, which includes fields such as memory architecture region identifiers and core type attributes. Then, the system creates each processor core object as a custom resource instance of the "ProcessorCore" type and stores it in the etcd database of Kubernetes. These processor core resource instances logically form a cluster-wide object pool for the container orchestration system to query and use.

[0088] During the process of storing the processor core objects, the system needs to pay attention to the version management and upgrade issues of the custom resources. Since the hardware configuration may change at any time, the definition of the processor core objects may also need to be adjusted accordingly. For this, the system can, through the version control mechanism of the CRD, define multiple versions for the "ProcessorCore" resource type and support automatic conversion and upgrade between versions. In addition, considering the expansion of the cluster scale, the access and management of the processor core object pool may also face performance challenges. For this, the system can explore mechanisms based on caching and sharding to improve the query efficiency and scalability of the object pool.

[0089] S105: Detect hardware changes by subscribing to device events of the system kernel;

[0090] The system detects hardware changes by subscribing to device events of the system kernel, specifically as follows: Create an event listening thread for the system kernel and establish a communication channel with the kernel device management subsystem; Only subscribe to processor core hot-plug events, memory topology change events, and device status change events to obtain event filtering rules; Register an event listening callback function with the kernel according to the event filtering rules; Receive device event messages sent by the system kernel, where the device event messages include event types, device identifiers, and status flags; Extract the hardware components and status changes involved in the device event messages; Obtain the hardware changes based on the event types and status changes.

[0091] In this step, the system monitors the changes in the hardware configuration in real time by subscribing to the device events of the kernel. The purpose of this step is to ensure that the processor core object pool can dynamically adapt to the changes in the underlying hardware and avoid the problem of inconsistent resource information.

[0092] When implementing specifically, the system first needs to create a dedicated kernel event listening thread and establish a communication channel with the device management subsystem of the kernel. Then, according to the predefined event filtering rules, the system selectively subscribes to events related to processor cores and memory topologies, such as processor core hot-plug events, memory topology change events, and device status change events. These event filtering rules can be set and adjusted through configuration files or API interfaces. Next, the system registers a callback function for event listening with the kernel so that it can be notified in a timely manner when relevant events occur. When the listening thread receives the device event message sent by the kernel, the system extracts the affected hardware components and status change information from it, and then determines whether there has been a change in the hardware configuration.

[0093] During the process of monitoring hardware changes, the system may need to handle abnormal situations such as event loss and out-of-order. For example, in a high-concurrency scenario, the device event messages generated by the kernel may exceed the processing capacity of the listening thread, resulting in some events being discarded. In this regard, the system can set an appropriate buffer at the source of kernel events or adopt the method of parallel processing with multiple listening threads to improve the reliability and real-time performance of event processing. In addition, since there may be dependencies between hardware change events, such as changes in the cache topology causing changes in the attributes of processor cores, etc., when the system analyzes the impact of events, it needs to consider the causal order relationship of events to ensure the correctness of processing.

[0094] S106. Update the processor core object pool according to the hardware change;

[0095] In this step, the system updates the affected objects in the processor core object pool according to the detected hardware changes. The purpose of this step is to keep the resource view synchronized with the physical hardware state and ensure that subsequent container scheduling decisions are based on the latest hardware information.

[0096] During specific implementation, the system first needs to locate the affected objects in the processor core object pool based on the device identifiers in the hardware change event. For example, if the event indicates that a processor core is hot-plugged, the system needs to find the corresponding "ProcessorCore" resource instance of this core in the object pool. Next, the system updates the attribute fields of the relevant processor core objects according to the reported hardware state changes in the event. For instance, if a core changes from a performance core to an energy-efficient core, the system needs to modify the "CoreType" attribute value of the corresponding object. After completing the attribute update, the system writes the modified processor core objects back to the etcd database of Kubernetes to ensure that other components in the cluster can see the latest resource status.

[0097] During the process of updating the processor core objects, the system needs to pay attention to the issues of concurrent modification and data consistency. In a cloud environment, multiple hardware events may occur simultaneously, resulting in concurrent updates to the same processor core object. In response, the system can utilize the resource version control and optimistic lock mechanism provided by Kubernetes to detect and handle concurrent modifications by comparing the version numbers of the objects. Additionally, since the scheduler and monitoring components of Kubernetes also access the processor core object pool, the system also needs to consider the visibility and consistency of data when updating the objects. To this end, the system can utilize the Watch mechanism of etcd to actively notify relevant components when the objects change, so as to reduce the inconsistent window period; or adopt the two-phase commit method to encapsulate the modification operation of the objects into a transaction to ensure atomicity.

[0098] S107. Receive a container group creation request and obtain the processor requirement configuration, where the processor requirement configuration includes the number of cores, core type, memory architecture region requirements, and thread binding relationship;

[0099] In this step, the system receives a container group (Pod) creation request from a user or an upper-layer system and extracts the requirement configuration information related to processor resources from the request. The purpose of this step is to clarify the specific requirements of the container group for processor resources and provide a basis for subsequent resource allocation decisions.

[0100] In specific implementation, the system can define a new container group annotation or label to describe the requirements of the container group for processor resources. For example, an annotation named "processor-request" can be defined, and its value is a JSON-formatted string containing fields such as the number of processor cores, core type, memory architecture region, and thread binding relationship. When the user creates a container group, this annotation can be added to the metadata of the container group to clarify its requirements for processor resources. After receiving the container group creation request, the system can parse the value of the "processor-request" annotation from the metadata of the request object to obtain the processor requirement configuration.

[0101] During the process of obtaining the processor requirement configuration, the system needs to perform a legality check on the configuration information provided by the user. For example, the number of cores cannot exceed the total number of cores of the node, and the identifier of the memory architecture region must actually exist in the cluster. In this regard, the system can pre-define a set of configuration specifications and constraints and check them item by item when parsing the configuration. If the configuration information is found to be illegal, the system can directly reject the container group creation request and return the corresponding error message to the user. In addition, considering the allocation granularity and limitations of processor resources, the system can also appropriately align and adjust the number of cores and memory regions configured by the user to better adapt to the characteristics of the underlying hardware.

[0102] S108. Select a set of processor cores that meet the processor requirement configuration from the updated processor core object pool, and use the set of processor cores as the container group.

[0103] In this step, the system selects the core resources that meet the requirements from the processor core object pool according to the processor requirement configuration of the container group, and binds the selected set of cores to the container group. This step aims to achieve precise matching and isolation of the container group to the physical cores, improving the performance and reliability of container applications.

[0104] In specific implementation, the system first needs to construct an algorithm for selecting processor cores according to the requirement configuration of the container group. This algorithm needs to comprehensively consider various factors such as the number of cores, core type, memory region, and NUMA locality to find the optimal core combination that meets the conditions from the processor core object pool. Among them, matching the core type can ensure that the container group obtains the required computing power, matching the memory region can minimize remote memory access, and matching the NUMA locality can improve the cache hit rate and memory bandwidth utilization. After finding the set of processor cores that meet the conditions, the system binds it to the requested container group and records the allocated core resources in the status information of the container group.

[0105] S109. Remove the selected set of processor cores from the updated processor core object pool to obtain a new container group;

[0106] In this step, the system removes the set of processor cores allocated to the container group from the core object pool to ensure exclusive use of these core resources. The purpose of this step is to achieve strong isolation between the container group and the host machine cores and avoid interference between different containers.

[0107] When specifically implemented, after the system completes the binding of the core resources and the container group, it needs to mark the corresponding "ProcessorCore" resource instances in the processor core object pool as allocated. This can be achieved by updating the status field of the resource object, such as setting the "Allocated" field to true. At the same time, the system also needs to record the information of the container group bound to these resource instances, such as the name and namespace of the container group, for subsequent status query and unbinding operations. After marking, the system writes the updated "ProcessorCore" resource object back to the etcd database of Kubernetes. In this way, when other container groups apply for processor resources, they will no longer select these allocated core objects, ensuring the exclusivity of the core resources.

[0108] S110. Create an independent resource group configuration for each thread in the new container group;

[0109] In this step, the system creates an independent resource group (cgroup) configuration for each thread in the newly created container group. The purpose of this step is to achieve finer-grained resource isolation and control and avoid resource competition among the threads inside the container.

[0110] When specifically implemented, the system first needs to determine the thread scope for which resource groups need to be created within the container group. This can be inferred from the metadata information of the container image, such as Entrypoint, CMD, etc., to deduce the thread structure of the main process inside the container. In addition, the system can also provide corresponding annotations or labels for the user to explicitly specify the threads that need to be isolated. Next, based on the binding relationship between the threads and the cores specified in the processor requirement configuration and the resource constraints of each core, the system generates corresponding cgroup parameters for each thread. These parameters can include CPU time quota, memory limit, cpuset, etc. After generation, the system calls the cgroup-related APIs, such as libcgroup or systemd, to create the corresponding cgroup hierarchies and subsystems on the host machine. Finally, the system adds the main process of the container group and its child threads to these cgroups to achieve refined resource management and control.

[0111] S111. According to the correspondence between threads and processor core numbers in the processor demand configuration, each thread is bound to the processor core with the corresponding number in the new container group through the configuration of an independent resource group.

[0112] In this step, the system binds each thread to the processor core with the corresponding number in the new container group through the configuration of an independent resource group according to the correspondence between threads and processor cores specified in the processor demand configuration. The purpose of this step is to further enhance the affinity between container applications and hardware and achieve predictable high performance.

[0113] When specifically implemented, when the system starts the container group, in addition to the conventional configurations such as environment variables and volume mounts, it is also necessary to set the processor affinity for the main process of the container. This can be achieved by modifying the startup parameters of the container. For example, add the "--cpuset-cpus" and "--cpuset-mems" options in the startup command of the Docker container, or set the "nodeSelector" and "affinity" fields in the definition of the Kubernetes Pod. The values of these parameters can be extracted from the processor demand configuration and mapped to the physical core numbers and NUMA nodes allocated by the system. After the container starts, its main process and child threads will automatically inherit these processor affinity configurations to achieve static binding of threads and cores.

[0114] In the process of implementing thread binding, the system needs to consider the differences between container runtimes and processor topologies. Different container runtimes, such as Docker, containerd, rkt, etc., may have differences in the interfaces and semantics of processor affinity configuration. Therefore, the system needs to provide an adaptation layer to map the processor demand configuration to the configuration parameters of different runtimes. In addition, in the NUMA architecture, threads can obtain lower memory access latency when running on cores within the local node. However, in some scenarios, threads may need to access the memory of remote nodes, such as cross-slot global synchronization operations, etc. In this regard, when generating the binding configuration, the system needs to take into account data locality and load balancing, and appropriately disperse threads to different NUMA nodes. At the same time, the system also needs to provide a certain degree of flexibility to allow users to manually adjust or cancel the thread binding settings for specific application scenarios.

[0115] In the above embodiments, core attributes are read from the processor information file and device information is read from the system device subsystem, a processor core object pool is established, and the object pool information is updated in real time based on hardware changes. When creating a container group, a suitable set of cores is selected from the object pool according to the processor requirements configuration, and an independent resource group configuration is created for each thread. This technical solution enables the container to more precisely allocate computing tasks to the most suitable processor cores for execution, reducing the memory access latency caused by cross-node access. At the same time, through the pooled management of processor core objects, the situation where multiple containers compete for the same processor core is avoided, improving the utilization efficiency of processor resources. Considering the memory architecture region identifiers of network cards and disk devices, the system can schedule network-intensive or storage-intensive containers to the processor cores close to the corresponding devices, reducing the path length of data transmission and the latency of I / O operations. The fine-grained core binding at the thread level further ensures that the processor resources of critical threads are not interfered with by other tasks, enhancing the performance stability of the application program.

[0116] Based on the above embodiments, in addition to improving the container performance through the processor core object pool and the fine-grained thread core binding mechanism, the present application also provides a processor frequency optimization method for specific computing tasks. Considering the common encryption operation requirements in cloud-native applications, the system realizes the intelligent adjustment of the working frequency of processor cores by identifying and analyzing the execution characteristics of encryption operation threads. This targeted frequency optimization method can not only improve the execution efficiency of encryption operations, but also reasonably control the energy consumption while ensuring performance. The following combines Figure 2 to describe a processor frequency optimization method for specific computing tasks in the embodiments of the present application:

[0117] Please refer to Figure 2 which is a flowchart of a processor frequency optimization method for specific computing tasks in the embodiments of the present application.

[0118] S201. Obtain the instruction execution sequences of each thread in the new container group, and detect the encryption operation instructions in the instruction execution sequences;

[0119] In this step, the system first obtains the instruction execution sequence of each thread in the newly created container group. This step aims to provide the necessary data basis for subsequent encryption operation detection and analysis. The instruction execution sequence reflects the actual running trajectory of the thread and records information such as the type and address of each instruction during the thread execution process.

[0120] In specific implementation, the system can adopt techniques such as dynamic instrumentation or static analysis to obtain the instruction sequence of a thread. Among them, dynamic instrumentation is a method of recording the execution of instructions in real time during program operation. The system can use tools such as ftrace or perf provided by the Linux kernel to insert probes at the system calls or function entrances of the container, capture the instruction stream information of each thread and record it. The advantage of this method is that it can obtain a detailed and accurate instruction sequence, but it may introduce a certain amount of runtime overhead. Another static analysis method is to disassemble or parse the symbol table of the executable file in the container image during the container image construction or deployment, and extract the instruction stream information in the code. The advantage of this method is that it does not need to modify the program and has less overhead, but the analysis accuracy may be affected by factors such as code obfuscation.

[0121] S202. Statistically calculate the time proportion of the encryption operation instructions, and mark the thread with a time proportion greater than the preset threshold as an encryption operation thread;

[0122] In this step, after the system detects the encryption operation instructions of the thread, it further statistically calculates the execution time proportion of these instructions. The purpose of this step is to identify the thread mainly engaged in encryption calculation and provide a targeted object for subsequent frequency tuning.

[0123] In specific implementation, the system can calculate the execution time of each encryption instruction based on the timestamp information in the instruction sequence, and accumulate it to obtain the total encryption operation time. At the same time, the system also needs to statistically calculate the total execution time of the thread, which can be obtained by subtracting the timestamp of the first instruction from the timestamp of the last instruction. Dividing the two can obtain the time proportion of the encryption operation instructions. To distinguish between the encryption operation thread and the ordinary thread, the system can set a threshold for the time proportion, such as 50%. When the time proportion of the encryption instructions of a thread exceeds this threshold, it is marked as an encryption operation thread.

[0124] S203. Before the encryption operation thread executes, adjust the working frequency of the processor core bound by the encryption operation thread to the maximum frequency threshold supported by the processor;

[0125] In this step, the system performs frequency tuning on the identified encryption operation thread. The purpose of this step is to accelerate the execution of the encryption instructions by increasing the working frequency of the processor core, thereby shortening the completion time of the encryption task.

[0126] During specific implementation, the system first needs to intercept the encryption operation thread before it executes. This can be achieved by setting a hook function in the thread scheduler, that is, triggering a callback function before the thread is scheduled to run on the CPU. In the callback function, the system can check the attributes of the thread to determine whether it is an encryption operation thread. If it is, the system continues with the frequency adjustment operation; if not, it simply returns without any processing. Next, the system needs to determine the processor core to which the encryption operation thread is bound. This can be obtained by reading the CPU affinity mask of the thread. Based on the bit numbers set in the mask, the system can know on which core the thread is currently running. After obtaining the target core, the system calls the DVFS (Dynamic Voltage and Frequency Scaling) interface of the processor, such as Intel's P-State or AMD's Cool'n'Quiet technology, to set the operating frequency of this core to the maximum supported threshold. This threshold can be obtained from the processor's specification document or queried in real time through the CPUID instruction.

[0127] During the process of adjusting the processor core frequency, the system needs to balance performance and power consumption. On the one hand, the higher the frequency, the faster the execution speed of the encryption operation, but it also means higher power consumption and heat generation. Therefore, when setting the maximum frequency threshold, the system needs to comprehensively consider factors such as the processor's heat dissipation capacity and power budget. In addition, blindly setting the frequency to the highest point is not always the optimal strategy. This is because increasing the frequency is usually accompanied by an increase in voltage, and too high a voltage may cause stability problems, leading to system crashes or data loss. To this end, the system can find an optimal frequency point that balances performance and stability by testing different frequency-voltage combinations. In addition, the system can also dynamically adjust the frequency threshold according to the urgency of the encryption task. For example, for some encryption tasks with high real-time requirements, such as financial transactions and security certifications, the system can appropriately increase the threshold to ensure performance; while for some offline batch processing tasks, such as data backup and log encryption, the system can appropriately lower the threshold to save energy consumption.

[0128] S204. Collect the execution completion time of the encryption operation thread under the maximum frequency threshold, and use the execution completion time as the baseline execution time;

[0129] In this step, after the system adjusts the core frequency of the encryption operation thread to the maximum threshold, it collects the execution completion time and uses this time as the reference value for subsequent optimization. The purpose of this step is to establish a reference standard for measuring the performance of the encryption operation and provide a feedback basis for dynamic frequency adjustment.

[0130] During specific implementation, the system can record the current timestamp immediately after the encryption operation thread finishes execution. Since the execution time of the thread may be very short, the accuracy of the timestamp needs to be high enough, at least reaching the microsecond level. To obtain more reliable statistical results, the system can measure the execution time of the thread multiple times continuously and calculate its average value. Considering that there may be multiple encryption operation threads in the system, and the computational amount and execution path of each thread may be different, the system needs to record the reference execution time for each thread separately. In addition, since the operating frequency of the processor may be affected by factors such as power consumption management and temperature control, resulting in the actual operating frequency being lower than the set maximum threshold. To eliminate the interference of these factors, the system can collect the actual operating frequency of the processor core while recording the reference execution time for subsequent analysis reference.

[0131] S205. Determine the predicted execution interval of the next encryption operation thread according to the reference execution time, and adjust the operating frequency of the processor core bound to the encryption operation thread to the maximum frequency threshold in advance before the predicted execution interval.

[0132] The system determines the predicted execution interval of the next encryption operation thread according to the reference execution time, specifically including: recording the execution time intervals of the encryption operation thread for N consecutive times, where N is an integer greater than 1; calculating the periodic regularity index of the encryption operation based on the execution time intervals; when the periodic regularity index is greater than the preset periodic threshold, determining the next predicted execution interval as one period after the completion moment of the previous execution; when the periodic regularity index is less than the preset periodic threshold, determining the next predicted execution interval as twice the reference execution time after the completion moment of the previous execution. Adjust the operating frequency of the processor core bound to the encryption operation thread to the maximum frequency threshold in advance before the predicted execution interval.

[0133] In this step, the system predicts the time interval of the next execution of the encryption operation thread according to the reference execution time obtained in the previous step, and adjusts the frequency of the processor core bound to the thread to the maximum threshold in advance before the interval arrives. The purpose of this step is to reduce the running time of the encryption operation thread in the low-frequency state through prediction and early frequency adjustment, and further improve the performance of encryption calculation.

[0134] In specific implementation, the system first needs to analyze the execution cycle pattern of the encryption operation thread based on historical execution information. To this end, the system can record the time intervals of consecutive executions of the thread and calculate statistical metrics such as the mean and variance. If the variance is small, it indicates that the execution of the thread has strong periodicity, and a fixed time interval can be used to predict the next execution time point. If the variance is large, it means that the execution of the thread is relatively random and it is difficult to make accurate predictions. In this case, the system can estimate a relatively loose time interval with reference to the benchmark execution time. After obtaining the predicted execution interval, the system needs to adjust the frequency of the processor core to the maximum threshold in advance by a certain amount of time before the arrival of this interval. The length of this advance amount depends on the time consumed by the frequency adjustment itself. Generally speaking, the frequency adjustment takes several milliseconds to dozens of milliseconds, so the advance amount should also be on this order of magnitude. To avoid performance degradation caused by frequent frequency adjustments, the system can also introduce a limit on the minimum adjustment interval, that is, there must be at least a certain time interval between two adjustments.

[0135] In the above embodiment, by analyzing the instruction execution sequence of the thread, the execution efficiency of the encryption operation thread is identified and optimized. The system accurately identifies the encryption operation threads that require high-performance support by counting the time proportion of encryption operation instructions. After identifying the encryption operation threads, an accurate benchmark execution time is established by adjusting the working frequency of the core to which they are bound to the maximum frequency threshold. Based on the execution time prediction mechanism, the system can adjust the processor core frequency in advance to ensure that the encryption operation thread obtains the best processor performance support during execution, reducing the energy waste caused by the processor continuously working at a high frequency state, and at the same time ensuring that the encryption operation can obtain sufficient computing power when needed. Through the precise identification and targeted optimization of the encryption operation thread, the execution efficiency of the encryption operation is improved, and the impact of encryption and decryption operations on the application performance is reduced.

[0136] Further, in another embodiment, before adjusting the working frequency of the processor core to which the encryption operation thread is bound to the maximum frequency threshold in advance before the predicted execution interval, it further includes:

[0137] Obtain the predicted execution intervals of the encryption operation threads of all container groups;

[0138] Calculate the overlapping time period of the predicted execution intervals;

[0139] When it is detected that the number of encryption operation threads in the overlapping time period exceeds the preset concurrency threshold, allocate the encryption operation threads to different time slices according to their priorities. Specifically: obtain the service level identifiers of the container groups to which each encryption operation thread belongs;

[0140] Sort the encryption operation threads according to their priorities based on the service level identifiers to obtain the sorting result;

[0141] Allocate the top pre-set number of encrypted operation threads to the original expected execution intervals according to the sorting results;

[0142] Allocate the encrypted operation threads other than the top pre-set number to the nearest idle time slice;

[0143] According to the allocation results of the time slices, sequentially adjust the working frequencies of the processor cores bound by each encrypted operation thread.

[0144] In this embodiment, the system further optimizes the frequency adjustment strategy of the encrypted operation threads to adapt to the scenario where multiple container groups concurrently execute encrypted tasks. The purpose of this optimization is to avoid resource competition problems caused by too many threads adjusting frequencies simultaneously while ensuring the performance of high-priority tasks.

[0145] When specifically implemented, the system first needs to obtain the expected execution intervals of the encrypted operation threads in all container groups. This can be achieved by deploying a monitoring agent in each container group. The agent is responsible for collecting the execution information of the encrypted threads in this group and regularly reporting it to the global scheduler. The scheduler aggregates the data from each container group to form a global view of the encrypted task execution. Next, the system performs an overlap analysis on the expected execution intervals of all threads to find the set of threads with time conflicts. If the number of encrypted threads exceeds the pre-set concurrency threshold (such as half of the number of processor cores) during a certain overlapping time period, it indicates a risk of resource competition, and it is necessary to perform sharding scheduling on the execution time of the threads.

[0146] After determining that sharding scheduling is required, the system further sorts the threads according to the service level of the container groups to which they belong. The service level can be specified by the user when creating the container group or automatically divided according to the importance of the business carried by the container group. The result of the priority sorting determines the resource allocation order of the threads in the sharding scheduling. To balance the timeliness of the tasks, the system allocates the top N high-priority threads to the original expected execution intervals to ensure that they can obtain processor resources within the expected time. The remaining low-priority threads are allocated to the nearest idle time slices. If all time slices are occupied, the system can also delay some low-priority threads to the next scheduling cycle according to the priority.

[0147] According to the allocation result of time segments, the system finally adjusts the frequencies of the processor cores bound to each thread. For high-priority threads allocated to the original execution interval, the system only needs to increase the frequency of the corresponding core to the maximum threshold before their expected start time. For low-priority threads allocated to other time segments, the system needs to dynamically adjust the core frequency before the threads actually execute. Since the frequency adjustment itself also has a certain time overhead, the system needs to reserve a certain buffer for frequency adjustment when specifying the start time of a time segment. In addition, if multiple threads are allocated to the same time segment, the system also needs to control the order of core frequency adjustment to reduce the performance loss caused by frequent frequency modulation. For example, the time segment can be further divided into multiple sub-segments, each sub-segment corresponding to a thread, and the frequencies are adjusted in turn according to the priorities of the threads.

[0148] In the above embodiment, by obtaining the expected execution intervals of all container group encryption operation threads and calculating the overlapping time periods, the system can discover the situation where multiple encryption operation threads execute concurrently in the same time period. When the number of encryption operation threads in the overlapping time period exceeds the preset concurrency threshold, allocating the threads to different time segments according to the priorities can avoid too many encryption operation threads competing for processor resources at the same time. Adjusting the working frequencies of the processor cores in a staggered manner enables each encryption operation thread to execute within different high-frequency time windows, reducing the performance loss caused by resource competition, ensuring the execution efficiency of the encryption operation, reducing the overall power consumption of the processor, and at the same time avoiding the problem of the processor temperature rising sharply due to intensive concurrent encryption operations.

[0149] The system in the embodiment of the present invention application will be described from the perspective of hardware processing. Please refer to Figure 3 , which is a schematic structural diagram of an entity device of a CPU core binding system for cloud-native applications provided by an embodiment of the present application.

[0150] It should be noted that Figure 3 the structure of the system shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0151] As Figure 3As shown, the system includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 302 or the program loaded from the storage section 308 into the Random Access Memory (RAM) 303, such as executing the methods in the above embodiments. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0152] The following components are connected to the I / O interface 305: an input section 306 including a camera, an infrared sensor, etc.; an output section 307 including a Liquid Crystal Display (LCD), a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read from it can be installed into the storage section 308 as needed.

[0153] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the Central Processing Unit (CPU) 301, various functions defined in the present invention are executed.

[0154] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above.

[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0156] As another aspect, the present invention also provides a computer-readable storage medium, which may be included in the system described in the above embodiments; or may exist separately without being assembled into the system. The above storage medium carries one or more computer programs, and when the one or more computer programs are executed by a processor of a system, the system implements the method provided in the above embodiments.

[0157] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0158] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as "if...", or "after...", or "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted as "if determining...", or "in response to determining...", or "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".

[0159] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.

[0160] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by relevant hardware instructed by a computer program. This program can be stored in a computer-readable storage medium. When this program is executed, it can include the processes of the above method embodiments. The aforementioned storage medium includes various media that can store program codes, such as ROM, random access memory (RAM), magnetic disks, or optical discs.

Claims

1. A CPU core binding method for cloud-native applications, characterized in that, including: Read the memory architecture region identifier and core type attribute of the allocable processor cores from the processor information file; Read the memory architecture region identifiers of the network card and disk devices from the system device subsystem; Generate a processor core object based on the memory architecture region identifier, core type attribute of the processor core, and the memory architecture region identifiers of the network card and disk devices; Store the processor core object in the container cluster in the form of a custom resource to obtain a processor core object pool; Detect hardware changes by subscribing to device events of the system kernel, specifically including: Create an event listening thread for the system kernel and establish a communication channel with the kernel device management subsystem; Only subscribe to processor core hot-plug events, memory topology change events, and device status change events to obtain an event filtering rule; Register an event listening callback function with the kernel according to the event filtering rule; Receive the device event message sent by the system kernel, where the device event message includes an event type, a device identifier, and a status flag; Extract the hardware components and status changes involved in the device event message; Obtain the hardware change based on the event type and the status change; Update the processor core object pool according to the hardware change; Receive a container group creation request and obtain the processor requirement configuration, where the processor requirement configuration includes the core number, core type, memory architecture region requirements, and thread binding relationship; Select a set of processor cores that meet the processor requirement configuration from the updated processor core object pool and use the set of processor cores as the container group; Remove the selected set of processor cores from the updated processor core object pool to obtain a new container group; Create an independent resource group configuration for each thread in the new container group; According to the correspondence between the threads and the processor core numbers in the processor requirement configuration, bind each thread to the corresponding numbered processor core in the new container group through the independent resource group configuration.

2. The method according to claim 1, wherein The step of reading the memory architecture region identifier and core type attribute of the allocable processor cores from the processor information file specifically includes: Obtain the physical identification numbers and physical location information of all processor cores in the system; Read the cache hierarchy structure information of each processor core according to the physical identification number, where the cache hierarchy structure information includes the size of the first-level cache, the size of the second-level cache, and the third-level cache sharing relationship; Construct a physical topology relationship graph of the processor cores based on the physical location information and the cache hierarchy structure information; Divide the processor cores with the same third-level cache sharing relationship into the same memory architecture region according to the physical topology relationship graph to obtain the memory architecture region identifier; Read the static operating parameters of each processor core, where the static operating parameters include the maximum operating frequency, the minimum operating frequency, and the peak power consumption; Determine the processor cores with a maximum operating frequency greater than a preset frequency threshold and a peak power consumption greater than a preset power consumption threshold as performance cores according to the static operating parameters, and determine the remaining processor cores as energy-efficient cores to obtain the core type attribute.

3. The method according to claim 1, wherein After binding each of the threads to the corresponding numbered processor core in the new container group through the independent resource group configuration, the method further includes: Obtain the instruction execution sequences of the threads in the new container group, and detect the encryption operation instructions in the instruction execution sequences; Statistically calculate the time proportion of the encryption operation instructions, and mark the threads with a time proportion greater than a preset threshold as encryption operation threads; Before the encryption operation thread is executed, adjust the operating frequency of the processor core to which the encryption operation thread is bound to the maximum frequency threshold supported by the processor; Collect the execution completion time of the encryption operation thread at the maximum frequency threshold, and use the execution completion time as the reference execution time; Determine the expected execution interval of the next encryption operation thread according to the reference execution time, and adjust the operating frequency of the processor core to which the encryption operation thread is bound to the maximum frequency threshold in advance before the expected execution interval.

4. The method according to claim 3, wherein The determining the expected execution interval of the next encryption operation thread according to the reference execution time specifically includes: Record the execution time intervals of the encryption operation thread for N consecutive times, where N is an integer greater than 1; calculate the periodic regularity index of the encryption operation based on the execution time intervals; When the periodic regularity index is greater than a preset period threshold, determine the next expected execution interval as one period after the previous execution completion time; When the periodic regularity index is less than the preset period threshold, determine the next expected execution interval as twice the reference execution time after the previous execution completion time.

5. The method according to claim 3 or 4, characterized in that Before adjusting the operating frequency of the processor core to which the encryption operation thread is bound to the maximum frequency threshold in advance before the expected execution interval, the method further includes: Obtain the expected execution intervals of the encryption operation threads of all container groups; calculate the overlapping time periods of the expected execution intervals; When it is detected that the number of encryption operation threads in the overlapping time period exceeds a preset concurrency threshold, allocate the encryption operation threads to different time slices according to the priorities; According to the allocation result of the time slices, sequentially adjust the operating frequencies of the processor cores to which the encryption operation threads are bound.

6. The method according to claim 5, wherein The allocating the encryption operation threads to different time slices according to the priorities specifically includes: Obtain the service level identifiers of the container groups to which the encryption operation threads belong; Perform priority sorting on the encryption operation threads according to the service level identifiers to obtain a sorting result; Allocate the top preset number of encryption operation threads to the original expected execution intervals according to the sorting result; allocate the encryption operation threads other than the top preset number to the nearest idle time slices.

7. A CPU core binding system for cloud-native applications, characterized in that, The system includes: One or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the system to execute the method according to any one of claims 1-6.

8. A computer-readable storage medium, comprising instructions, characterized in that, When the instructions are run on the system, cause the system to execute the method according to any one of claims 1-6.

9. A computer program product, characterized in that, When the computer program product is run on the system, cause the system to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Low-power-consumption lightweight virtualization method for edge computing

    CN112162826A

  • Heterogeneous hardware-based uniform resource pooling container scheduling engine and scheduling method thereof

    CN112363820A