CPU core binding method and system for cloud native application

By generating a processor core object pool and real-time update based on hardware changes, combined with the processor requirements configuration in container group creation requests, the precise binding of the processor core in cloud-native applications is achieved, cross-NUMA access problems are solved, and application performance and stability are improved.

CN119938341AActive Publication Date: 2025-05-06北京志凌海纳科技股份有限公司

Patent Information

Application Number
CN202510413519.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-06
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

In a cloud-native environment, it is difficult for the existing technology to effectively perceive and utilize the attribute characteristics of multi-core processors, resulting in cross-NUMA access problems and affect application performance.

Method used

By reading the memory architecture area identification and core type attributes from the processor information files and the system device subsystem, a processor core object pool is generated and updated in real time based on hardware changes. Combining the processor requirements configuration in the container group creation request, select the appropriate processor core collection, and create an independent resource group configuration for each thread to achieve accurate processor core binding.

Benefits of technology

Improve the performance and stability of cloud-native applications, reduce memory access latency caused by cross-node access, avoid multiple containers competing for the same processor core, and optimize the utilization efficiency of processor resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938341A_ABST
    Figure CN119938341A_ABST
Patent Text Reader

Abstract

The invention discloses a CPU core binding method and system for a cloud native application, and relates to the field of electrical digital data processing. Generating a processor core object; storing the processor core object in a container cluster to obtain a processor core object pool; updating the affected processor core object in the processor core object pool according to the hardware change; selecting a processor core set meeting processor demand configuration from the processor core object pool, and taking the processor core set as a container group; removing the selected processor core set from the processor core object pool to obtain a new container group; creating an independent resource group configuration for each thread in the new container group; and binding each thread to the processor core with the corresponding serial number in the new container group according to the corresponding relationship between the threads in the processor demand configuration and the serial number of the processor core. According to the method, the performance and the stability of the cloud native application are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of electronic digital data processing, and in particular, to a CPU core binding method and system for cloud native applications. Background Art

[0002] With the continuous development of computer technology, multi-core modern processors (CPU) have been widely used in personal computers and data center servers. In the traditional Linux software running environment, processes are generally managed through systemd or run directly in the background of the operating system through the command line. In this way, processes will frequently migrate between different cores, resulting in a large number of context switches, increasing system overhead, and increasing memory access latency due to cache failure, seriously affecting system performance.

[0003] In related technologies, Kubernetes provides the Control CPU Management function to implement process CPU core binding in a cloud-native environment. This function divides all CPU cores into a shared core pool and an exclusive core pool. When a user creates a Pod and declares the required number of cores in the configuration file, Kubernetes will allocate the corresponding number of cores from the exclusive core pool to the Pod for independent use, thereby avoiding frequent migration of processes between different cores.

[0004] However, since this function treats all CPU cores equally, it is difficult to perceive the attribute characteristics of different CPU cores (such as NUMA areas, performance cores, or energy-efficiency cores). When allocating cores, it often leads to problems such as cross-NUMA access. Especially in performance-sensitive applications, this indiscriminate allocation method will reduce application performance. Summary of the invention

[0005] The present application provides a CPU core binding method and system for cloud native applications, which are used to improve the performance and stability of cloud native applications.

[0006] In a first aspect, the present application provides a CPU core binding method for a cloud native application, which reads a memory architecture region identifier and a core type attribute of an allocatable processor core from a processor information file; Read the memory architecture region identifiers of network cards and disk devices from the system device subsystem; Generate a processor core object according to the memory architecture region identifier of the processor core, the core type attribute, and the memory architecture region identifiers of the network card and the disk device; The processor core objects are stored in the container cluster as custom resources to obtain a processor core object pool; Detect hardware changes by subscribing to device events in the system kernel; Update the processor core object pool according to hardware changes; Receive container group creation requests and obtain processor requirement configuration, including the number of cores, core type, memory architecture area requirements, and thread binding relationships; Selecting a processor core set that meets the processor requirement configuration from the updated processor core object pool, and using the processor core set as a container group; removing the selected processor core set from the updated processor core object pool to obtain a new container group; Create a separate resource group configuration for each thread in the new container group; According to the correspondence between threads and processor core numbers in the processor requirement configuration, each thread is bound to a processor core with a corresponding number in the new container group through an independent resource group configuration.

[0007] By adopting the above technical solution, the core attributes are read from the processor information file and the device information is read from the system device subsystem, a processor core object pool is established, and the object pool information is updated in real time based on hardware changes. When the container group is created, the appropriate core set is selected from the object pool according to the processor demand configuration, and an independent resource group configuration is created for each thread. This technical solution enables the container to more accurately assign computing tasks to the most suitable processor core for execution, reducing the memory access delay caused by cross-node access. At the same time, by managing the processor core object pool, it avoids the situation where multiple containers compete for the same processor core, and improves the utilization efficiency of processor resources. Considering the memory architecture area identification of the network card and disk device, the system can schedule network-intensive or storage-intensive containers to the processor core close to the corresponding device, reducing the path length of data transmission and reducing the delay of I / O operations. The refined binding of cores at the thread level further ensures that the processor resources of key threads are not interfered with by other tasks, improving the performance stability of the application.

[0008] In conjunction with some embodiments of the first aspect, in some embodiments, reading the memory architecture region identifier and core type attribute of the allocatable processor core from the processor information file specifically includes: Get the physical identification number and physical location information of all processor cores in the system; Reading cache hierarchy information of each processor core according to the physical identification number, the cache hierarchy information including the first-level cache size, the second-level cache size and the third-level cache sharing relationship; Construct a physical topology diagram of the processor core based on the physical location information and cache hierarchy information; According to the physical topology relationship diagram, the processor cores having the same three-level cache sharing relationship are divided into the same memory architecture region to obtain a memory architecture region identifier; Read the static operating parameters of each processor core, including the maximum operating frequency, minimum operating frequency and peak power consumption; According to the static operating parameters, the processor cores whose maximum operating frequency is greater than a preset frequency threshold and whose peak power consumption is greater than a preset power consumption threshold are determined as performance cores, and the remaining processor cores are determined as energy efficiency cores to obtain core type attributes.

[0009] By adopting the above technical solution, the memory architecture area is divided based on the physical topology relationship diagram of the processor core, which improves the accuracy of identifying the processor core group that shares the third-level cache. The performance and energy efficiency of the processor cores are classified in combination with static operating parameters, and a fine-grained core profile is established, so that the container scheduling system can fully understand the physical architecture and performance characteristics of the processor, so that the cache sharing advantage of the processor can be maximized when allocating processor cores. By identifying performance cores and energy-efficient cores, the system can prioritize the scheduling of computing-intensive tasks to the performance cores and schedule background tasks to the energy-efficient cores, realizing the reasonable allocation of processor resources and reducing energy consumption while ensuring performance. Threads with strong correlation can be scheduled to cores that share caches, reducing cache synchronization overhead and improving cache hit rates.

[0010] In conjunction with some embodiments of the first aspect, in some embodiments, detecting hardware changes by subscribing to device events of the system kernel specifically includes: Create an event listening thread for the system kernel and establish a communication channel with the kernel device management subsystem; Only subscribe to processor core hot-plug events, memory topology change events, and device status change events to obtain event filtering rules; Register event monitoring callback functions to the kernel according to event filtering rules; Receive a device event message sent by the system kernel, the device event message including an event type, a device identifier and a status flag; Extract the hardware components and status changes involved in the device event message; Get hardware changes based on event type and state changes.

[0011] By adopting the above technical solution, an event monitoring mechanism is used to capture the hardware change information of the system kernel in real time. Through precise event filtering rules, only key events such as processor core hot plugging, memory topology changes and device status changes are focused on. The established communication channel ensures that hardware change information can be transmitted to the container scheduling system in a timely manner, so that the system can be notified immediately when a processor core fails, is added or has performance degradation, and update the information in the processor core object pool accordingly. By adjusting the available processor core resources in real time, the system avoids scheduling containers to processor cores that have failed or have degraded performance, thereby improving the reliability of container operation. The event filtering mechanism reduces the system's overhead in processing hardware events, reduces resource waste caused by irrelevant events, and improves the real-time accuracy of the processor resource view.

[0012] In combination with some embodiments of the first aspect, in some embodiments, after binding each thread to a processor core with a corresponding number in the new container group through independent resource group configuration, the method further includes: Obtaining the instruction execution sequence of each thread in the new container group, and detecting the encryption operation instructions in the instruction execution sequence; Count the time proportion of encryption operation instructions, and mark the threads whose time proportion is greater than a preset threshold as encryption operation threads; Before the encryption operation thread is executed, the operating frequency of the processor core to which the encryption operation thread is bound is adjusted to the maximum frequency threshold supported by the processor; Collect the execution completion time of the encryption operation thread under the maximum frequency threshold, and use the execution completion time as the benchmark execution time; The estimated execution interval of the next encryption operation thread is determined according to the benchmark execution time, and the operating frequency of the processor core bound to the encryption operation thread is adjusted to the maximum frequency threshold in advance before the estimated execution interval.

[0013] By adopting the above technical solution, the execution efficiency of the encryption operation thread is identified and optimized by analyzing the instruction execution sequence of the thread. The system accurately identifies the encryption operation thread that requires high-performance support by counting the time proportion of the encryption operation instruction. After identifying the encryption operation thread, an accurate benchmark execution time is established by adjusting the operating frequency of its bound core to the maximum frequency threshold. Based on the execution time prediction mechanism, the system can adjust the processor core frequency in advance to ensure that the encryption operation thread obtains the best processor performance support during execution, reducing the energy waste caused by the processor continuously working in a high-frequency state, while ensuring that the encryption operation can obtain sufficient computing power when needed. Through the accurate identification and targeted optimization of the encryption operation thread, the execution efficiency of the encryption operation is improved, and the impact of encryption and decryption operations on application performance is reduced.

[0014] In conjunction with some embodiments of the first aspect, in some embodiments, determining the expected execution interval of the next encryption operation thread according to the benchmark execution time specifically includes: Record the execution time interval of N consecutive encryption operation threads, where N is an integer greater than 1; Calculate the periodic regularity indicator of the cryptographic operation based on the execution time interval; When the periodic regularity index is greater than the preset period threshold, the next expected execution interval is determined as a period after the last execution completion time; When the period regularity index is less than the preset period threshold, the next estimated execution interval is determined to be twice the benchmark execution time after the last execution completion time.

[0015] By adopting the above technical solution, when the periodic regularity index is greater than the preset periodic threshold, it indicates that the encryption operation presents a stable periodic feature. After setting the expected execution interval to one period, the next execution time can be accurately predicted. When the periodic regularity index is less than the preset periodic threshold, it indicates that the execution of the encryption operation is uncertain. Setting the expected execution interval to twice the benchmark execution time can reserve sufficient execution time for irregular encryption operations. Dynamically adjusting the expected execution interval according to the actual execution characteristics of the encryption operation not only reduces the waste of resources caused by reserving too much execution time for periodic encryption operations, but also ensures that non-periodic encryption operations have sufficient execution time, thereby improving the utilization efficiency of processor resources.

[0016] In conjunction with some embodiments of the first aspect, in some embodiments, before the operating frequency of the processor core bound to the encryption operation thread is adjusted to the maximum frequency threshold in advance before the expected execution interval, the method further includes: Obtain the estimated execution interval of the encryption operation threads of all container groups; Calculate the overlapping time periods of the estimated execution intervals; When it is detected that the number of encryption operation threads in the overlapping time period exceeds the preset concurrency threshold, the encryption operation threads are allocated to different time segments according to their priorities; According to the allocation result of the time slice, the operating frequency of the processor core bound to each encryption operation thread is adjusted in turn.

[0017] By adopting the above technical solution, the estimated execution intervals of all container group encryption operation threads are obtained and the overlapping time periods are calculated. The system can discover the situation where multiple encryption operation threads are executed concurrently in the same time period. When the number of encryption operation threads in the overlapping time period exceeds the preset concurrency threshold, allocating threads to different time segments according to priority can avoid too many encryption operation threads competing for processor resources at the same time. The operating frequency of the processor core is adjusted to a staggered level so that each encryption operation thread is executed in a different high-frequency time window, reducing the performance loss caused by resource competition, ensuring the execution efficiency of encryption operations, reducing the overall power consumption of the processor, and also avoiding the problem of a sharp increase in processor temperature due to intensive concurrent encryption operations.

[0018] In conjunction with some embodiments of the first aspect, in some embodiments, allocating encryption operation threads to different time segments according to priorities specifically includes: Obtain the service level identifier of the container group to which each encryption operation thread belongs; Prioritize the encryption operation threads according to the service level identifier to obtain a ranking result; Allocate a preset number of encryption operation threads ranked first to the original expected execution interval according to the sorting result; Allocate the encryption operation threads other than the preset number of the top ranked ones to the nearest idle time segment.

[0019] By adopting the above technical solution, the top-ranked encryption operation threads are arranged in the original expected execution interval, ensuring the timely response of high-priority services. The remaining encryption operation threads are scheduled to the nearest idle time segment, which can make full use of the idle computing resources of the processor. While ensuring the performance of key services, it also realizes the reasonable reuse of processor resources, improves the system's processing ability for encryption operation tasks of different priorities, and enables limited high-performance computing resources to give priority to important services while taking into account the execution needs of other services.

[0020] In the second aspect, an embodiment of the present application provides a CPU core binding system for a cloud-native application, wherein the CPU core binding system for a cloud-native application comprises: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code comprises computer instructions, and one or more processors call the computer instructions to enable the system to execute the method described in the first aspect and any possible implementation method of the first aspect.

[0021] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, comprising instructions, which, when executed on a system, causes the system to execute the method described in the first aspect and any possible implementation of the first aspect.

[0022] In a fourth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a system, the system executes the method described in any possible implementation manner in the first aspect.

[0023] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. The present application provides a CPU core binding method for cloud native applications, which reads core attributes from a processor information file and device information from a system device subsystem, establishes a processor core object pool, and updates the object pool information in real time based on hardware changes. When a container group is created, a suitable core set is selected from the object pool according to the processor demand configuration, and an independent resource group configuration is created for each thread. This technical solution enables the container to more accurately assign computing tasks to the most suitable processor core for execution, reducing memory access delays caused by cross-node access. At the same time, by managing the processor core objects in a pool, it avoids the situation where multiple containers compete for the same processor core, thereby improving the utilization efficiency of processor resources. Considering the memory architecture area identifiers of network cards and disk devices, the system can schedule network-intensive or storage-intensive containers to processor cores close to the corresponding devices, reducing the path length of data transmission and reducing the latency of I / O operations. The refined binding of cores at the thread level further ensures that the processor resources of key threads are not interfered with by other tasks, thereby improving the performance stability of the application.

[0024] 2. This application provides a CPU core binding method for cloud native applications, which identifies and optimizes the execution efficiency of encryption operation threads by analyzing the instruction execution sequence of the thread. The system accurately identifies the encryption operation threads that require high-performance support by counting the time proportion of encryption operation instructions. After identifying the encryption operation thread, an accurate benchmark execution time is established by adjusting the operating frequency of its bound core to the maximum frequency threshold. Based on the execution time prediction mechanism, the system can adjust the processor core frequency in advance to ensure that the encryption operation thread obtains the best processor performance support during execution, reducing the energy waste caused by the processor continuously working in a high-frequency state, while ensuring that the encryption operation can obtain sufficient computing power when needed. Through the accurate identification and targeted optimization of encryption operation threads, the execution efficiency of encryption operations is improved, and the impact of encryption and decryption operations on application performance is reduced.

[0025] 3. The present application provides a CPU core binding method for cloud native applications. It obtains the expected execution intervals of all container group encryption operation threads and calculates the overlapping time periods. The system can find that multiple encryption operation threads are executed concurrently in the same time period. When the number of encryption operation threads in the overlapping time period exceeds the preset concurrency threshold, allocating threads to different time segments according to priority can avoid too many encryption operation threads competing for processor resources at the same time. The operating frequency of the processor core is adjusted to a staggered frequency so that each encryption operation thread is executed in a different high-frequency time window, reducing the performance loss caused by resource competition, ensuring the execution efficiency of encryption operations, reducing the overall power consumption of the processor, and also avoiding the problem of a sharp increase in processor temperature due to intensive concurrent encryption operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flow chart of a CPU core binding method for a cloud native application in an embodiment of the present application.

[0027] Figure 2 It is a flowchart of a method for optimizing processor frequency for a specific computing task in an embodiment of the present application.

[0028] Figure 3 This is a schematic diagram of the physical device structure of a CPU core binding system for a cloud native application provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to be used as limitations to the present application. As used in the specification and appended claims of the present application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include plural expressions, unless there is a clear indication to the contrary in the context. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations comprising one or more listed items.

[0030] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more.

[0031] The following uses an embodiment and combines Figure 1 , a CPU core binding method for a cloud native application in an embodiment of the present application is described: See also Figure 1, which is a flow chart of a CPU core binding method for a cloud native application in an embodiment of the present application.

[0032] S101, reading the memory architecture region identifier and core type attribute of the allocatable processor core from the processor information file; The system reads the memory architecture region identifier and core type attributes of the allocatable processor core from the processor information file, specifically: obtaining the physical identification number and physical location information of all processor cores in the system; reading the cache hierarchy structure information of each processor core according to the physical identification number, the cache hierarchy structure information including the first-level cache size, the second-level cache size and the third-level cache sharing relationship; constructing a physical topology relationship graph of the processor core based on the physical location information and the cache hierarchy structure information; dividing the processor cores with the same third-level cache sharing relationship into the same memory architecture region according to the physical topology relationship graph, and obtaining the memory architecture region identifier; reading the static operating parameters of each processor core, the static operating parameters including the maximum operating frequency, the minimum operating frequency and the peak power consumption; determining the processor cores with a maximum operating frequency greater than a preset frequency threshold and a peak power consumption greater than a preset power consumption threshold as performance cores, and determining the remaining processor cores as energy efficiency cores according to the static operating parameters, and obtaining the core type attributes.

[0033] In this step, the system reads the memory architecture region identifier and core type attributes of the allocable processor core from the processor information file. This step aims to obtain detailed information about the processor core in the system and provide a basis for subsequent resource allocation and management. The system can obtain processor information in various ways, such as directly accessing hardware registers and calling API interfaces provided by the operating system.

[0034] Specifically, the system first obtains the physical identification number and physical location information of all processor cores in the system. The physical identification number can be a unique identification of the processor core, such as APIC ID, etc.; the physical location information can include the processor socket number, the location of the core in the physical package, etc. Then, the system reads the cache hierarchy structure information of each processor core according to the physical identification number, including the first-level cache size, the second-level cache size, and the third-level cache sharing relationship. This information can be obtained from the CPUID instruction return value of the processor. Next, the system builds a physical topology relationship diagram of the processor core based on the physical location information and the cache hierarchy structure information to reflect the physical connection and cache sharing relationship between the cores. On this basis, the system divides the processor cores with the same third-level cache sharing relationship into the same memory architecture area according to the physical topology relationship diagram, and obtains the memory architecture area identification. Finally, the system reads the static operating parameters of each processor core, including the maximum operating frequency, the minimum operating frequency, and the peak power consumption, and divides the processor core into a performance core and an energy efficiency core according to the preset frequency and power consumption thresholds to obtain the core type attributes.

[0035] S102, reading memory architecture area identifiers of network cards and disk devices from the system device subsystem; In this step, the system reads the memory architecture region identifiers of the network card and disk device from the device subsystem. The purpose of this step is to understand the physical location relationship between the I / O device and the processor core to provide support for NUMA-aware resource scheduling.

[0036] In specific implementation, the system can call the device management interface provided by the operating system kernel, such as the Linux sysfs file system, to obtain the NUMA Node information to which each I / O device is connected. For PCIe devices, the system can also obtain the NUMA Node ID by parsing the configuration space register of the device. This information can be mapped with the memory architecture area identifier of the processor core to establish the corresponding relationship between the I / O device and the processor core.

[0037] In some scenarios, I / O devices may be connected to the processor through components such as PCIe switches, making the physical location relationship between the device and the processor complicated. In this case, the system can further analyze the PCIe topology to identify the processor socket and memory controller to which the device is actually connected, thereby correctly determining its NUMA affiliation. In addition, for old devices that do not support NUMA awareness, the system can classify them into the default memory architecture area based on the bus number and device number to which they are connected.

[0038] S103, generating a processor core object according to the memory architecture region identifier of the processor core, the core type attribute, and the memory architecture region identifiers of the network card and the disk device; In this step, the system generates processor core objects for management and allocation based on the previously acquired attribute information of processor cores and I / O devices. This step aims to aggregate and abstract the scattered hardware resource information to form a unified resource view to facilitate subsequent resource management operations.

[0039] In specific implementation, the system can define a data structure of a processor core object, including fields such as memory architecture region identifier, core type attribute, and associated local I / O device list. Then, the system traverses each processor core and fills in the corresponding fields according to its memory architecture region identifier and core type attribute. At the same time, the system also needs to traverse each network card and disk device, associate its memory architecture region identifier with the processor core object, and build a complete resource topology structure.

[0040] In the process of generating processor core objects, the system may need to handle some special cases. For example, in some asymmetric architecture processors, different cores may have different instruction set extensions or functional units, which makes them significantly different in performance and power consumption characteristics. In this regard, when generating core objects, the system can introduce more fine-grained attribute descriptions, such as instruction set architecture version, special function flags, etc., to fully characterize the heterogeneous characteristics of the core. In addition, considering the diversity of node hardware configurations in cloud environments, the system can also support dynamic expansion and customization of processor core objects, allowing users to add custom attribute fields as needed to adapt to different application scenarios.

[0041] S104, storing the processor core object in the container cluster in the form of a custom resource to obtain a processor core object pool; In this step, the system stores the generated processor core objects in the container cluster as custom resources to build a global processor core object pool. The purpose of this is to unify the hardware resource information with the resource model of the container cluster so that the container orchestration system can perceive and schedule the underlying CPU resources.

[0042] In specific implementation, the system can use the Custom Resource Definition (CRD) mechanism provided by Kubernetes to define a new resource type to describe processor core objects. For example, you can define a CRD named "ProcessorCore" that contains fields such as memory architecture region identifier and core type attributes. Then, the system creates each processor core object as a custom resource instance of type "ProcessorCore" and stores it in the etcd database of Kubernetes. These processor core resource instances logically form a cluster-wide object pool for query and use by the container orchestration system.

[0043] In the process of storing processor core objects, the system needs to pay attention to the version management and upgrade issues of custom resources. Since the hardware configuration may change at any time, the definition of the processor core object may also need to be adjusted accordingly. To this end, the system can define multiple versions for the "ProcessorCore" resource type through the version control mechanism of CRD, and support automatic conversion and upgrade between versions. In addition, considering the expansion of the cluster scale, the access and management of the processor core object pool may also face performance challenges. To this end, the system can explore caching and sharding-based mechanisms to improve the query efficiency and scalability of the object pool.

[0044] S105, detecting hardware changes by subscribing to device events of the system kernel; The system detects hardware changes by subscribing to device events of the system kernel, specifically: creating an event listening thread of the system kernel and establishing a communication channel with the kernel device management subsystem; subscribing only to processor core hot-plug events, memory topology change events and device status change events to obtain event filtering rules; registering event listening callback functions with the kernel according to the event filtering rules; receiving device event messages sent by the system kernel, which include event type, device identification and status flags; extracting the hardware components and status changes involved in the device event messages; and obtaining hardware changes according to the event type and status changes.

[0045] In this step, the system monitors the changes of hardware configuration in real time by subscribing to the kernel's device events. The purpose of this step is to ensure that the processor core object pool can dynamically adapt to changes in the underlying hardware and avoid inconsistent resource information.

[0046] In the specific implementation, the system first needs to create a dedicated kernel event listening thread and establish a communication channel with the kernel's device management subsystem. Then, according to the predefined event filtering rules, the system selectively subscribes to events related to the processor core and memory topology, such as processor core hot plug events, memory topology change events, and device status change events. These event filtering rules can be set and adjusted through configuration files or API interfaces. Next, the system registers the event listening callback function with the kernel so that it can be notified in time when related events occur. When the listening thread receives the device event message sent by the kernel, the system extracts the affected hardware components and status change information from it, and then determines whether a hardware configuration change has occurred.

[0047] In the process of monitoring hardware changes, the system may need to handle abnormal situations such as event loss and disordered order. For example, in a high-concurrency scenario, the device event messages generated by the kernel may exceed the processing capacity of the listening thread, resulting in some events being discarded. In this regard, the system can set an appropriate buffer at the source of the kernel event, or use multiple listening threads to process in parallel to improve the reliability and real-time performance of event processing. In addition, since there may be dependencies between hardware change events, such as changes in cache topology causing changes in the properties of the processor core, the system needs to consider the causal order of events when analyzing the impact of events to ensure the correctness of processing.

[0048] S106, updating the processor core object pool according to the hardware change; In this step, the system updates the affected objects in the processor core object pool according to the detected hardware changes. This step is to keep the resource view synchronized with the physical hardware status and ensure that subsequent container scheduling decisions are based on the latest hardware information.

[0049] In specific implementation, the system first needs to locate the affected objects in the processor core object pool based on the device identifier in the hardware change event. For example, if the event indicates that a processor core is hot-plugged, the system needs to find the "ProcessorCore" resource instance corresponding to the core in the object pool. Next, the system updates the attribute fields of the relevant processor core objects based on the hardware status changes reported in the event. For example, if a core changes from a performance core to an energy-efficiency core, the system needs to modify the "CoreType" attribute value of the corresponding object. After completing the attribute update, the system writes the modified processor core object back to the etcd database of Kubernetes to ensure that other components in the cluster can see the latest resource status.

[0050] In the process of updating the processor core object, the system needs to pay attention to the issues of concurrent modification and data consistency. In a cloud environment, multiple hardware events may occur at the same time, resulting in concurrent updates to the same processor core object. To this end, the system can use the resource version control and optimistic locking mechanism provided by Kubernetes to detect and handle concurrent modifications by comparing the version numbers of objects. In addition, since the scheduler and monitoring components of Kubernetes also access the processor core object pool, the system also needs to consider data visibility and consistency when updating objects. To this end, the system can use the Watch mechanism of etcd to actively notify related components when objects change to reduce the inconsistent window period; or use a two-phase commit method to encapsulate the object modification operation into a transaction to ensure atomicity.

[0051] S107, receiving a container group creation request, and obtaining a processor requirement configuration, where the processor requirement configuration includes the number of cores, core type, memory architecture area requirements, and thread binding relationships; In this step, the system receives container group (Pod) creation requests from users or upper-layer systems and extracts the required configuration information related to processor resources from the requests. The purpose of this step is to clarify the specific requirements of the container group for processor resources and provide a basis for subsequent resource allocation decisions.

[0052] In specific implementation, the system can define a new container group annotation or label to describe the container group's demand for processor resources. For example, you can define an annotation named "processor-request", whose value is a string in JSON format, containing fields such as the number of processor cores, core type, memory architecture area, and thread binding relationship. When a user creates a container group, he can add this annotation to the metadata of the container group to clarify its demand for processor resources. After receiving a container group creation request, the system can parse the value of the "processor-request" annotation from the metadata of the request object to obtain the processor demand configuration.

[0053] In the process of obtaining the processor requirement configuration, the system needs to verify the legitimacy of the configuration information provided by the user. For example, the number of cores cannot exceed the total number of cores of the node, and the identifier of the memory architecture area must actually exist in the cluster. To this end, the system can predefine a set of configuration specifications and constraints, and check them item by item when parsing the configuration. If the configuration information is found to be illegal, the system can directly reject the request to create the container group and return the corresponding error message to the user. In addition, considering the allocation granularity and restrictions of processor resources, the system can also appropriately align and adjust the number of cores and memory areas configured by the user to better adapt to the characteristics of the underlying hardware.

[0054] S108, selecting a processor core set that meets the processor requirement configuration from the updated processor core object pool, and using the processor core set as a container group; In this step, the system selects core resources that meet the requirements from the processor core object pool according to the processor requirement configuration of the container group, and binds the selected core set to the container group. This step aims to achieve accurate matching and isolation of container groups to physical cores, and improve the performance and reliability of container applications.

[0055] In the specific implementation, the system first needs to build a processor core selection algorithm based on the required configuration of the container group. The algorithm needs to comprehensively consider multiple factors such as the number of cores, core types, memory areas, and NUMA locality, and find the optimal core combination that meets the conditions from the processor core object pool. Among them, matching the core type can ensure that the container group obtains the required computing power, matching the memory area can minimize remote memory access, and matching NUMA locality can improve cache hit rate and memory bandwidth utilization. After finding a processor core set that meets the conditions, the system binds it to the requested container group and records the allocated core resources in the status information of the container group.

[0056] S109, removing the selected processor core set from the updated processor core object pool to obtain a new container group; In this step, the system removes the processor core set assigned to the container group from the core object pool to ensure exclusive use of these core resources. The purpose of this step is to achieve strong isolation between the container group and the host core to avoid interference between different containers.

[0057] In specific implementation, after the system completes the binding of core resources to container groups, it needs to mark the corresponding "ProcessorCore" resource instance from the processor core object pool as allocated. This can be achieved by updating the status field of the resource object, such as setting the "Allocated" field to true. At the same time, the system also needs to record the information of the container group bound to it in these resource instances, such as the name and namespace of the container group, for subsequent status query and unbinding operations. After the marking is completed, the system writes the updated "ProcessorCore" resource object back to the etcd database of Kubernetes. In this way, other container groups will no longer select these allocated core objects when applying for processor resources, ensuring the exclusivity of core resources.

[0058] S110, creating an independent resource group configuration for each thread in the new container group; In this step, the system creates an independent resource group (cgroup) configuration for each thread in the newly created container group. This step is intended to achieve finer-grained resource isolation and control to prevent threads within the container from competing for resources.

[0059] In the specific implementation, the system first needs to determine the thread range within the container group where the resource group needs to be created. This can be done based on the metadata information of the container image, such as Entrypoint, CMD, etc., to infer the thread structure of the main process in the container. In addition, the system can also provide users with corresponding annotations or tags, allowing users to explicitly specify the threads that need to be isolated. Next, the system generates corresponding cgroup parameters for each thread based on the binding relationship between threads and cores specified in the processor requirement configuration, as well as the resource constraints of each core. These parameters can include CPU time quotas, memory limits, cpusets, etc. After the generation is complete, the system calls cgroup-related APIs, such as libcgroup or systemd, to create corresponding cgroup hierarchies and subsystems on the host. Finally, the system adds the main process of the container group and its child threads to these cgroups to achieve refined resource management and control.

[0060] S111 . According to the correspondence between threads and processor core numbers in the processor requirement configuration, each thread is bound to a processor core with a corresponding number in the new container group through an independent resource group configuration.

[0061] In this step, the system binds each thread to the corresponding numbered processor core in the new container group through independent resource group configuration according to the correspondence between threads and processor cores specified in the processor requirement configuration. The purpose of this step is to further enhance the affinity between container applications and hardware and achieve predictable high performance.

[0062] In specific implementation, when the system starts the container group, in addition to the conventional environment variables, volume mounts and other configurations, it is also necessary to set the processor affinity for the main process of the container. This can be achieved by modifying the startup parameters of the container, such as adding the "--cpuset-cpus" and "--cpuset-mems" options in the startup command of the Docker container, or setting the "nodeSelector" and "affinity" fields in the definition of the KubernetesPod. The values ​​of these parameters can be extracted from the processor demand configuration and mapped to the physical core number and NUMA node assigned by the system. After the container is started, its main process and child threads will automatically inherit these processor affinity configurations to achieve static binding of threads and cores.

[0063] In the process of implementing thread binding, the system needs to consider the differences between container runtimes and processor topologies. Different container runtimes, such as Docker, containerd, rkt, etc., may have different interfaces and semantics for processor affinity configuration. Therefore, the system needs to provide an adaptation layer to map the processor requirement configuration to the configuration parameters of different runtimes. In addition, under the NUMA architecture, threads can obtain lower memory access latency when running on cores within the local node. However, in some scenarios, threads may need to access the memory of remote nodes, such as global synchronization operations across slots. In this regard, when generating the binding configuration, the system needs to take into account data locality and load balancing, and appropriately distribute the threads to different NUMA nodes. At the same time, the system also needs to provide a certain degree of flexibility, allowing users to manually adjust or cancel the thread binding settings for specific application scenarios.

[0064] In the above embodiment, core attributes are read from the processor information file and device information is read from the system device subsystem, a processor core object pool is established, and the object pool information is updated in real time based on hardware changes. When a container group is created, a suitable core set is selected from the object pool according to the processor demand configuration, and an independent resource group configuration is created for each thread. This technical solution enables the container to more accurately assign computing tasks to the most suitable processor core for execution, reducing memory access delays caused by cross-node access. At the same time, by managing the processor core objects in a pool, the situation where multiple containers compete for the same processor core is avoided, and the utilization efficiency of processor resources is improved. Considering the memory architecture area identification of network cards and disk devices, the system can schedule network-intensive or storage-intensive containers to processor cores close to the corresponding devices, reducing the path length of data transmission and reducing the delay of I / O operations. The refined binding of cores at the thread level further ensures that the processor resources of key threads are not interfered with by other tasks, thereby improving the performance stability of the application.

[0065] On the basis of the above embodiments, in addition to improving container performance through processor core object pools and refined thread core binding mechanisms, this application also provides a processor frequency optimization method for specific computing tasks. Taking into account the common encryption computing requirements in cloud-native applications, the system realizes intelligent adjustment of the processor core operating frequency by identifying and analyzing the execution characteristics of encryption computing threads. This targeted frequency optimization method can not only improve the execution efficiency of encryption computing, but also achieve reasonable control of energy consumption while ensuring performance. The following is combined with Figure 2 , a processor frequency optimization method for a specific computing task in an embodiment of the present application is described: See also Figure 2 , which is a flow chart of a method for optimizing processor frequency for a specific computing task in an embodiment of the present application.

[0066] S201, obtaining the instruction execution sequence of each thread in the new container group, and detecting the encryption operation instructions in the instruction execution sequence; In this step, the system first obtains the instruction execution sequence of each thread in the newly created container group. This step is to provide the necessary data basis for subsequent encryption operation detection and analysis. The instruction execution sequence reflects the actual running track of the thread and records the type, address and other information of each instruction during the thread execution process.

[0067] In specific implementation, the system can use dynamic instrumentation or static analysis techniques to obtain the instruction sequence of the thread. Among them, dynamic instrumentation is a method of recording the execution of instructions in real time when the program is running. The system can use tools such as ftrace or perf provided by the Linux kernel to insert probes at the system call or function entry of the container to capture and record the instruction flow information of each thread. The advantage of this method is that it can obtain a detailed and accurate instruction sequence, but it may introduce certain runtime overhead. Another static analysis method is to disassemble or parse the symbol table of the executable file in the container when the container image is built or deployed, and extract the instruction flow information in the code. The advantage of this method is that there is no need to modify the program and the overhead is small, but the accuracy of the analysis may be affected by factors such as code obfuscation.

[0068] S202, counting the time proportion of encryption operation instructions, and marking threads whose time proportion is greater than a preset threshold as encryption operation threads; In this step, after detecting the encryption operation instructions of the thread, the system further counts the execution time proportion of these instructions. The purpose of this step is to identify the threads that are mainly used for encryption calculations and provide targeted objects for subsequent frequency tuning.

[0069] In specific implementation, the system can calculate the execution time of each encryption instruction based on the timestamp information in the instruction sequence, and accumulate the total encryption operation time. At the same time, the system also needs to count the total execution time of the thread, which can be obtained by subtracting the timestamp of the first instruction from the timestamp of the last instruction. Dividing the two, the time proportion of the encryption operation instruction can be obtained. In order to distinguish between encryption operation threads and ordinary threads, the system can set a time proportion threshold, such as 50%. When the encryption instruction time proportion of a thread exceeds this threshold, it is marked as an encryption operation thread.

[0070] S203, before the encryption operation thread is executed, adjusting the operating frequency of the processor core to which the encryption operation thread is bound to a maximum frequency threshold supported by the processor; In this step, the system performs frequency tuning on the identified encryption operation thread. The purpose of this step is to speed up the execution of encryption instructions by increasing the operating frequency of the processor core, thereby shortening the completion time of the encryption task.

[0071] In the specific implementation, the system first needs to intercept the encryption operation thread before it is executed. This can be achieved by setting a hook function in the thread scheduler, that is, triggering a callback function before the thread is scheduled to run on the CPU. In the callback function, the system can check the properties of the thread to determine whether it is an encryption operation thread. If it is, it continues to perform the frequency adjustment operation; if not, it returns directly without doing anything. Next, the system needs to determine the processor core to which the encryption operation thread is bound. This can be obtained by reading the CPU affinity mask of the thread. According to the bit number set in the mask, the system can know which core the thread is currently running on. After obtaining the target core, the system sets the operating frequency of the core to the maximum supported threshold by calling the processor's DVFS (dynamic voltage frequency adjustment) interface, such as Intel's P-State or AMD's Cool'n'Quiet technology. This threshold can be obtained from the processor's specification document or queried in real time through the CPUID instruction.

[0072] In the process of adjusting the processor core frequency, the system needs to balance performance and power consumption. On the one hand, the higher the frequency, the faster the encryption operation is executed, but it also means higher power consumption and heat generation. Therefore, when setting the maximum frequency threshold, the system needs to comprehensively consider factors such as the processor's heat dissipation capacity and power budget. In addition, blindly adjusting the frequency to the highest point is not always the best strategy. This is because increasing the frequency is usually accompanied by an increase in voltage, and excessive voltage may cause stability problems, leading to system crashes or data loss. To this end, the system can find an optimal frequency point that takes into account both performance and stability by testing different frequency-voltage combinations. In addition, the system can also dynamically adjust the frequency threshold according to the urgency of the encryption task. For example, for some encryption tasks with high real-time requirements, such as financial transactions and security authentication, the system can appropriately increase the threshold to ensure performance; while for some offline batch tasks, such as data backup and log encryption, the system can appropriately lower the threshold to save energy.

[0073] S204, collecting the execution completion time of the encryption operation thread under the maximum frequency threshold, and using the execution completion time as the benchmark execution time; In this step, after the system adjusts the core frequency of the encryption operation thread to the maximum threshold, it collects the time it takes to complete the execution and uses this time as the benchmark value for subsequent optimization. The purpose of this step is to establish a reference standard for measuring encryption operation performance and provide feedback for dynamic frequency adjustment.

[0074] In specific implementation, the system can record the current timestamp immediately after the execution of the encryption operation thread is completed. Since the execution time of the thread may be very short, the accuracy of the timestamp needs to be high enough, at least at the microsecond level. In order to obtain more reliable statistical results, the system can measure the execution time of the thread multiple times in a row and calculate its average value. Considering that there may be multiple encryption operation threads in the system, the calculation amount and execution path of each thread may be different, so the system needs to record the benchmark execution time for each thread separately. In addition, since the operating frequency of the processor may be affected by factors such as power consumption management and temperature control, the actual operating frequency is lower than the set maximum threshold. In order to eliminate the interference of these factors, the system can collect the actual operating frequency of the processor core while recording the benchmark execution time as a reference for subsequent analysis.

[0075] S205. Determine an estimated execution interval of the next encryption operation thread according to the benchmark execution time, and adjust the operating frequency of the processor core bound to the encryption operation thread to a maximum frequency threshold in advance before the estimated execution interval.

[0076] The system determines the estimated execution interval of the next encryption operation thread according to the benchmark execution time, specifically including: recording the execution time intervals of N consecutive encryption operation threads, where N is an integer greater than 1; calculating the periodic regularity index of the encryption operation based on the execution time interval; when the periodic regularity index is greater than the preset period threshold, determining the next estimated execution interval as a period after the last execution completion time; when the periodic regularity index is less than the preset period threshold, determining the next estimated execution interval as twice the benchmark execution time after the last execution completion time. Before the estimated execution interval, the operating frequency of the processor core bound to the encryption operation thread is adjusted to the maximum frequency threshold in advance.

[0077] In this step, the system predicts the next execution time interval of the encryption operation thread based on the benchmark execution time obtained in the previous step, and adjusts the processor core frequency bound to the thread to the maximum threshold in advance before the interval arrives. The purpose of this step is to reduce the running time of the encryption operation thread in the low-frequency state by predicting and adjusting the frequency in advance, and further improve the performance of encryption calculations.

[0078] In specific implementation, the system first needs to analyze the execution cycle of the encryption operation thread based on historical execution information. To this end, the system can record the time intervals between multiple consecutive executions of the thread and calculate statistical indicators such as its mean and variance. If the variance is small, it means that the execution of the thread has a strong periodicity, and a fixed time interval can be used to predict the time point of the next execution. If the variance is large, it means that the execution of the thread is relatively random and it is difficult to make an accurate prediction. In this case, the system can refer to the benchmark execution time and estimate a relatively loose time interval. After obtaining the expected execution interval, the system needs to adjust the frequency of the processor core to the maximum threshold in advance a certain amount of time before the arrival of the interval. The length of this advance depends on the time consumption of the frequency adjustment itself. Generally speaking, the frequency adjustment takes several milliseconds to tens of milliseconds, so the advance should also be in this order of magnitude. In order to avoid performance loss caused by frequent frequency adjustment, the system can also introduce a minimum adjustment interval limit, that is, there must be at least a certain time interval between two adjustments.

[0079] In the above embodiment, the execution efficiency of the encryption operation thread is identified and optimized by analyzing the instruction execution sequence of the thread. The system accurately identifies the encryption operation thread that requires high-performance support by counting the time proportion of the encryption operation instruction. After identifying the encryption operation thread, an accurate benchmark execution time is established by adjusting the operating frequency of its bound core to the maximum frequency threshold. Based on the execution time prediction mechanism, the system can adjust the processor core frequency in advance to ensure that the encryption operation thread obtains the best processor performance support during execution, reducing the energy waste caused by the processor continuously working in a high-frequency state, while ensuring that the encryption operation can obtain sufficient computing power when needed. Through the accurate identification and targeted optimization of the encryption operation thread, the execution efficiency of the encryption operation is improved, and the impact of the encryption and decryption operations on the application performance is reduced.

[0080] Furthermore, in another embodiment, before the operating frequency of the processor core bound to the encryption operation thread is adjusted to the maximum frequency threshold in advance before the expected execution interval, it also includes: Obtain the estimated execution interval of the encryption operation threads of all container groups; Calculate the overlapping time periods of the estimated execution intervals; When it is detected that the number of encryption operation threads in the overlapping time period exceeds the preset concurrency threshold, the encryption operation threads are allocated to different time segments according to the priority. Specifically: the service level identifier of the container group to which each encryption operation thread belongs is obtained; Prioritize the encryption operation threads according to the service level identifier to obtain a ranking result; Allocate a preset number of encryption operation threads ranked first to the original expected execution interval according to the sorting result; Allocate the encryption operation threads other than the preset number of threads ranked first to the nearest idle time segment; According to the allocation result of the time slice, the operating frequency of the processor core bound to each encryption operation thread is adjusted in turn.

[0081] In this embodiment, the system further optimizes the frequency adjustment strategy of the encryption operation thread to adapt to the scenario where multiple container groups concurrently execute encryption tasks. The purpose of this optimization is to ensure the performance of high-priority tasks while avoiding resource competition caused by too many threads adjusting the frequency at the same time.

[0082] In specific implementation, the system first needs to obtain the expected execution intervals of the encryption operation threads in all container groups. This can be achieved by deploying a monitoring agent in each container group. The agent is responsible for collecting the execution information of the encryption threads in this group and reporting it to the global scheduler regularly. The scheduler aggregates the data from each container group to form a global encryption task execution view. Next, the system performs an overlap analysis on the expected execution intervals of all threads to find the set of threads with time conflicts. If the number of encryption threads exceeds the preset concurrency threshold (such as half the number of processor cores) during a certain overlapping time period, it means that there is a risk of resource competition and the thread execution time needs to be sliced ​​and scheduled.

[0083] After determining that shard scheduling is required, the system further prioritizes the threads according to the service level of the container group to which they belong. The service level can be specified by the user when the container group is created, or automatically divided according to the importance of the business carried by the container group. The result of the priority sorting determines the resource allocation order of the threads in the shard scheduling. In order to take into account the timeliness of the task, the system allocates the top N high-priority threads to the original expected execution interval to ensure that they can obtain processor resources within the expected time. The remaining low-priority threads are assigned to the nearest idle time segments. If all time segments are occupied, the system can also delay the execution of some low-priority threads to the next scheduling cycle according to their priority.

[0084] According to the allocation results of the time slice, the system finally adjusts the frequency of the processor core bound to each thread. For high-priority threads assigned to the original execution interval, the system only needs to increase the frequency of the corresponding core to the maximum threshold before its expected start time. For low-priority threads assigned to other time slices, the system needs to dynamically adjust the core frequency before the thread is actually executed. Since the frequency adjustment itself also has a certain time overhead, the system needs to reserve a certain buffer for frequency adjustment when specifying the start time of the time slice. In addition, if multiple threads are assigned to the same time slice, the system also needs to control the adjustment order of the core frequency to reduce the performance loss caused by frequent frequency adjustment. For example, the time slice can be further divided into multiple sub-segments, each sub-segment corresponds to a thread, and the frequency is adjusted in sequence according to the thread priority.

[0085] In the above embodiment, the expected execution intervals of all container group encryption operation threads are obtained and the overlapping time periods are calculated. The system can find that multiple encryption operation threads are executed concurrently in the same time period. When the number of encryption operation threads in the overlapping time period exceeds the preset concurrency threshold, allocating threads to different time segments according to priority can avoid too many encryption operation threads competing for processor resources at the same time. The operating frequency of the processor core is adjusted to a staggered level so that each encryption operation thread is executed in a different high-frequency time window, reducing the performance loss caused by resource competition, ensuring the execution efficiency of encryption operations, reducing the overall power consumption of the processor, and also avoiding the problem of a sharp increase in processor temperature due to intensive concurrent encryption operations.

[0086] The following describes the system in the embodiment of the present invention from the perspective of hardware processing. Figure 3 , which is a schematic diagram of the physical device structure of a CPU core binding system for a cloud native application provided in an embodiment of the present application.

[0087] It should be noted that Figure 3 The structure of the system shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0088] like Figure 3As shown, the system includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage part 308 to the random access memory (RAM) 303, such as executing the method in the above embodiment. In RAM 303, various programs and data required for system operation are also stored. CPU 301, ROM 302 and RAM 303 are connected to each other through a bus 304. Input / output (I / O) interface 305 is also connected to bus 304.

[0089] The following components are connected to the I / O interface 305: an input section 306 including a camera, an infrared sensor, etc.; an output section 307 including a liquid crystal display (LCD) and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read therefrom is installed into the storage section 308 as needed.

[0090] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 309, and / or installed from a removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, various functions defined in the present invention are performed.

[0091] It should be noted that the computer-readable medium shown in the embodiment of the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing.

[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. Among them, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0093] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the system described in the above embodiment; or may exist independently without being assembled into the system. The above storage medium carries one or more computer programs, and when the above one or more computer programs are executed by a processor of a system, the system implements the method provided in the above embodiment.

[0094] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0095] As used in the above embodiments, the term "when..." may be interpreted as "if..." or "after..." or "in response to determining..." or "in response to detecting...", depending on the context. Similarly, the phrases "upon determining..." or "if (the stated condition or event) is detected" may be interpreted as "if determining..." or "in response to determining..." or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)", depending on the context.

[0096] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk), etc.

[0097] Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiments, the processes can be completed by computer programs to instruct related hardware, and the programs can be stored in computer-readable storage media. When the programs are executed, they can include the processes of the above-mentioned method embodiments. The aforementioned storage media include: ROM or random access memory RAM, magnetic disk or optical disk and other media that can store program codes.

Claims

1. A CPU core binding method for cloud native applications, characterized in that: include: Read the memory architecture region identifier and core type attribute of the allocatable processor core from the processor information file; Read the memory architecture region identifiers of network cards and disk devices from the system device subsystem; Generate a processor core object according to the memory architecture region identifier of the processor core, the core type attribute, and the memory architecture region identifiers of the network card and the disk device; Storing the processor core object in a container cluster in the form of a custom resource to obtain a processor core object pool; Detect hardware changes by subscribing to device events in the system kernel; updating the processor core object pool according to the hardware change; Receive a container group creation request and obtain a processor requirement configuration, wherein the processor requirement configuration includes the number of cores, core type, memory architecture area requirements, and thread binding relationships; Selecting a processor core set that meets the processor requirement configuration from the updated processor core object pool, and using the processor core set as the container group; removing the selected processor core set from the updated processor core object pool to obtain a new container group; Creating an independent resource group configuration for each thread in the new container group; According to the correspondence between threads and processor core numbers in the processor requirement configuration, each of the threads is bound to a processor core with a corresponding number in the new container group through the independent resource group configuration.

2. The method according to claim 1, characterized in that The step of reading the memory architecture region identifier and core type attribute of the allocatable processor core from the processor information file specifically includes: Get the physical identification number and physical location information of all processor cores in the system; Reading cache hierarchy structure information of each processor core according to the physical identification number, the cache hierarchy structure information including a first-level cache size, a second-level cache size, and a third-level cache sharing relationship; Constructing a physical topology relationship diagram of the processor core based on the physical location information and the cache hierarchy information; According to the physical topology relationship diagram, the processor cores having the same three-level cache sharing relationship are divided into the same memory architecture region to obtain a memory architecture region identifier; Reading static operating parameters of each of the processor cores, the static operating parameters including maximum operating frequency, minimum operating frequency and peak power consumption; According to the static operating parameters, the processor core whose maximum operating frequency is greater than a preset frequency threshold and whose peak power consumption is greater than a preset power consumption threshold is determined as a performance core, and the remaining processor cores are determined as energy efficiency cores to obtain core type attributes.

3. The method according to claim 1, characterized in that The detection of hardware changes by subscribing to device events of the system kernel specifically includes: Create an event listening thread for the system kernel and establish a communication channel with the kernel device management subsystem; Only subscribe to processor core hot-plug events, memory topology change events, and device status change events to obtain event filtering rules; Register an event monitoring callback function with the kernel according to the event filtering rule; Receiving a device event message sent by the system kernel, wherein the device event message includes an event type, a device identifier, and a status flag; Extracting hardware components and state changes involved in the device event message; A hardware change is obtained according to the event type and the state change.

4. The method according to claim 1, characterized in that: After binding each of the threads to the processor cores with corresponding numbers in the new container group through the independent resource group configuration, the method further includes: Obtaining the instruction execution sequence of each of the threads in the new container group, and detecting the encryption operation instructions in the instruction execution sequence; Counting the time proportion of the encryption operation instruction, and marking the thread whose time proportion is greater than a preset threshold as an encryption operation thread; Before the encryption operation thread is executed, adjusting the operating frequency of the processor core to which the encryption operation thread is bound to a maximum frequency threshold supported by the processor; Collecting the execution completion time of the encryption operation thread under the maximum frequency threshold, and using the execution completion time as the benchmark execution time; The estimated execution interval of the encryption operation thread next time is determined according to the benchmark execution time, and the operating frequency of the processor core bound to the encryption operation thread is adjusted to the maximum frequency threshold in advance before the estimated execution interval.

5. The method according to claim 4, characterized in that Determining the estimated execution interval of the encryption operation thread next time according to the benchmark execution time specifically includes: Recording N consecutive execution time intervals of the encryption operation thread, where N is an integer greater than 1; Calculating a periodic regularity index of the encryption operation based on the execution time interval; When the periodic regularity index is greater than a preset period threshold, the next expected execution interval is determined to be a period after the last execution completion time; When the periodic regularity index is less than the preset periodic threshold, the next estimated execution interval is determined to be twice the reference execution time after the last execution completion time.

6. The method according to claim 4 or 5, characterized in that: Before the estimated execution interval, the operating frequency of the processor core bound to the encryption operation thread is adjusted to the maximum frequency threshold in advance, the method further includes: Obtain the estimated execution interval of the encryption operation threads of all container groups; Calculating the overlapping time periods of the estimated execution intervals; When it is detected that the number of encryption operation threads in the overlapping time period exceeds a preset concurrency threshold, the encryption operation threads are allocated to different time segments according to priority; According to the allocation result of the time segments, the operating frequency of the processor core bound to each of the encryption operation threads is adjusted in sequence.

7. The method according to claim 6, characterized in that The allocating the encryption operation threads to different time segments according to priorities specifically includes: Obtaining a service level identifier of the container group to which each of the encryption operation threads belongs; Prioritize the encryption operation threads according to the service level identifier to obtain a ranking result; Allocating a preset number of encryption operation threads ranked first to the original expected execution interval according to the sorting result; Allocate the encryption operation threads other than the preset number of the top ranked ones to the nearest idle time segment.

8. A CPU core binding system for cloud native applications, characterized in that: The system comprises: One or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the system to execute the method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on a system, the system is caused to execute the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that When the computer program product is run on a system, the system is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Low-power-consumption lightweight virtualization method for edge computing

    CN112162826A

  • Heterogeneous hardware-based uniform resource pooling container scheduling engine and scheduling method thereof

    CN112363820A

  • Distributed unit cloud deployment method and system, equipment and storage medium

    CN113992688A

  • Method and device for acquiring network throughput data in real time

    CN117155827A

  • Distributed embedded software for a switch

    US20070061813A1

Cited By

  • Processor topology mapping management method and device, program and storage medium

    CN120670229A

  • Process processing method

    CN121455693A