A storage system performance optimization method and electronic device
Patent Information
- Application Number
- CN202610874844.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-17
AI Technical Summary
[0004]有鉴于此,本申请提供了一种存储系统性能优化方法及电子设备,以至少解决相关技术中无法根据实际负载动态调整资源分配,造成处理器资源利用率低下,难以充分发挥硬件性能潜力的问题,实现了网卡线程与处理器核心之间的动态迁移,可以适应动态变化的业务负载,并且实现了负载感知的资源分配,因此提高了处理器核心的资源利用率,减少了资源竞争与闲置,可以充分发挥多核硬件的潜力,从而提升了存储系统的存储性能
[0007] This application first collects performance monitoring metrics for different processor cores in the storage system. Then, based on the collected performance monitoring data, it calculates the current comprehensive load index for each processor core and selects a first processor core from multiple processor cores based on the current comprehensive load index. Next, based on the historical comprehensive load index, it selects a second processor core from each of the first processor cores and scans the network interface card (NIC) threads running on the second processor core. Then, it evaluates the first processor cores after removing the second processor core and determines the target core to be migrated based on the core evaluation results. Finally, it migrates the NIC threads to the target core. Through the above method, dynamic migration between NIC threads and processor cores is realized, which can adapt to dynamically changing business loads. Furthermore, this application performs migration based on the comprehensive load index obtained from real-time collected performance monitoring metrics, realizing load-aware resource allocation. Therefore, it improves the resource utilization of the cores, reduces resource contention and idleness, and can fully utilize the potential of multi-core hardware, thereby improving the storage performance of the storage system, such as data access efficiency.
Smart Images

Figure CN122450683B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method for optimizing the performance of a storage system and an electronic device. Background Technology
[0002] Currently, in IP SAN storage systems with multi-core CPU architecture, the competition for CPU resources between the storage-side network interface card (NIC) and system applications is a prominent issue. Critical tasks such as NIC interrupt handling threads and disk I / O threads may be randomly assigned to different CPU cores, leading to a decrease in CPU cache hit rate and an increase in cross-core communication overhead, which in turn affects overall storage performance.
[0003] However, current solutions often employ static CPU binding, such as fixing the network card driver thread to a specific CPU core, allocating interrupt handling for different network card queues to different CPU cores, optimizing network data processing by setting CPU affinity, and binding network transceiver threads and storage processing threads to independent cores. These methods fail to dynamically adjust resource allocation based on actual load, resulting in low CPU resource utilization and difficulty in fully realizing the hardware's performance potential. Furthermore, they do not consider the locality of CPU cache, leading to significant overhead for cross-core data transfer. Additionally, the lack of effective performance verification and rollback mechanisms makes it difficult to guarantee optimization results. Summary of the Invention
[0004] In view of this, this application provides a storage system performance optimization method and electronic device to at least solve the problem in related technologies that the inability to dynamically adjust resource allocation according to actual load results in low processor resource utilization and difficulty in fully realizing the hardware performance potential. It realizes dynamic migration between network card threads and processor cores, which can adapt to dynamically changing business loads, and realizes load-aware resource allocation. Therefore, it improves the resource utilization of processor cores, reduces resource contention and idleness, and can fully realize the potential of multi-core hardware, thereby improving the storage performance of the storage system.
[0005] This application provides a method for optimizing storage system performance, including: The performance monitoring metrics of each processor core in the storage system are collected to obtain the first performance monitoring data. Based on the first performance monitoring data, the comprehensive load index of the corresponding processor core is determined to obtain the first comprehensive load index. The storage system is a storage system based on a multi-core processor architecture. Based on the first comprehensive load index, the first processor core is selected from each processor core to obtain a candidate list, and the historical comprehensive load index of each first processor core in the candidate list is obtained. The second processor cores are selected from each first processor core based on the historical comprehensive load index, and the network card threads running on the second processor cores are scanned. The second processor core is removed from the candidate list to obtain the filtered list. Each first processor core in the filtered list is evaluated according to the preset evaluation indicators to obtain the core evaluation results. The preset evaluation indicators include hardware affinity, load balancing and cache locality. The target core is determined based on the core evaluation results, and the network card thread is migrated to the target core.
[0006] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described storage system performance optimization methods when executing the computer program.
[0007] This application first collects performance monitoring metrics for different processor cores in the storage system. Then, based on the collected performance monitoring data, it calculates the current comprehensive load index for each processor core and selects a first processor core from multiple processor cores based on the current comprehensive load index. Next, based on the historical comprehensive load index, it selects a second processor core from each of the first processor cores and scans the network interface card (NIC) threads running on the second processor core. Then, it evaluates the first processor cores after removing the second processor core and determines the target core to be migrated based on the core evaluation results. Finally, it migrates the NIC threads to the target core. Through the above method, dynamic migration between NIC threads and processor cores is realized, which can adapt to dynamically changing business loads. Furthermore, this application performs migration based on the comprehensive load index obtained from real-time collected performance monitoring metrics, realizing load-aware resource allocation. Therefore, it improves the resource utilization of the cores, reduces resource contention and idleness, and can fully utilize the potential of multi-core hardware, thereby improving the storage performance of the storage system, such as data access efficiency. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0009] Figure 1 This is a flowchart of a storage system performance optimization method disclosed in this application; Figure 2 This is a flowchart of a specific storage system performance optimization method disclosed in this application. Detailed Implementation
[0010] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0011] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0012] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0013] This application discloses a method for optimizing storage system performance. (See also...) Figure 1 As shown, the method includes: Step S11: Collect performance monitoring indicators of each processor core in the storage system to obtain the first performance monitoring data, and determine the comprehensive load index of the corresponding processor core based on the first performance monitoring data to obtain the first comprehensive load index; the storage system is a storage system based on a multi-core processor architecture.
[0014] It should be noted that the storage system performance optimization method proposed in this application is specifically applied to storage systems based on multi-core processor architecture, such as IP SAN (Storage Area Network, a storage area network technology based on IP network for transmitting block-level data) storage systems with multi-core CPU (Central Processing Unit) architecture. Furthermore, before performing performance optimization on the storage system, it is necessary to perform corresponding initialization operations to establish a basic operating environment for subsequent performance optimization operations.
[0015] Specifically, during storage system startup, a complete system architecture mapping can be established first through hardware topology detection. This can include CPU core distribution, cache hierarchy relationships, and NUMA (Non-Uniform Memory Access) node topology. This initialization process provides accurate hardware foundation data for subsequent affinity computing.
[0016] In one specific implementation, the initialization of the storage system may include the following four steps: hardware topology detection, parameter configuration loading, performance benchmark establishment, and monitoring infrastructure startup.
[0017] The specific execution process of hardware topology detection is as follows: First, by reading the / proc / cpuinfo file, the number of physical cores (cpu cores field), the number of logical cores (siblings field), and the hyper-threading status (HT flag in the flags field) are parsed to construct a basic CPU architecture information table; then, / sys / devices / system / cpu / cpu The ` / cache` directory is used to extract the level, size, and associated cores (shared_cpu_list) of each cache to generate a cache hierarchy graph (e.g., L1 cache is associated with 1 logical core, and L2 cache is associated with 2 logical cores). Finally, the node information such as node0 and node1 in the ` / sys / devices / system / node / ` directory is parsed, and the distance matrix between nodes is calculated through the node distances file (the core distance within the same node is 1, and the core distance across nodes is 2), thus providing a topological basis for subsequent affinity calculations.
[0018] The specific execution process of parameter configuration loading is as follows: First, the storage system loads the running parameters from the configuration file, that is, reads the key parameters from the / etc / ipsan_optimizer.conf configuration file. Specifically, it may include: cpu_threshold=0.75 (CPU comprehensive load index threshold, if exceeded, the performance optimization process is triggered); reserved_cores=[0,1] (the system reserves the core, used to run the kernel process, sshd (Secure Shell Daemon, the server daemon process of the SSH protocol) and other key services, and thread binding is prohibited); nic_thread_patterns=["irq / [0-9]+-nic", "net_rx[0-9]+","net_tx[0-9]+", "eth[0-9]+"] (regular expression used for network card thread identification, covering interrupt handling, protocol stack, and driver thread); monitor_interval=100ms (monitoring data collection interval); verify_window=5s (window duration for verifying optimization effect).
[0019] The specific execution process for establishing the performance benchmark is as follows: First, the storage system runs for a preset duration (e.g., 2 minutes) in an initial stable state (e.g., no peak traffic, CPU load <50%), and collects key performance indicators at preset time intervals (e.g., 100ms), such as network throughput (calculated by the difference in the bytes field of / proc / net / dev), IOPS (Input / Output Operations Per Second, the number of read / write operations per second on the hard drive, which can be obtained by iostat), and request latency (captured by tcpdump to obtain the round-trip time of data packets). Next, a sliding window averaging algorithm (window size = 60 sampling points, i.e., 6 seconds) is used to process the collected data. The specific formula is: Benchmark value = Σ(value of the i-th sampling point) / 60, to smooth out instantaneous fluctuations, thereby generating a performance benchmark (e.g., benchmark throughput = 950MB / s, benchmark latency = 2.5ms, benchmark CPU efficiency = 65%).
[0020] The specific execution process for starting the monitoring infrastructure is as follows: First, a shared memory circular buffer (e.g., size = 1024KB, a buffer storing the most recent 1000 monitoring data entries) is initialized, and a lock-free queue mechanism is used to avoid data contention. Next, the monitoring daemon process ipsan_monitor is started, and its scheduling policy is set to SCHED_FIFO with a priority of 90 (higher than ordinary processes) through the sched_setscheduler function, thereby ensuring that monitoring data collection is not affected by system load. At the same time, the logging module is initialized, and the initialization status, parameter configuration, and performance benchmark values are recorded, thereby providing a basis for subsequent troubleshooting.
[0021] After the above initialization operations, the performance of the storage system can be monitored, and performance optimization can be performed when the performance does not meet the preset performance requirements. Specifically, performance monitoring indicators of each processor core (such as CPU core) in the storage system can be periodically collected at preset time intervals (such as monitor_interval=100ms) to obtain the first performance monitoring data. The first performance monitoring data includes, but is not limited to, any one or more of the following: core utilization rate, cache hit rate, translation lookaside buffer (TLB) hit rate, network interface card queue depth, packet processing rate, and packet loss rate. Then, based on the first performance monitoring data, the comprehensive load index of the corresponding processor core is calculated to obtain the first comprehensive load index.
[0022] In this embodiment, to improve the efficiency of performance monitoring metric collection, multiple (e.g., three) monitoring threads (such as the read_cpu_stats() thread, the read_perf_counters() thread, and the read_netlink_stats() thread) can be used to collect performance monitoring metrics in parallel. The read_cpu_stats() thread, by parsing the / proc / stat file, can extract the user (user-mode runtime), nice (low-priority user-mode runtime), system (kernel-mode runtime), and idle (idle) fields for each CPU core to calculate core utilization. The read_perf_counters() thread can use the Linux perf subsystem to call the perf_event_open function to obtain L1 / L2 / L3 cache hit rates (cache hit rate = number of cache hits / (number of hits + number of misses)) and TLB hit rates. The read_netlink_stats() thread communicates with the kernel network subsystem via netlink sockets to obtain network interface card queue depth (number of packets currently waiting to be processed), packet processing rate (number of packets sent and received per second), and packet loss rate (number of lost packets / total number of packets). It should be noted that all collected data can be added with precise timestamps (precision = 1μs) to ensure timing consistency.
[0023] The process of obtaining core utilization can specifically include: obtaining the user-mode runtime, kernel-mode runtime, and idle runtime of each processor core; calculating the sum of the user-mode runtime and kernel-mode runtime to obtain a first sum; calculating the sum of the user-mode runtime, kernel-mode runtime, and idle runtime to obtain a second sum; and calculating the ratio of the first sum to the second sum to obtain the core utilization of the corresponding processor core. For example, obtaining the user (user-mode runtime), nice (low-priority user-mode runtime), system (kernel-mode runtime), and idle (idle) fields for each CPU core, then calculating user + nice + system and user + nice + system + idle to obtain the first and second sums, and then calculating (user + nice + system) / (user + nice + system + idle). 100% is the CPU core utilization rate.
[0024] In this embodiment, the comprehensive load index of the corresponding processor core is determined based on the first performance monitoring data to obtain the first comprehensive load index. Specifically, this may include: performing sliding window filtering on the first performance monitoring data corresponding to each processor core according to a preset window size to obtain filtered state data; identifying and removing outliers in the filtered state data to obtain preprocessed state data, and obtaining the load weight coefficients corresponding to the first performance monitoring data; calculating the product of each load weight coefficient and the corresponding first performance monitoring data to obtain the initial load index; and calculating the sum of multiple initial load indices corresponding to a single processor core to obtain the first comprehensive load index of the corresponding processor core. In this embodiment, the first performance monitoring data corresponding to each processor core (such as a CPU core) can be filtered by a sliding window according to a preset window size (such as 10 sampling points, i.e., 1 second) to obtain filtered state data. The specific formula is: filtered state data = Σ (value of the i-th sampling point) / 10. Through the above filtering process, instantaneous interference can be eliminated. Then, anomaly detection is performed on the filtered state data to identify and remove abnormal points in the filtered state data to obtain preprocessed state data. For example, the Z-score algorithm is used to detect abnormal points. The specific steps are: calculate the mean μ and standard deviation σ of the most recent 100 sampling points, and then calculate the Z-score value of the current data point. The specific calculation formula is: Z = (current value - μ) / σ. If |Z| > 3 (i.e., the deviation from the mean exceeds 3 times the standard deviation), the sampling point is determined to be an abnormal point. At this time, the value of the previous normal sampling point can be used to replace it, thereby ensuring the continuity of the data. For example, when the utilization of a CPU core suddenly jumps from 60% to 99% (Z-score=3.2), it is replaced with 62% from the previous sampling point.
[0025] Furthermore, obtain the load weight coefficients pre-set for different performance monitoring metrics in the first performance monitoring data. For example, the load weight coefficient for CPU utilization (40%), the load weight coefficient for cache efficiency (25%, cache efficiency = 1 - cache hit rate), the load weight coefficient for memory bandwidth utilization (20%, memory bandwidth utilization = current bandwidth / maximum bandwidth), and the load weight coefficient for network interface card queue depth ratio (15%, network interface card queue depth ratio = current queue depth / maximum queue depth). Then calculate the product of each load weight coefficient and the corresponding first performance monitoring data, for example, 0.4. CPU utilization, 0.25 Cache efficiency, 0.2 Memory bandwidth utilization, 0.15 The network interface card queue depth ratio is then calculated to be 0.4. CPU utilization +0.25 Cache efficiency +0.2 Memory bandwidth utilization +0.15 The network interface card (NIC) queue depth ratio yields the first comprehensive load index for the corresponding CPU core. The preprocessing operations described above, including filtering and outlier removal, improve data accuracy, making performance optimization more aligned with expectations. By assigning different load weight coefficients to different performance monitoring metrics and integrating multiple dimensions (such as CPU, cache, memory, NUMA, etc.), the current load status of the CPU core can be more accurately reflected, further enhancing performance optimization. Furthermore, multi-dimensional sensing (such as CPU, cache, memory, NUMA) facilitates more accurate identification of problematic CPU cores in subsequent steps.
[0026] Step S12: Based on the first comprehensive load index, filter the first processor cores from each processor core to obtain a candidate list, and obtain the historical comprehensive load index of each first processor core in the candidate list.
[0027] In this embodiment, after determining the comprehensive load index of each processor core, multiple processor cores (such as CPU cores) in the storage system can be filtered based on the first comprehensive load index to obtain a candidate list containing multiple first processor cores. Then, the historical comprehensive load index of each first processor core in the candidate list is obtained (the specific calculation method is the same as the calculation method of the first comprehensive load index).
[0028] Specifically, selecting candidate processor cores from each processor core based on a first comprehensive load index can include: selecting first processor cores whose first comprehensive load index is greater than a first threshold from multiple processor cores. In this embodiment, all CPU cores (excluding reserved_cores) can be traversed, and the first comprehensive load index corresponding to each CPU core can be compared with the first threshold (e.g., 0.75). If the first comprehensive load index is greater than the first threshold (i.e., first comprehensive load index > 0.75), then the CPU core is regarded as the first processor core (potential high-load core), and all the first processor cores obtained after screening are saved to the candidate list. For example, if the comprehensive load index of core 2 is 0.78 and the comprehensive load index of core 5 is 0.82, then these two cores are saved to the candidate list.
[0029] Step S13: Select the second processor core from each first processor core based on the historical comprehensive load index, and scan the network card threads running on the second processor core.
[0030] In this embodiment, multiple first processor cores in the candidate list are first screened again based on the historical comprehensive load index to obtain second processor cores, and then the network card threads running on the second processor cores are scanned.
[0031] Specifically, selecting second processor cores from each first processor core based on historical comprehensive load index can include: selecting second processor cores from multiple first processor cores whose historical comprehensive load index is greater than a first threshold. In this embodiment, the first processor cores (i.e., potential high-load cores > 0.75) are further screened. Specifically, a historical load queue of length 3 can be maintained for each first processor core (recording the comprehensive load index of the most recent 3 monitoring periods, i.e., 300ms). Only when all 3 values in the queue are > 0.75 is the core determined to be a continuously high-load core (i.e., the second processor core). For example, the historical queue values of core 2 are [0.78, 0.79, 0.80], which meets the condition that all values are > 0.75, so the core is confirmed to be a continuously high-load core; the queue values of core 5 are [0.82, 0.73, 0.76], which does not meet the condition that all values are > 0.75, so the core is removed from the candidate list.
[0032] In this embodiment, scanning the network interface card (NIC) threads running on the second processor core may specifically include: calculating the average of the first comprehensive load indices corresponding to multiple processor cores, obtaining the first average, and determining whether the first average is greater than a second threshold; if the second threshold is greater than the first threshold; if the first average is not greater than the second threshold, then obtaining the comprehensive load indices of neighboring cores related to the second processor core, obtaining the second comprehensive load indices, and determining whether multiple second comprehensive load indices are all greater than a third threshold; if the third threshold is less than the first threshold; if there is an index among the multiple second comprehensive load indices that is not greater than the third threshold, then using a depth-first traversal strategy to scan the NIC threads running on the second processor core. In this embodiment, the overall load of the storage system is first detected. Specifically, the average value of the first comprehensive load index corresponding to all processor cores in the storage system is calculated, and it is determined whether it is >0.85. If it is, it is determined to be system-level pressure, and optimizing a single core has limited effect, so the optimization process is not triggered at this time. If it is not, the comprehensive load index of the neighboring cores (such as the same NUMA node or CPU cores sharing L3 cache) of the high-load core (i.e., the second processor core) is obtained to obtain the second comprehensive load index, and it is determined whether the second comprehensive load index is all >0.7. If they are all greater than >0.7, it is determined to be a high-load core. At this time, there is local pressure in the system, and the optimization process is not triggered at this time. Only when the overall system load is ≤0.85 and there are cores with ≤0.7 among the neighboring cores, does the next step of the optimization process (i.e., scanning the network card threads running on the second processor core using a deep traversal strategy) proceed; otherwise, the monitoring and collection process in step S11 is returned, and data is collected periodically. This application uses a comprehensive load index and a three-level verification mechanism to determine whether to trigger the optimization process. First, it screens out potential high-load cores, then confirms the real high-load state through continuous verification, and finally detects the overall system load level. This ensures that the optimization process is only triggered in local performance bottleneck scenarios. In this way, it can avoid overreacting to instantaneous load fluctuations. Furthermore, the multi-level verification mechanism can accurately assess the system load state.
[0033] Specifically, when a high-load core requiring optimization is identified, a depth-first search strategy is used to scan all threads. Then, regular expression pattern matching is used to accurately identify network interface card (NIC) related threads (i.e., NIC threads), such as interrupt handling threads, network protocol stack threads, and NIC driver threads. Specifically, a depth-first search strategy can be used to scan the `task / ` subdirectories of all processes under the ` / proc / ` directory (each subdirectory corresponds to one thread), and obtain the thread PID, thread name, currently running core (via the `Cpus_allowed_list` field in ` / proc / [pid] / status`), and CPU affinity configuration. Next, regular expression matching is used to match the thread name. Threads matching any of the preset patterns `nic_thread_patterns` (such as `irq / 12-nic`, `net_rx0`, `eth0`) are identified as NIC-related threads. Furthermore, it can be verified whether the core currently running this network thread is the confirmed high-load core (i.e., the second processor core). If a match is found, it is added to the list of threads to be optimized. By scanning thread information using a depth-first search strategy, NIC threads running on the second processor core can be accurately identified.
[0034] In this embodiment, after scanning the network interface card (NIC) threads running on the second processor core, the process may further include: performing interrupt affinity analysis on the NIC device to obtain the processor cores with interrupt binding relationships to the NIC device, thus obtaining interrupt-bound cores; establishing a mapping relationship between interrupt-bound cores and the NIC device, thus obtaining a NIC core mapping table; evaluating the importance of each NIC thread to obtain a thread importance evaluation value, and sorting multiple NIC threads according to their thread importance evaluation values from highest to lowest, thus obtaining sorted threads; and designating the top-ranked NIC threads in the sorted threads as critical threads, and the other NIC threads as non-critical threads. In this embodiment, after scanning the NIC threads running on the second processor core, the interrupt affinity settings of the NIC device can also be analyzed to ensure a comprehensive understanding of the distribution of network processing tasks. Specifically, the first step is to identify the network interface card (NIC) interrupt binding relationships. This can be done by parsing the ` / proc / interrupts` file to extract the interrupt numbers associated with the NIC devices (e.g., interrupt numbers 12 and 13 for eth0). Then, by reading the ` / proc / irq / [irq] / smp_affinity` file (using a hexadecimal mask), the interrupt numbers are converted to their corresponding CPU cores (e.g., mask 0x04 corresponds to core 2). Next, the current bound core for each NIC interrupt is obtained (i.e., the interrupt-bound core). Based on the mapping relationship between interrupt numbers and interrupt-bound cores, a NIC core mapping table is established, providing data support for subsequent collaborative optimization. Additionally, the importance of each NIC thread can be evaluated to obtain thread importance assessment values. Then, multiple NIC threads are sorted according to their importance assessment values from highest to lowest. The top 80% of the sorted NIC threads are designated as critical threads, while other NIC threads are designated as non-critical threads. By dividing all network interface card (NIC) threads into critical and non-critical threads, it can be ensured that critical threads have priority in obtaining resources, thereby guaranteeing the priority execution of critical business operations and improving user experience.
[0035] In this embodiment, the importance of each network interface card (NIC) thread is evaluated to obtain a thread importance evaluation value. Specifically, this may include: obtaining the thread evaluation weights corresponding to preset thread importance evaluation indicators; the preset thread importance evaluation indicators include thread type, interrupt frequency, and data throughput; evaluating the importance of the NIC threads based on the thread evaluation weights to obtain a thread importance evaluation value; NIC threads include any one or more of interrupt handling threads, network protocol stack threads, and NIC driver threads. In this embodiment, first, the thread evaluation weights corresponding to the preset thread importance evaluation indicators (such as thread type, interrupt frequency, and data throughput) are obtained (for example, the thread evaluation weight for thread type is 0.6, the thread evaluation weight for interrupt frequency is 0.3, and the thread evaluation weight for data throughput is 0.1). Then, the importance of each NIC thread is evaluated using the thread importance evaluation indicators to obtain a thread importance evaluation value. The specific formula is: Thread importance evaluation value = 0.6 Thread type +0.3 Interrupt frequency +0.1 Throughput. Interrupt frequency = current thread interrupt frequency / system maximum interrupt frequency (value range 0-1), and throughput = current thread processing throughput / system maximum network throughput (value range 0-1). It should be noted that different threads may have different evaluation weights. For example, for thread type evaluation weights, interrupt handling thread = 1.0, network protocol stack thread = 0.8, and network card driver thread = 0.6. By using the above multi-indicator weighted fusion method to evaluate the importance of network card threads, we can more accurately determine their importance and ensure the effectiveness of storage system performance optimization.
[0036] Step S14: Remove the second processor core from the candidate list to obtain the filtered list, and evaluate each first processor core in the filtered list according to the preset evaluation indicators to obtain the core evaluation results; the preset evaluation indicators include hardware affinity indicators, load balancing indicators and cache locality indicators.
[0037] In this embodiment, high-load cores (i.e., second processor cores) are first removed from the candidate list to obtain a filtered list. Then, each first processor core in the filtered list is evaluated according to preset evaluation indicators to obtain core evaluation results. The preset evaluation indicators include, but are not limited to, hardware affinity indicators, load balancing indicators, and cache locality indicators.
[0038] In this embodiment, the second processor core is removed from the candidate list to obtain a filtered list, including: removing the reserved cores and the second processor core from the candidate list to obtain the removed cores; removing the cores whose first comprehensive load index exceeds the fourth threshold from the removed cores to obtain the filtered list; the fourth threshold is less than the third threshold. In this embodiment, the system reserved cores (i.e., reserved_cores) and the confirmed high-load cores (i.e., the second processor core) are first excluded from the candidate list to obtain the removed cores, and then the cores with a comprehensive load index > 0.6 are excluded from the removed cores. The remaining cores constitute the usable candidate list (i.e., the filtered list). For example, when the storage system has a total of 16 CPU cores, the reserved_cores[0,1], the high-load cores[2], and the cores with a comprehensive load index > 0.6[3,4] are excluded, and the filtered list is [5-15]. By removing the cores with a comprehensive load index > 0.6, the migration of network card threads to busy cores can be avoided.
[0039] Furthermore, each candidate core in the filtered list is evaluated from multiple dimensions. Specifically, the weights of each evaluation metric are first obtained, such as hardware affinity (40%), load balancing (40%), and cache locality (20%). Then, the core evaluation result for each candidate core is generated using the following formula: core evaluation result = 0.4 Hardware affinity score +0.4 Load balancing score +0.2 Cache locality score; where, hardware affinity score is obtained based on NUMA node distance matrix, score = 1 / NUMA distance (e.g., same node distance = 1, score = 1.0; cross node distance = 2, score = 0.5); load balancing score = 1 - comprehensive load index of candidate core (the lower the comprehensive load index, the higher the score, e.g., when the comprehensive load index is 0.3, the corresponding score is 0.7); cache locality score can be obtained by analyzing the historical cache access records of the network card thread to be optimized, for example, cache locality score = average historical cache hit rate of network card thread on candidate core (e.g., when the historical hit rate is 0.85, the corresponding score is 0.85).
[0040] Step S15: Determine the target core based on the core evaluation results and migrate the network card threads to the target core.
[0041] In this embodiment, the optimal target core can be determined based on the core evaluation results. For example, the core with the highest score in the core evaluation results can be selected as the target core. For instance, if the evaluation result of candidate core 6 is 0.82 and the evaluation result of core 7 is 0.79, core 6 is selected as the target core, and then the network card thread is migrated to the target core (i.e., core 6).
[0042] In this embodiment, migrating the network interface card (NIC) thread to the target core can specifically include: obtaining the state of each thread in the NIC thread to obtain multiple first thread state information; the first thread state information includes the bound core, processor affinity mask, thread scheduling policy, and thread priority; saving the first thread state information to a preset structure and pausing the execution of the NIC thread; adjusting the interrupt affinity of the NIC thread based on the first thread state information to bind critical threads to the target core, and then binding non-critical threads to the target core. In this embodiment, the migration of the NIC thread can be achieved through an atomic thread rebinding operation. Specifically, the process can begin by iterating through the list of threads to be optimized. The `prctl` function (a system call used to control and query process behavior and attributes) is used to obtain the raw state of each thread, such as its current core binding, CPU affinity mask, thread scheduling policy, and thread priority. This raw state information is then stored in a `thread_backup` structure (for rollback). Subsequently, the `kill` function (an important system call in Unix / Linux systems used to send signals to processes or process groups) is used to send a `SIGSTOP` signal to each thread to pause its execution (preventing state changes during migration). Simultaneously, the pause status (e.g., success / failure) is recorded, and threads in a failed state are removed from the list of threads to be optimized.
[0043] Furthermore, based on the first thread status information, the interrupt affinity of the network interface card (NIC) threads is adjusted. For example, for NIC threads that have been successfully paused, the CPU affinity is set in batches using the `sched_setaffinity` function (a system call function used to set the CPU affinity of a process or thread), binding them to the target core (such as core 6). For example, the affinity mask can be set to 0x40 (corresponding to core 6). Simultaneously, the scheduling policy of the network data processing thread can be changed to `SCHED_FIFO` with a priority of 99 (highest real-time priority) using the `sched_setscheduler` function, ensuring priority processing of network packets. For example, the interrupt handling thread PID=1234 can be bound to core 6, with the scheduling policy changed to `SCHED_FIFO` and the priority set to 99. For NIC threads identified as critical threads, they can be preferentially bound to the target core, i.e., resources are allocated preferentially, thereby improving the execution efficiency of the corresponding services. Additionally, adjusting interrupt affinity ensures service continuity.
[0044] Furthermore, after migrating the network interface card (NIC) thread to the target core, the process may further include: restoring the target core's operation and, after a preset running time, obtaining the NIC thread's status to acquire second thread status information; determining whether the first thread status information and the second thread status information are consistent; if the first thread status information and the second thread status information are consistent, then the NIC thread binding is considered successful, and the interrupt binding core corresponding to the NIC thread in the NIC core mapping table is synchronized to the target core. In this embodiment, a SIGCONT signal can be sent to all successfully migrated NIC threads to restore their operation. Then, after waiting for one scheduling cycle (e.g., 10ms), the current running core of the thread is verified through the / proc / [pid] / status file to determine whether it is the target core, whether the CPU affinity mask is correct, and whether the scheduling policy has been updated. If the verification is successful, the thread binding is considered successful, and the corresponding NIC thread is included in the optimization completion list. In addition, interrupt affinity can be synchronized. For example, based on the network card core mapping table mentioned above (which includes interrupt number-interrupt bound core), the bound core of the network card device interrupt can be synchronized to the target core by writing to the / proc / irq / [irq] / smp_affinity file (for example, setting the mask of interrupt bound core 12 to 0x40, thus binding it to core 6). By synchronizing interrupt affinity, it can be ensured that interrupt handling and thread run on the same core, eliminating the context switching overhead caused by cross-core interrupts and reducing the number of switching.
[0045] In this embodiment, the method may further include: if the first thread state information is inconsistent with the second thread state information, then the network card thread binding is determined to have failed, and a migration failure warning log is generated; the state of the network card thread that failed to bind is rolled back using the first thread state information recorded in the preset structure. In this embodiment, if the verification fails (i.e., the first thread state information is inconsistent with the second thread state information), the network card thread binding is determined to have failed, a migration failure warning log is generated, and a partial rollback is triggered (only restoring the original state of the thread). Specifically, the state of the network card thread that failed to bind can be rolled back using the first thread state information recorded in the thread_backup structure, such as rolling back to the current bound core, CPU affinity mask, thread scheduling policy, and thread priority recorded in the first thread state information. Through the verification mechanism and fast rollback operation, the optimization process can be made smooth and safe, business continuity is guaranteed, and the reliability of the storage system is enhanced. At the same time, the optimization process can be ensured to be safe and adaptive, ultimately achieving a synergistic improvement in storage performance and processor resource utilization.
[0046] As can be seen, this embodiment first collects performance monitoring metrics of different processor cores in the storage system, then calculates the current comprehensive load index of each processor core based on the collected performance monitoring data, and selects a first processor core from multiple processor cores based on the current comprehensive load index. Then, it selects a second processor core from each first processor core based on the historical comprehensive load index, and scans the network interface card (NIC) threads running on the second processor core. The first processor cores after removing the second processor core are then evaluated, and the target core to be migrated is determined based on the core evaluation results. Finally, the NIC threads are migrated to the target core. Through the above method, dynamic migration between NIC threads and processor cores is realized, which can adapt to dynamically changing business loads. Furthermore, this embodiment performs migration based on the comprehensive load index obtained from real-time collected performance monitoring metrics, realizing load-aware resource allocation. Therefore, it improves the resource utilization of the cores, reduces resource contention and idleness, and can fully utilize the potential of multi-core hardware, thereby improving the storage performance of the storage system, such as data access efficiency.
[0047] This application discloses a specific method for optimizing storage system performance; see [link to relevant documentation]. Figure 2 As shown, the method includes: Step S21: Collect performance monitoring indicators of each processor core in the storage system to obtain the first performance monitoring data, and determine the comprehensive load index of the corresponding processor core based on the first performance monitoring data to obtain the first comprehensive load index; the storage system is a storage system based on a multi-core processor architecture.
[0048] Step S22: Based on the first comprehensive load index, filter the first processor cores from each processor core to obtain a candidate list, and obtain the historical comprehensive load index of each first processor core in the candidate list.
[0049] Step S23: Select the second processor core from each first processor core based on the historical comprehensive load index, and scan the network card threads running on the second processor core.
[0050] Step S24: Remove the second processor core from the candidate list to obtain the filtered list, and evaluate each first processor core in the filtered list according to the preset evaluation indicators to obtain the core evaluation results; the preset evaluation indicators include hardware affinity indicators, load balancing indicators and cache locality indicators.
[0051] Step S25: Determine the target core based on the core evaluation results, estimate the load of the target core after migrating the network card threads to the target core, obtain the estimated load, and determine whether the estimated load is less than or equal to the preset load threshold.
[0052] In this embodiment, after determining the optimal target core, the load of the target core after the network card thread is migrated to the target core is estimated to obtain the estimated load, and then it is determined whether the estimated load is greater than or equal to a preset load threshold (for example, whether the estimated load is ≤0.85).
[0053] In this embodiment, the estimated load on the target core after migrating the network interface card (NIC) threads to the target core can be obtained by: acquiring the load of the target core and the load of each thread within the NIC threads to obtain the core load and the load of multiple threads, and calculating the average of the loads of multiple threads to obtain a second average; and summing the second average with the core load to obtain the estimated load of the target core. That is, the current load of the target core after the migration of the NIC threads to be migrated, and the average load of all the NIC threads to be migrated are acquired, and then the current load plus the average load is calculated to obtain the estimated load of the target core.
[0054] Step S26: If the estimated load is less than or equal to the preset load threshold, analyze the cache access mode of the network card thread to obtain the cache miss rate after migration, and determine whether the cache miss rate is greater than the preset miss rate.
[0055] In this embodiment, if the estimated load is ≤0.85, cache consistency verification is performed. Specifically, the cache access pattern of the network card thread to be migrated can be analyzed using the perf tool (a performance analysis tool based on hardware performance counters) to obtain the cache miss rate after migration, and then it is determined whether the cache miss rate is greater than the preset miss rate (e.g., 10%).
[0056] Step S27: If the cache miss rate is not greater than the preset miss rate, then determine the memory access node where the network card thread's memory data is located and the memory access node where the target core is located, and obtain the first memory access node and the second memory access node.
[0057] In this embodiment, if the cache miss rate after migration is ≤10%, the memory access node (such as NUMA node) where the memory data of the network card thread to be optimized is located and the memory access node where the target core is located are further determined to obtain the first memory access node (NUMA1 node) and the second memory access node (NUMA2 node).
[0058] Step S28: Determine whether the first memory access node and the second memory access node belong to the same node.
[0059] Step S29: If the first memory access node and the second memory access node belong to the same node, then migrate the network card thread to the target core.
[0060] In this embodiment, if NUMA1 and NUMA2 belong to the same node, the network interface card (NIC) thread is migrated to the target core; if NUMA1 and NUMA2 do not belong to the same node, no migration operation is performed.
[0061] In this embodiment, after migrating the network interface card (NIC) thread to the target core, the process may further include: collecting performance monitoring metrics of the target core after migration at preset time intervals to obtain second performance monitoring data; using a two-tailed t-test to verify whether the difference between the first performance monitoring data and the second performance monitoring data exceeds a preset difference threshold; if the difference between the first performance monitoring data and the second performance monitoring data exceeds the preset difference threshold, it is determined that the performance of the storage system after migration has been significantly improved. In this embodiment, after completing the thread migration operation, performance monitoring metrics (50 samples in total) of the target core after migration can be collected at preset time intervals (e.g., 100ms), and the collected metrics are consistent with the performance monitoring metrics in step S11, such as network throughput, IOPS, request latency, CPU utilization, cache hit rate, etc. For example, the collected network throughput samples are [1100, 1150, ..., 1250] MB / s, and the latency samples are [2.1, 2.0, ..., 1.7] ms.
[0062] Next, a two-tailed t-test is used to verify the significance of the performance improvement, with a confidence level set at 95%. Specifically, the mean μ1 and standard deviation σ1 of the 50 optimized samples are first calculated, and the baseline value μ0 before optimization is obtained. Then, the t-statistic is calculated using the following formula: Next, consulting the t-distribution table (degrees of freedom = 49), the critical value at the 95% confidence level is ±2.01. If t > 2.01 and p < 0.05, then the performance improvement is considered statistically significant (excluding the influence of random fluctuations). For example, if the mean throughput after optimization is 1216 MB / s, the baseline is 950 MB / s, t = 15.3, and p < 0.001, then it can be determined that the performance of the storage system has significantly improved after the migration.
[0063] Furthermore, business metrics can be collected throughout the performance optimization process, such as connection interruption counts (using netstat statistics), packet loss rate, and timeout event counts (obtained through application log analysis). If the number of connection interruptions is 0, the packet loss rate is less than 0.1%, and the number of timeout events is less than 5, it is determined that the optimization has not had a negative impact on the business; otherwise, it is determined to be a business anomaly (i.e., it has had a negative impact on the business), and a rollback process is triggered. By evaluating the actual effect of the optimization operation, the authenticity of the performance improvement and the security of the business can be ensured.
[0064] In addition, the success of the optimization can be determined from multiple dimensions, and then the experience can be recorded or rolled back. Specifically, the criteria for successful optimization can be that the following four conditions are met simultaneously: Condition 1: Network throughput increase ≥ 10% ((optimized average - baseline) / baseline ≥ 0.1); Condition 2: Request latency reduction ≥ 8% ((baseline - optimized average) / baseline ≥ 0.08); Condition 3: CPU efficiency improvement ≥ 5% (CPU efficiency = throughput / CPU utilization, improvement ≥ 0.05); Condition 4: Performance improvement is statistically significant (p < 0.05) and there are no business anomalies. For example, if throughput increases by 28%, latency decreases by 32%, and CPU efficiency increases by 22%, and all conditions are met simultaneously, the optimization is considered successful; if throughput increases by 8% but less than 10%, the optimization is considered a failure.
[0065] Next, based on the above judgment results, corresponding operations are performed. For example, when the optimization is judged to be successful, the successful experience is recorded; when the optimization is judged to be unsuccessful (no success condition is met or there is a business anomaly), a rollback operation is performed. Furthermore, the complete decision context can be saved, recording in detail the decision context of each optimization, including timestamps, decision results, verification data, optimization parameters, and system status. This information can form a complete decision trajectory, facilitating subsequent problem analysis and algorithm improvement. Simultaneously, the storage system can maintain real-time decision statistics to track optimization success rates and optimization effect trends, thereby providing data support for the system's self-learning.
[0066] When an optimization is deemed successful, a process for recording successful experience information can be executed. Complete optimization context information is saved to a knowledge base, such as source cores, target cores, a list of threads to be optimized, the degree of performance improvement, and system status. Based on this successful experience information, effective optimization patterns can be directly identified, and optimization algorithm parameters can be fine-tuned, thereby improving the overall efficiency of subsequent performance optimizations. Specifically, the complete successful experience information for this optimization can be stored in a knowledge base (such as an SQLite database). This successful experience information can include: timestamps, source cores (e.g., high-load cores), target cores, a list of threads to be optimized (PID, name, importance weight, etc.), performance benchmark values, post-optimization metric values, changes in the overall load index, and optimization parameters used (e.g., thresholds, weight allocation). The knowledge base supports vector similarity retrieval; by calculating the similarity between the current system status and historical records (e.g., CPU core count, NUMA topology, load characteristics, etc.), the optimal performance optimization solution can be quickly matched. Next, based on the successful experience information recorded in the knowledge base, cluster analysis is performed according to four dimensions: workload type (e.g., random read / write / sequential read / write / mixed read / write), core combination (source core-target core pairing), thread composition (percentage of interrupt handling threads), and NUMA topology (same node / cross node). This analysis uses the K-means algorithm (K=5) to extract optimization patterns effective across multiple scenarios. For example, if a pattern using random read / write load + core migration within the same NUMA node + priority binding of interrupt handling threads achieves a 92% success rate, this pattern is included in the pattern library. Furthermore, the weight coefficients of the optimization algorithm can be dynamically adjusted based on successful experience information, such as using a moving average method (window size = 10 successful optimization records), with the formula: New weight = 0.9. Old weight +0.1 The current success weight contribution. For example, if multiple success records show that cache locality significantly contributes to latency improvement, its weight is adjusted from 20% to 25%; if load balancing has the greatest impact on throughput improvement, its weight remains unchanged at 40%.
[0067] When optimization is deemed a failure, meaning the optimization effect does not meet expectations, an automatic rollback operation is immediately executed to restore the network interface card (NIC) to its pre-optimization state. The cause of the failure is analyzed to avoid repeated, ineffective attempts. Specifically, an atomic process consistent with the migration process can be used first: all migrated NIC threads are paused, and their original CPU affinity, scheduling policy, and priority are restored using the `thread_backup` structure. Then, a SIGCONT signal is sent to resume thread execution, and the rollback success rate is recorded (target: 100%). It should be noted that the entire rollback process should be controlled within 50ms to minimize the impact on business operations (e.g., application response latency increases should not exceed 10ms). Next, interrupt affinity is restored. The saved original interrupt binding configuration is read first, and then written to the ` / proc / irq / [irq] / smp_affinity` file to restore the NIC interrupt to its pre-optimization core binding state, thereby ensuring consistency between interrupt handling and thread execution. Additionally, log analysis tools can be used to extract key failure information to analyze the reasons for failure. For example, improper target core selection (e.g., incorrect load estimation, resulting in a load > 0.85 after migration); incompatible thread combinations (e.g., high-priority threads competing for the same core); and unsuitable system conditions (e.g., the presence of other high-load processes during optimization). Based on the identified causes, the corresponding source core, target core, and thread combinations can be added to a temporary blacklist (valid for 1 hour) to avoid repeated attempts in the short term. Simultaneously, the optimization algorithm can be adjusted, such as increasing load redundancy when selecting target cores (estimated load ≤ 0.8), thereby improving decision accuracy.
[0068] It should be noted that regardless of whether the optimization is successful, the storage system will enter a parameter update phase, dynamically adjusting operating parameters based on the experience gained from this optimization, thereby achieving self-learning capabilities. Specifically, the threshold can be adaptively adjusted. First, the success rate (number of successful optimizations / total number of optimizations) and the average improvement ((optimized metric - baseline value) / average of baseline values) of the most recent 10 optimizations are calculated. If the success rate is ≥80% and the average improvement is ≥15%, then cpu_threshold is reduced by 0.05 (down to a minimum of 0.60) to adopt a more aggressive optimization strategy; if the success rate is ≤40% or the average improvement is <5%, then cpu_threshold is increased by 0.05 (up to a maximum of 0.90) to adopt a conservative strategy; otherwise, the threshold remains unchanged. For example, if the success rate of the most recent 10 optimizations is 85% and the average improvement is 20%, the threshold is reduced from 0.75 to 0.70. Furthermore, the weights can be optimized. For example, based on the most recent 20 optimization records (successful records + failure records), linear regression analysis can be used to analyze the correlation between each scoring dimension (such as hardware affinity, load balancing, cache locality) and optimization success, and the weight coefficients can be dynamically adjusted. The formula is: New weight = Old weight + Correlation coefficient The total weight should be kept at 1.0. For example, if the correlation coefficient between load balancing and success is 0.8 and cache locality is 0.3, then the load balancing weight should be increased by 0.04, cache locality by 0.01, and hardware affinity by 0.05. Additionally, monitoring strategies can be adjusted; for example, the variance of the comprehensive load index over the most recent minute can be calculated (variance = Σ(i-th sampling point - mean)). 2 The monitoring interval is then assessed based on the number of samplings. If the variance is less than 0.01 (stable load), the monitoring interval is increased from 100ms to 500ms to save system resources. If the variance is greater than or equal to 0.05 (fluctuating load), the monitoring interval is shortened to 50ms to improve response speed. Otherwise, the 100ms interval is maintained. Additionally, data collection dimensions can be added under fluctuating loads (such as adding memory page error rate) to ensure the capture of critical performance changes. Once the parameters (such as weighting coefficients) are updated, a new round of monitoring and optimization is initiated to achieve continuous adaptive storage system performance optimization.
[0069] For more detailed processing procedures of steps S21 to S24, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0070] As can be seen, the embodiments of this application calculate the current comprehensive load index of each processor core based on the performance monitoring indicators of each processor core, and perform core screening by combining the historical comprehensive load index to obtain high-load cores. Then, cores that do not contain high-load cores are evaluated from multiple dimensions to obtain target cores. Network card threads running on high-load cores are then migrated to target cores. Through the above dynamic migration method, the resource utilization of the processor is greatly improved, resource contention and idleness are reduced, the potential of multi-core hardware is fully utilized, and the storage performance of the storage system is improved. Moreover, the entire optimization process does not require manual intervention and the binding strategy can be adjusted in real time, thus adapting to various business scenarios and reducing operation and maintenance costs. In addition, before migration, this application estimates the load of the target core after migration. If the estimated load is less than or equal to a preset load threshold, it determines whether the cache miss rate after migration is greater than the preset miss rate. If it is not greater, it further determines whether the memory access node where the network card thread's memory data is located and the memory access node where the target core is located belong to the same node. If they belong to the same node, the migration operation is performed. That is, only when all three verifications above pass, the migration is confirmed. Otherwise, the second-best candidate core is selected again for verification. In this way, cross-core overhead can be eliminated.
[0071] Accordingly, embodiments of this application also disclose a storage system performance optimization apparatus, which includes: The acquisition module is used to collect performance monitoring metrics of each processor core in the storage system to obtain the first performance monitoring data; the storage system is a storage system based on a multi-core processor architecture. The first determining module is used to determine the comprehensive load index of the corresponding processor core based on the first performance monitoring data, and obtain the first comprehensive load index; The first filtering module is used to filter the first processor cores from each processor core based on the first comprehensive load index, obtain a candidate list, and obtain the historical comprehensive load index of each first processor core in the candidate list. The second filtering module is used to filter the second processor cores from each first processor core based on the historical comprehensive load index, and scan the network card threads running on the second processor cores. The removal module is used to remove the second processor core from the candidate list, resulting in a filtered list. The evaluation module is used to evaluate each first processor core in the filtered list according to preset evaluation indicators to obtain core evaluation results; the preset evaluation indicators include hardware affinity indicators, load balancing indicators and cache locality indicators. The second determination module is used to determine the target core based on the core evaluation results and migrate the network card thread to the target core.
[0072] The specific workflow of each of the above modules can be found in the relevant content disclosed in the foregoing embodiments, and will not be repeated here.
[0073] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the storage system performance optimization method.
[0074] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the storage system performance optimization method when running.
[0075] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0076] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described storage system performance optimization method embodiments.
[0077] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described storage system performance optimization method embodiments.
[0078] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be executed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-a-chip (SoC), a complex programmable logic device (CPLD), a microprocessor (MCU), etc. The terms "system," "computing device," or "apparatus" as used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or one or more combinations thereof. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.
[0079] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0080] The above provides a detailed description of a storage system performance optimization method and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A method for optimizing the performance of a storage system, characterized in that, include: The performance monitoring metrics of each processor core in the storage system are collected to obtain the first performance monitoring data, and the comprehensive load index of the corresponding processor core is determined based on the first performance monitoring data to obtain the first comprehensive load index; the storage system is a storage system based on a multi-core processor architecture. Based on the first comprehensive load index, the first processor cores are selected from each processor core to obtain a candidate list, and the historical comprehensive load index of each first processor core in the candidate list is obtained. Based on the historical comprehensive load index, the second processor core is selected from each of the first processor cores, and the network card threads running on the second processor core are scanned. The second processor core is removed from the candidate list to obtain a filtered list. Each first processor core in the filtered list is evaluated according to a preset evaluation index to obtain a core evaluation result. The preset evaluation index includes hardware affinity index, load balancing index, and cache locality index. The hardware affinity index is determined based on the NUMA node distance matrix, the load balancing index is determined based on the current comprehensive load index of each first processor core, and the cache locality index is determined based on the historical cache hit rate of the network card thread in the corresponding first processor core. Based on the core evaluation results, the target core is determined, and the network interface card (NIC) thread is migrated to the target core. The step of determining the comprehensive load index of the corresponding processor core based on the first performance monitoring data to obtain the first comprehensive load index includes: performing sliding window filtering on the first performance monitoring data corresponding to each processor core according to a preset window size to obtain filtered state data; identifying and removing outliers in the filtered state data to obtain preprocessed state data, and obtaining the load weight coefficient corresponding to the first performance monitoring data; calculating the product of each load weight coefficient and the corresponding first performance monitoring data to obtain an initial load index; and calculating the sum of multiple initial load indices corresponding to a single processor core to obtain the first comprehensive load index of the corresponding processor core. The step of scanning network interface card (NIC) threads running on the second processor core includes: calculating the average of the first comprehensive load index corresponding to multiple processor cores to obtain a first average, and determining whether the first average is greater than a second threshold; the second threshold is greater than the first threshold; if the first average is not greater than the second threshold, then obtaining the comprehensive load index of neighboring cores related to the second processor core to obtain a second comprehensive load index, and determining whether multiple second comprehensive load indices are all greater than a third threshold; the third threshold is less than the first threshold; if there is an index among the multiple second comprehensive load indices that is not greater than the third threshold, then a depth-first traversal strategy is used to scan the NIC threads running on the second processor core. The step of scanning the network interface card (NIC) threads running on the second processor core using a depth-first traversal strategy includes: scanning all threads running on the second processor core using a depth-first traversal strategy, and then accurately identifying NIC-related threads by regular expression pattern matching to obtain the NIC threads; the NIC threads include any one or more of interrupt handling threads, network protocol stack threads, and NIC driver threads.
2. The storage system performance optimization method according to claim 1, characterized in that, The first performance monitoring data includes any one or more of the following: core utilization, cache hit rate, translation back buffer hit rate, network interface card queue depth, packet processing rate, and packet loss rate. The process of obtaining the core utilization rate includes: Obtain the user-mode runtime, kernel-mode runtime, and idle runtime of each processor core; Calculate the sum of the user-mode runtime and the kernel-mode runtime to obtain a first sum value; Calculate the sum of the user-mode runtime, the kernel-mode runtime, and the idle runtime to obtain a second sum value; Calculate the ratio of the first sum to the second sum to obtain the core utilization rate of the corresponding processor core.
3. The storage system performance optimization method according to claim 1, characterized in that, The process of selecting a candidate list of processor cores from each processor core based on the first comprehensive load index includes: From multiple processor cores, a candidate list is obtained by selecting the first processor core whose first comprehensive load index is greater than the first threshold. Accordingly, the step of selecting the second processor core from each of the first processor cores based on the historical comprehensive load index includes: Select a second processor core from multiple first processor cores whose historical comprehensive load index is greater than the first threshold.
4. The storage system performance optimization method according to claim 3, characterized in that, After scanning the network interface card (NIC) threads running on the second processor core, the process also includes: Interrupt affinity analysis is performed on the network interface card (NIC) device to identify the processor cores that have interrupt binding relationships with the NIC device, thus obtaining the interrupt-bound cores; Establish the mapping relationship between the interrupt binding core and the network interface card (NIC) device to obtain the NIC core mapping table; The importance of each network interface card (NIC) thread is evaluated to obtain a thread importance evaluation value. The NIC threads are then sorted in descending order of their thread importance evaluation values to obtain the sorted threads. The network card threads that rank at the top of the sorted threads according to a predetermined ratio are designated as critical threads, while the other network card threads are designated as non-critical threads.
5. The storage system performance optimization method according to claim 4, characterized in that, The evaluation of the importance of each network interface card (NIC) thread to obtain a thread importance evaluation value includes: Obtain the thread evaluation weights corresponding to preset thread importance evaluation indicators; the preset thread importance evaluation indicators include thread type, interrupt frequency, and data throughput; The importance of the network card threads is evaluated based on the thread evaluation weights to obtain thread importance evaluation values.
6. The storage system performance optimization method according to claim 4, characterized in that, The step of migrating the network interface card (NIC) thread to the target core includes: The state of each thread in the network interface card (NIC) thread is obtained to obtain multiple first thread state information; the first thread state information includes core binding, processor affinity mask, thread scheduling policy, and thread priority. Save the first thread state information to a preset structure and pause the operation of the network card thread; Based on the first thread state information, the interrupt affinity of the network card thread is adjusted to bind the critical thread to the target core, and then bind the non-critical thread to the target core.
7. The storage system performance optimization method according to claim 6, characterized in that, After migrating the network interface card (NIC) thread to the target core, the process further includes: Restore the operation of the target core, and after running for a preset time, obtain the status of the network card thread to obtain the second thread status information; Determine whether the first thread state information is consistent with the second thread state information; If the first thread status information is consistent with the second thread status information, it is determined that the network card thread binding is successful, and the interrupt binding core corresponding to the network card thread in the network card core mapping table is synchronized to the target core.
8. The storage system performance optimization method according to claim 7, characterized in that, Also includes: If the first thread status information is inconsistent with the second thread status information, the network card thread binding is determined to have failed, and a migration failure warning log is generated. Using the first thread state information recorded in the preset structure, the state of the network card thread that failed to bind is rolled back.
9. The storage system performance optimization method according to claim 1, characterized in that, The step of removing the second processor core from the candidate list to obtain the filtered list includes: Remove the reserved core and the second processor core from the candidate list to obtain the removed core; After removing the cores, remove the cores whose first comprehensive load index exceeds the fourth threshold to obtain the filtered list; the fourth threshold is less than the third threshold.
10. The storage system performance optimization method according to claim 1, characterized in that, The step of migrating the network interface card (NIC) thread to the target core includes: Estimate the load of the target core after migrating the network interface card (NIC) threads to the target core, obtain the estimated load, and determine whether the estimated load is less than or equal to a preset load threshold. If the estimated load is less than or equal to the preset load threshold, the cache access mode of the network card thread is analyzed to obtain the cache miss rate after migration, and it is determined whether the cache miss rate is greater than the preset miss rate. If the cache miss rate is not greater than the preset miss rate, then the memory access node where the memory data of the network card thread is located and the memory access node where the target core is located are determined to obtain the first memory access node and the second memory access node. Determine whether the first memory access node and the second memory access node belong to the same node; If the first memory access node and the second memory access node belong to the same node, then the network card thread is migrated to the target core.
11. The storage system performance optimization method according to claim 10, characterized in that, The estimated load is obtained by migrating the network interface card (NIC) threads to the target core, including: The load of the target core and the load of each thread in the network card thread are obtained respectively to obtain the core load and the load of multiple threads, and the average of the multiple thread loads is calculated to obtain the second average. The estimated load of the target core is obtained by summing the second mean and the core load.
12. The storage system performance optimization method according to any one of claims 1 to 11, characterized in that, After migrating the network interface card (NIC) thread to the target core, the process further includes: The performance monitoring metrics of the target core after migration are collected at preset time intervals to obtain the second performance monitoring data. The two-tailed t-test was used to verify whether the difference between the first performance monitoring data and the second performance monitoring data exceeded a preset difference threshold. If the difference between the first performance monitoring data and the second performance monitoring data exceeds a preset difference threshold, it is determined that the performance of the storage system has been significantly improved after the migration.
13. An electronic device, characterized in that, It includes a processor and a memory; wherein, when the processor executes a computer program stored in the memory, it implements the storage system performance optimization method as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Interrupt request processing method and device, equipment and storage medium
CN121326524A
Multi-core scheduling system and method based on interrupt affinity and storage medium
CN121681116A