Read and write request response method and device, electronic device, and computer program product

By dividing memory access domains under the CXL memory architecture, obtaining performance parameters, calculating weights and dynamically selecting buffers, the memory resource allocation problem of the database system under the CXL memory architecture is solved, and efficient read and write request responses and resource utilization are achieved.

CN120448137BActive Publication Date: 2025-09-23INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510944400.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-09-23
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

In a server environment based on the CXL memory architecture, the question is how the database system can effectively allocate memory resources to respond to read and write requests based on real-time performance parameters and memory access patterns, thereby maximizing performance and resource utilization.

Method used

By dividing multiple memory modules into different access domains, obtaining the performance parameters of each domain, calculating the weight parameters, and dynamically selecting the target memory buffer to respond to read and write requests, the memory resource allocation is optimized by combining real-time performance monitoring and weight parameter updates.

Benefits of technology

It improves the operational efficiency and user experience of the database system, ensures the intelligence and efficiency of resource allocation, and maximizes performance and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448137B_ABST
    Figure CN120448137B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for responding to read and write requests, an electronic device, and a computer program product, relating to the field of databases, comprising: dividing a plurality of memory modules into a plurality of memory access domains according to division parameters corresponding to the plurality of memory modules, wherein the plurality of memory modules are located in a computer system, and the access delays of the plurality of memory access domains are different; upon receiving a read and write request for a database of the computer system, obtaining a first performance parameter of the plurality of memory access domains within a historical time period; calculating a weight parameter of the plurality of memory access domains of the database according to the first performance parameter; determining a target memory buffer in the plurality of memory access domains according to the plurality of the weight parameters, and responding to the read and write request according to the target memory buffer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of storage, and in particular to a method and device for responding to read and write requests, an electronic device, and a computer program product. Background Art

[0002] With the rapid expansion of data volumes and the increasing complexity of computing tasks, computer systems are placing increasingly stringent demands on memory performance and capacity. This led to the emergence of Compute Express Link (CXL) technology, which effectively integrates memory resources of varying performance levels through a tiered memory architecture, providing an innovative solution for improving overall system performance. Among CXL's diverse application scenarios, database application optimization is a key focus. The CXL tiered memory architecture integrates memory at multiple performance levels. The upper tier typically consists of high-performance main memory (such as DRAM), while the lower tier consists of larger but lower-performance CXL memory. In this architecture, a unified memory address space encompasses multiple tiers, and the proper allocation of data across these tiers is crucial for improving system performance.

[0003] Regarding the related technology, in a server environment based on the CXL memory architecture, how can the database system effectively allocate memory resources to respond to read and write requests based on real-time performance parameters and memory access patterns, thereby maximizing performance and resource utilization? No effective solution has been proposed yet. Summary of the Invention

[0004] The present application provides a method and apparatus for responding to read and write requests, an electronic device, and a computer program product to at least address the related art problem of how a database system, in a server environment based on a CXL memory architecture, can effectively allocate memory resources to respond to read and write requests based on real-time performance parameters and memory access patterns, thereby maximizing performance and resource utilization.

[0005] The present application provides a method for responding to read and write requests, comprising: dividing the multiple memory modules into multiple memory access domains according to division parameters corresponding to the multiple memory modules, wherein the multiple memory modules are located in a computer system, and the access delays of the multiple memory access domains are different; upon receiving a read and write request to a database of the computer system, obtaining first performance parameters of the multiple memory access domains within a historical time period; calculating weight parameters of the multiple memory access domains of the database according to the first performance parameters; determining a target memory buffer in the multiple memory access domains according to the multiple weight parameters, and responding to the read and write request according to the target memory buffer.

[0006] The present application also provides a device for responding to read and write requests, including: a partitioning module, used to divide the multiple memory modules into multiple memory access domains according to the partitioning parameters corresponding to the multiple memory modules, wherein the multiple memory modules are located in a computer system, and the access delays of the multiple memory access domains are different; an acquisition module, used to obtain first performance parameters of the multiple memory access domains within a historical time period when receiving a read and write request to a database of the computer system; a calculation module, used to calculate weight parameters of the multiple memory access domains of the database according to the first performance parameters; a response module, used to determine a target memory buffer in the multiple memory access domains according to the multiple weight parameters, and respond to the read and write request according to the target memory buffer.

[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned methods for responding to read and write requests when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for responding to any read or write request are implemented.

[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned methods for responding to read and write requests when the computer program is executed by a processor.

[0010] This application ensures intelligent and efficient resource allocation by acquiring real-time performance parameters, calculating multi-dimensional weight parameters, and dynamically selecting target memory buffers, significantly improving the operational efficiency and user experience of database systems based on the CXL memory architecture. This addresses the related technical issue of how database systems in CXL memory architecture-based server environments can effectively allocate memory resources to respond to read and write requests based on real-time performance parameters and memory access patterns, thereby maximizing performance and resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 1 is a hardware structure block diagram of a computer system according to a method for responding to a read / write request in an embodiment of the present application;

[0013] Figure 2 is a flowchart of a method for responding to a read / write request according to an embodiment of the present application;

[0014] Figure 3 is a flowchart of a database optimization method based on CXL according to an embodiment of the present application;

[0015] Figure 4 is a schematic diagram of an improved NUMA architecture according to an embodiment of the present application;

[0016] Figure 5 This is a structural block diagram of a read / write request response device according to an embodiment of the present application. DETAILED DESCRIPTION

[0017] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0018] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0019] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0020] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the response method for the read and write request depends, the specific application environment architecture or specific hardware architecture is described here.

[0021] The method embodiments provided in the embodiments of the present application can be executed in a computer system or similar computing device. Taking running on a computer system as an example, Figure 1 This is a hardware structure diagram of a computer system according to a method for responding to a read / write request according to an embodiment of the present application. Figure 1 As shown, the computer system may include one or more ( Figure 1Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. The computer system may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer system. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0022] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the startup method of the operating system in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the computer system via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0023] Transmission device 106 is used to receive or transmit data via a network. A specific example of such a network may include a wireless network provided by a communications provider of the computer system. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0024] In this embodiment, a method for responding to a read / write request is provided, including but not limited to being applied to a computer system. Figure 2 is a flow chart of a method for responding to a read / write request according to an embodiment of the present application, such as Figure 2 As shown, the method includes the following steps S202-S206:

[0025] Step S202: dividing the plurality of memory modules into a plurality of memory access domains according to division parameters corresponding to the plurality of memory modules, wherein the plurality of memory modules are located in a computer system, and the plurality of memory access domains have different access delays;

[0026] Step S204, when a read or write request to a database of the computer system is received, obtaining first performance parameters of the plurality of memory access domains within a historical time period;

[0027] Step S206, calculating weight parameters of multiple memory access domains of the database according to the first performance parameter;

[0028] Step S208: determining a target memory buffer in the multiple memory access domains according to the multiple weight parameters, and responding to the read / write request according to the target memory buffer.

[0029] Through the above steps, by acquiring real-time performance parameters, calculating multi-dimensional weight parameters, and dynamically selecting the target memory buffer, intelligent and efficient resource allocation is ensured, significantly improving the operational efficiency and user experience of database systems based on the CXL memory architecture. This solves the related technical problem of how database systems in CXL memory architecture-based server environments can effectively allocate memory resources to respond to read and write requests based on real-time performance parameters and memory access patterns, thereby maximizing performance and resource utilization.

[0030] In an exemplary embodiment, the multiple memory modules are divided into multiple memory access domains according to the division parameters corresponding to the multiple memory modules, including: determining the physical distance between the multiple memory modules and the central processing unit of the computer system according to the hardware layout information and hardware specification information of the computer system, wherein the division parameters include the physical distance; and dividing the multiple memory modules into the multiple memory access domains according to the physical distance intervals in which the multiple physical distances are located, wherein the multiple memory access domains correspond to multiple different physical distance intervals respectively.

[0031] First, based on the computer system's hardware layout and specifications, the physical distance between each memory module and the central processing unit (CPU) is determined. Physical distance is a partitioning parameter that is related to signal propagation time and affects data access latency. Based on the physical distance intervals within which the memory modules are located, these memory modules are categorized into multiple memory access domains, each corresponding to a specific physical distance interval. This allows the system to intelligently place important data in the memory access domain closest to the CPU, where access latency is lowest, significantly enhancing data processing capabilities and response speeds.

[0032] Optionally, in another exemplary embodiment, the multiple memory modules are divided into multiple memory access domains according to the division parameters corresponding to the multiple memory modules, including: determining the interface type of the connection interface corresponding to the multiple memory modules, wherein the connection interface is used to connect the memory module to the computer system, and the division parameters include the interface type; dividing the multiple memory modules into the multiple memory access domains according to the interface type, wherein the multiple memory access domains correspond one-to-one to the multiple interface types.

[0033] First, the interface type used by all memory modules is identified and determined. The interface type determines the data transmission rate and latency, and serves as a partitioning parameter. Then, based on each interface type, the memory modules are assigned to corresponding memory access domains. Each access domain is directly associated with a specific interface type. This correspondence ensures that memory modules using the same interface type share optimized access paths, reducing data flow barriers and improving overall performance.

[0034] Specifically, based on this embodiment, in the existing NUMA (Non-Uniform Memory Access) architecture, each CPU socket is designed to reside in an independent NUMA domain (i.e., the aforementioned memory access domain). Each physical CPU socket typically integrates one or more memory controllers. The physical memory directly connected to the memory controller on that socket is referred to as the local memory of that NUMA node. This memory provides the lowest access latency and highest bandwidth for the CPU core on that socket. Multiple CPU sockets are connected via high-speed interconnect channels. The bandwidth and latency of these channels determine the performance of remote access (a core on one NUMA node accessing the memory of another NUMA node), which is slower than local access. CXL devices are identified as CPU-less NUMA nodes. While their local memory has large capacity, it also has higher latency than DRAM (approximately 100 nanoseconds higher). In this architecture, NUMA domains are divided based on access latency, which can be defined as NUMA domain 1, NUMA domain 2, and so on. Specifically, DRAM directly connected to the CPU socket forms a faster NUMA domain because its physical distance to the CPU core is shorter, reducing data transfer time. Conversely, CXL memory modules, because they communicate with the CPU through a slower interface, have higher access latency and thus act as a slower NUMA domain. This design allows the operating system and applications to optimize data placement and processing based on memory access latency, thereby improving overall system performance.

[0035] In an exemplary embodiment, responding to the read and write request according to the target memory buffer includes: in the process of responding to the read and write request, collecting in real time a second performance parameter of the database in response to the read and write request, wherein the first performance parameter and the second performance parameter are of different categories; dynamically updating the multiple weight parameters according to the second performance parameter, and redetermining the target memory buffer based on the multiple weight parameters after the dynamic update.

[0036] The read and write request response process integrates real-time performance monitoring and a dynamic weight parameter update mechanism, directly linking to database operation performance optimization. Secondary performance parameters include I / O latency, CPU utilization, memory access latency, particularly cross-NUMA node access latency, cache hits, and bandwidth utilization details, complementing the primary performance parameters to provide a comprehensive performance view.

[0037] Weight parameters are updated based on the second performance parameter, namely the actual system performance during runtime. This includes, but is not limited to, latency caused by data read and write operations, CPU resource consumption, latency differences between DRAM and CXL memory accesses, cache efficiency, and inter-NUMA communication bandwidth usage. The updated weight parameters are used to relocate the target memory buffer to ensure data is stored at the most appropriate memory tier, improving access efficiency.

[0038] I / O latency, the time difference between initiating an I / O request and fully processing the data, reflects the database system's responsiveness to read and write operations. CPU utilization, which indicates the extent of computing resource consumption, is crucial for assessing system performance bottlenecks. Memory access latency, which distinguishes the difference in memory access time between local and remote nodes, provides a basis for optimizing data placement. Cache hit rate, which measures cache utilization efficiency, is directly related to reducing direct memory accesses and improving database responsiveness. Bandwidth utilization, which monitors data traffic on communication paths between NUMA nodes, helps adjust load balancing strategies.

[0039] This mechanism dynamically adjusts weight parameters through real-time monitoring and feedback to ensure that database read and write operations are always executed at optimal performance. Even when the load changes or system resource conditions fluctuate, it can automatically optimize and maintain high performance.

[0040] Optionally, calculating the weight parameters of the plurality of memory access domains of the database according to the first performance parameter includes: using the formula Calculate the weight parameters W of the multiple memory access domains i , where W i is the weight parameter of the i-th memory access domain, F i is the data access frequency of the i-th memory access domain, L i is the delay sensitivity coefficient of the i-th memory access domain, Hi is the heat decay factor of the i-th memory access domain, S i is the system load rate of the i-th memory access domain, R i is the cache hit rate of the i-th memory access domain, is the sum of the data access frequencies of the multiple memory access domains, ε is a very small constant, α is an empirical coefficient, i is a positive integer, and the first performance parameter includes: the data access frequency, the delay sensitivity coefficient, the heat attenuation factor, the system load rate, and the cache hit rate.

[0041] The weight parameter calculation incorporates key indicators such as data access frequency, latency sensitivity, heat attenuation factor, system load rate, and cache hit rate. The formula design comprehensively considers data heat, latency sensitivity, system pressure, and cache mechanism efficiency, aiming to achieve optimal resource allocation for the database under the CXL memory architecture. ε is a very small constant introduced to avoid mathematical calculation anomalies with a denominator of zero, and α is an empirical coefficient used to adjust the cache hit rate R. i For weight W i degree of impact.

[0042] Data access frequency reflects the volume of data requests for a specific memory domain over a period of time and serves as the basis for measuring data heat. The latency sensitivity coefficient, based on business characteristics, quantifies the sensitivity to response time and has different settings for OLTP (online transaction processing) and OLAP (online analytical processing) businesses to accommodate their differing latency requirements. The heat decay factor smooths the natural decline in data heat over time, ensuring that bursty accesses do not excessively impact resource allocation. The system load rate monitors the resource utilization of memory domains and indicates the system pressure that needs to be considered when allocating memory resources. The cache hit rate demonstrates the effectiveness of the database caching mechanism and is directly related to reducing the number of direct memory accesses and improving data processing speed.

[0043] Overall, this formula enables dynamic evaluation of the performance advantages of different memory access domains through the interaction between parameters, ensuring that database read and write requests are responded to at the most appropriate memory level, thereby optimizing overall data processing efficiency and performance.

[0044] Optionally, obtaining a first performance parameter of the database within a first time period includes: calculating the data access frequency through an operating system tool, a hardware performance counter, a CXL management interface, and database built-in monitoring, wherein the operating system tool is used to obtain overall IO data of the operating system of the database, the hardware performance counter is used to count the number of read and write operations performed on the multiple memory access domains, the CXL management interface is used to count the access frequency of the CXL memory through the CXL switch, and the database built-in monitoring is used to collect the overall IO data of the database, and the multiple memory access domains include the CXL memory.

[0045] During the first time period, the acquisition of the first performance parameter involved multi-dimensional monitoring, encompassing the operating system, hardware performance counters, the CXL management interface, and internal database monitoring, all working together to calculate the frequency of critical data access. Operating system tools provide insight into the overall I / O activity of the database OS, hardware performance counters pinpoint the read and write frequencies of multiple memory access domains, the CXL management interface leverages the CXL switch to measure CXL memory access frequency, and the database's built-in monitoring system comprehensively captures database I / O dynamics. These four elements work together to ensure accurate and comprehensive assessment of data access heat.

[0046] Operating system tools provide a comprehensive overview of IO operations, including read and write rates, wait times, and buffer usage, which is key to understanding system-level IO load. Hardware performance counters delve into the CPU and memory subsystems, recording access events to specific memory regions and providing a microscopic perspective for analyzing changes in data popularity. The CXL management interface, through interaction with the CXL switch, obtains detailed access records of CXL memory, particularly information related to latency and throughput, which is crucial for optimizing the interaction between CXL memory and the database. Built-in database monitoring provides deep insights into internal database operations, tracking IO activity caused by SQL queries, assisting in identifying hot data sets and performance bottlenecks, and ensuring that optimization strategies are aligned with actual business needs.

[0047] By integrating this monitoring information, data access frequency calculations not only incorporate system-level I / O overviews, hardware-level memory read / write statistics, CXL memory access characteristics, and database operational analysis, but also consider the dynamic changes in data popularity within the time window, laying a solid data foundation for subsequent resource allocation strategies. Through meticulous monitoring and rigorous calculations, this mechanism accurately reflects the real-time database popularity distribution within the CXL memory system, providing a strong basis for performance optimization.

[0048] Optionally, the data access frequency is calculated through operating system tools, hardware performance counters, CXL management interface and database built-in monitoring, including: obtaining multidimensional access data of the database in real time through the operating system tools, the hardware performance counters, the CXL management interface and the database built-in monitoring; calculating the multidimensional access data through the exponentially weighted moving average method in a sliding window to obtain the data access frequency.

[0049] The technical solution of this embodiment describes an integrated monitoring and dynamic evaluation strategy for data access frequency calculation. It uses operating system tools, hardware performance counters, CXL management interfaces, and database built-in monitoring to capture the multi-dimensional access behavior of the database in real time. Subsequently, the exponentially weighted moving average method within a sliding window is used to accurately calculate the collected data to obtain the data access frequency that reflects the data popularity.

[0050] Operating system tools provide an overview of the database OS's overall IO activity, capturing comprehensive IO data such as database read and write rates and wait times, providing insights into the database's overall IO pressure. Hardware performance counters provide in-depth statistics on the number of read and write operations for each memory access domain, specifically counting accesses to CXL memory. The CXL management interface, working in conjunction with the CXL switch, specifically counts the frequency of CXL memory accesses, including direct and indirect accesses, to obtain information on CXL memory resource utilization. Built-in database monitoring records IO activity caused by SQL queries, providing deep insights into database access patterns and ensuring that data access frequency calculations are aligned with actual business needs.

[0051] By acquiring access data from multiple dimensions in real time, combined with the exponentially weighted moving average method within a sliding window, we achieve dynamic and accurate data popularity assessment. By assigning higher weight to recently accessed data while gradually reducing the influence of past access data, we ensure that data access frequency promptly reflects the database's current hot spots, providing a key basis for dynamic resource allocation strategies and optimizing overall database performance within the CXL in-memory architecture.

[0052] Optionally, obtaining the first performance parameter of the database in the first time period includes: determining the business type corresponding to the read and write request, wherein the business type includes: online transaction processing and online analytical processing; and determining the delay sensitivity coefficient according to the business type.

[0053] Obtaining the primary performance parameter of a database over a specific time period involves identifying the business characteristics behind read and write requests. Business types are categorized as either online transaction processing (OLTP) or online analytical processing (OLAP). The former emphasizes the immediacy and concurrency control of data operations, while the latter focuses on complex data analysis and report generation. OLTP, due to its pursuit of rapid response, has a higher coefficient. OLAP, given its batch processing characteristics, has a lower latency sensitivity coefficient to balance computation and latency.

[0054] Online transaction processing (OLTP) primarily handles daily operational tasks, such as banking transactions and order processing. These businesses require rapid data access and are extremely sensitive to latency. Any increase in latency may impact user experience or system stability. Therefore, its latency sensitivity coefficient is set high to ensure that the system can maintain efficient response even when facing a large number of concurrent transactions.

[0055] Online Analytical Processing (OLAP) is suitable for data warehouse environments and supports multidimensional analysis and complex queries, such as business intelligence reports and market trend analysis. Compared with OLTP, OLAP places more emphasis on data processing capabilities and the detailed nature of query results rather than the speed of individual operations. Therefore, its latency sensitivity coefficient is lower, which means that when retrieving and analyzing large-scale data, the system can appropriately sacrifice response time in exchange for more sufficient computing resources and data integrity.

[0056] The latency sensitivity coefficient is determined based on the business type, aiming to accurately match the database's response strategy to read and write requests. This ensures that both high-frequency small-scale transactions and low-frequency large-scale data analysis can find the optimal performance balance point under the CXL memory architecture, thereby improving overall system performance.

[0057] It should be noted that the heat decay factor is set according to the business model to prevent the impact of sudden business volume on the business; the system load rate also has a relatively large impact on performance. The system load rate can be obtained through system monitoring data. The cache hit rate is the result of the combined effect of multiple factors such as data characteristics, database architecture design, and resource allocation. The cache hit rate also directly affects the performance of the overall database.

[0058] Optionally, determining the target memory buffer in the multiple memory access domains based on multiple weight parameters includes: determining a first memory size corresponding to the read and write requests; allocating the first memory size according to the multiple weight parameters to obtain multiple second memory sizes, wherein the multiple second memory sizes correspond one-to-one to the multiple memory access domains; and determining the target memory buffer in the multiple memory access domains based on the multiple second memory sizes.

[0059] First, the first memory size required for read and write requests is determined. Then, based on the weight parameters of each memory access domain, the first memory size is proportionally allocated to generate the second memory size corresponding to each memory access domain. This allocation step ensures the rational layout of memory resources. Finally, based on these second memory sizes, the target memory buffer is implemented in each memory access domain, achieving optimal storage of data between different levels of memory.

[0060] First, memory size, that is, the amount of raw memory space required for read and write requests, is the basic unit of database operations and determines whether data can be loaded or stored smoothly. It is determined by factors such as the data block size and the number of records involved in the request, and is directly related to database performance and resource consumption.

[0061] Memory allocation guided by weight parameters dynamically adjusts the storage share of read and write data within each domain based on the performance parameter ratio of each memory access domain. This allocation mechanism takes into account multiple factors such as data access popularity, latency characteristics, and system load. The weight parameters calculated through a mathematical model determine the optimal distribution of the first memory size among different levels of memory, thereby improving data processing efficiency and system response speed.

[0062] Determining the target memory buffer is the final implementation of the resource allocation strategy. Based on the pre-calculated secondary memory size, specific memory space is allocated within the corresponding memory access domain. This step ensures that data is stored in the memory level that best suits its characteristics, reduces data access latency, and improves cache hit rates, playing a key role in enhancing database performance. Especially in the CXL memory architecture, a reasonable target memory buffer setting can significantly improve the database's utilization of large-capacity memory resources.

[0063] Furthermore, the target memory buffer is determined in the multiple memory access domains according to the multiple second memory sizes, including: allocating sub-memory buffers in the multiple memory access domains according to the multiple second memory sizes, wherein the sub-memory buffers are continuous memory spaces; when the sub-memory buffers are successfully allocated to the multiple memory access domains, the multiple sub-memory buffers are determined as the target memory buffers.

[0064] The strategy of demarcating contiguous sub-memory buffers within each memory access domain based on the second memory size to form the target memory buffer is designed to ensure efficient data storage and access. Specifically, based on the second memory size corresponding to each access domain, the system attempts to allocate contiguous sub-memory buffers within that domain. These contiguous memory spaces facilitate fast data reading and writing. Only when all memory access domains have been successfully allocated to corresponding sub-memory buffers are these sub-memory buffers integrated and confirmed as target memory buffers, thereby ensuring that database operations are optimized within the multi-level memory architecture and avoiding the performance loss that may be caused by data fragmentation.

[0065] Sub-memory buffers are contiguous memory spaces defined by the second memory size within a memory access domain. They serve as temporary storage for data and are crucial for improving database performance. Contiguous memory spaces reduce CPU memory access latency and improve data loading and storage efficiency. In the CXL memory architecture, this contiguous nature significantly impacts cross-domain data transfer optimization.

[0066] The final determination of the target memory buffer depends on the successful allocation of each sub-memory buffer. This mechanism ensures that the database can fully utilize the characteristics of each level of memory to achieve optimal data storage when executing read and write requests. At the same time, it also considers the performance differences of memory access domains. Through the allocation of continuous sub-memory buffers guided by weight parameters, it ensures the efficiency and balance of database operations under the CXL memory layered architecture, thereby improving the overall data processing speed and system responsiveness.

[0067] Based on the above steps, after allocating sub-memory buffers in the multiple memory access domains according to the multiple second memory sizes, the method also includes: when a sub-memory buffer of the third memory size is not allocated in the m-th memory access domain, determining the fourth memory size of the largest continuous memory space in the m-th memory access domain, and allocating the largest continuous memory space as the sub-memory buffer of the m-th memory access domain, wherein m is a positive integer; allocating a sub-memory buffer of the fifth memory size in the m+1-th memory access domain, wherein the fifth memory size = the third memory size - the fourth memory size + the sixth memory size, the sixth memory size is the second memory size corresponding to the m+1-th memory access domain, and the access delay of the m-th memory access domain is lower than the access delay of the m+1-th memory access domain.

[0068] This embodiment details the response strategy when sub-memory buffer allocation fails as expected. When an attempt to allocate space of the third memory size in the mth memory access domain fails, the system immediately switches to an adaptive path, automatically detecting the fourth memory size of the largest contiguous memory within that domain and quickly locking it as the new sub-memory buffer. Simultaneously, for the m+1th memory access domain, the originally planned fifth memory size is recalculated, adjusted to the third memory size minus the fourth memory size, plus the sixth memory size originally allocated to that domain. This dynamic adjustment ensures maximum utilization of existing resources and maintains data access continuity and efficiency even in the face of memory fragmentation challenges.

[0069] The maximum contiguous memory space refers to the longest contiguous memory segment that can be immediately allocated within the mth memory access domain. Its determination is crucial for resolving memory fragmentation issues, especially in high-performance database application scenarios. Contiguous memory space has significant advantages in reducing addressing time and improving data loading speed. The automatic detection and allocation mechanism in this embodiment further enhances the system's resource scheduling capabilities and reduces the risk of performance bottlenecks caused by memory discontinuity.

[0070] Access latency measures the time it takes to retrieve and store data within a specific memory access domain. Lower access latency for the mth domain indicates faster data access within that domain, making it more suitable for storing hot data. Higher access latency for the m+1th domain suggests it may be more suitable for storing cold or backup data. During resource allocation, considering access latency helps optimize data distribution across different memory tiers, improving overall system responsiveness and data processing efficiency. Precisely controlling access latency is crucial for fully leveraging the advantages of heterogeneous memory, particularly in the CXL memory architecture.

[0071] Based on the above steps, after allocating the maximum continuous memory space as a sub-memory buffer of the mth memory access domain, the method also includes: when the sub-memory buffer of the fifth memory size is not allocated in the m+1th memory access domain, allocating a discrete memory buffer of the fifth memory size in the m+1th memory access domain, or recalculating the weight coefficients of the multiple memory access domains and re-determining the target memory buffer in the multiple memory access domains based on the updated weight coefficients.

[0072] If the maximum contiguous memory space allocated to the mth memory access domain is insufficient to meet subsequent larger memory requirements, that is, if a matching contiguous memory block cannot be found in the m+1th memory access domain, one of the following two methods can be used to handle the problem:

[0073] 1. If there is no continuous space of the fifth memory size in the (m+1)th domain, but there is a sufficient amount of scattered memory, then these discrete memory blocks are combined to meet the space requirements of the fifth memory size, ensuring that database operations or other data-intensive tasks can continue unimpeded.

[0074] 2. Start the weight coefficient recalculation process, and adjust the allocation weight coefficient of each memory access domain based on the latest system status and resource usage. Then, based on the updated weight, the system will again try to determine the target memory buffer within all memory access domains. This process helps to distribute memory resources more evenly, avoid overloading a single memory access domain, and thus optimize overall system performance.

[0075] Optionally, the second performance parameter includes: the IO delay of the read / write request, the memory access delay corresponding to the read / write request, the CPU utilization corresponding to the read / write request, the cache hit rate corresponding to the read / write request, the bandwidth utilization corresponding to the read / write request, the memory allocation success rate corresponding to the read / write request, and the operation completion time of the read / write request.

[0076] The second set of performance parameters deeply analyzes the multi-dimensional performance impact of read and write requests, covering I / O latency, memory access latency, CPU utilization, cache hit rate, bandwidth utilization, memory allocation success rate, and operation completion time, together depicting the overall system performance when the database processes requests.

[0077] IO latency, which is the total waiting time from when the database issues a read or write instruction to when it receives a response, is directly related to user experience and system responsiveness. Low IO latency means efficient data interaction.

[0078] Memory access latency, broken down into local and remote access, reveals the actual time it takes the CPU to retrieve data. Low-latency memory access is crucial for improving data processing speed, especially in the CXL memory architecture, where remote access latency can increase significantly. Proper management can effectively avoid performance bottlenecks.

[0079] CPU utilization refers to how busy the CPU core is during the execution of read and write requests. High utilization may indicate CPU resource shortage, affecting overall system performance.

[0080] The cache hit rate reflects the availability of read and write request data in the cache. A high hit rate means fewer memory accesses, thereby improving system response speed.

[0081] Bandwidth utilization, for data transmission between different NUMA nodes, shows the efficiency of interconnection channel usage. High bandwidth utilization may indicate the need to optimize data distribution or adjust access patterns.

[0082] The memory allocation success rate measures the degree of implementation when allocating memory by weight. A high success rate indicates that the resource allocation strategy is consistent with the actual memory status, which helps improve data access efficiency.

[0083] Operation completion time, that is, the total time from the issuance of a read or write request to its complete processing, comprehensively considers the impact of all performance parameters and intuitively reflects the efficiency of the database in processing requests.

[0084] The comprehensive consideration of this series of parameters provides a comprehensive perspective for optimizing database performance within the CXL memory architecture. Continuous monitoring and analysis of these parameters accurately identifies performance bottlenecks and dynamically adjusts resource allocation strategies, ensuring database operations maintain high performance and efficiency while fully leveraging the advantages of CXL memory, providing users with a superior data service experience.

[0085] Obviously, the embodiments described above are only part of the embodiments of the present application, rather than all the embodiments. In order to better understand the above method, the above process is described below in conjunction with the embodiments, but it is not intended to limit the technical solutions of the embodiments of the present application. Specifically:

[0086] In an optional embodiment, the present application proposes a database optimization method based on CXL, the process of which is as follows: Figure 3 As shown, the following steps are included:

[0087] Step 1. Start;

[0088] Step 2: NUMA domain division;

[0089] In an optional embodiment, the present application improves the existing NUMA architecture to support the above-mentioned response method for read and write requests. The improved NUMA architecture is as follows: Figure 4 As shown, each CPU socket is designed to be in an independent NUMA domain, meaning that all cores of the CPU and their associated DRAM memory are in a single NUMA domain, forming a tightly coupled computing and memory unit. All memory modules connected through the CXL interface constitute another independent NUMA domain. In this architecture, NUMA domains are divided based on access latency. Specifically, DRAM directly connected to the CPU socket constitutes a NUMA domain with faster access speeds because they are physically closer to the CPU cores, reducing data transfer time. In contrast, because CXL memory modules communicate with the CPU through a relatively slow interface, they have higher access latency and therefore act as a relatively slow NUMA domain. This design allows the operating system and applications to optimize data placement and processing based on the latency characteristics of memory access, thereby improving overall system performance.

[0090] Step 3: Obtain performance parameters including at least data access frequency, current system load rate, cache hit rate, and delay sensitivity coefficient;

[0091] Through analysis of database business scenarios, the parameters that affect database performance are data access frequency and latency. The database itself also varies for different business types. For example, there are significant differences between OLTP and OLAP, and there are significant differences in IO requirements and latency sensitivity. In view of these situations, combined with actual application scenarios, data parameters such as data access frequency, latency sensitivity coefficient, heat attenuation factor, current system load rate, and cache hit rate are analyzed to evaluate and optimize the performance of the database in memory under CXL. The above performance parameters (equivalent to the first performance parameter in the above embodiment) include:

[0092] (1) Data access frequency. Considering that data access frequency can be obtained in a variety of ways, in order to ensure the accuracy of the data, the data access frequency is obtained through a comprehensive evaluation using operating system tools, hardware performance counters, CXL management interface, and database built-in monitoring. Operating system tools can obtain the IO data of the entire operating system, hardware performance counters can count data such as CXL_mem_read and CXL_mem_write, and the CXL management interface can obtain IO data and latency data through the CXL switch; the database built-in monitoring can obtain the IO data within the database.

[0093] (2) Delay sensitivity coefficient: By analyzing and comparing the OLTP and OLAP business models and combining them with business needs, we can get the empirical value data of 1.5 for OLTP and 0.8 for OLAP.

[0094] (3) The heat attenuation factor is set according to the business model to prevent the impact of sudden business volume on the business.

[0095] (4) The system load rate has a significant impact on performance, and the system load rate can be obtained through system monitoring data.

[0096] (5) The cache hit rate is the result of the combined effects of multiple factors, including data characteristics, database architecture design, and resource allocation. The cache hit rate also directly affects the performance of the overall database.

[0097] Step 4: Dynamic resource allocation;

[0098] By combining the obtained parameter information with the core characteristics of the CXL protocol and the dynamic changes in database access volume, a weight formula is used to combine key factors such as data popularity, latency sensitivity, system load rate, and cache hit rate. Combining the CXL protocol mechanism, resource allocation strategy, and load balancing algorithm, the weight coefficient W of each NUMA domain is obtained. i The dynamic adjustment method is as follows: , the definition and data source of each parameter are shown in Table 1:

[0099] Table 1

[0100]

[0101] The data access frequency F can be calculated by counting the number of accesses in a time window (e.g., 5 minutes), and then using the exponentially weighted moving average method.

[0102] Example: 0.7 × Fprev (number of visits in the previous time window) + 0.3 × Fcurrent (number of visits in the current time window).

[0103] L: Delay sensitivity coefficient, dynamically configured based on the CXL device type. The value is 1.5 for OLTP (online transaction processing) and 0.8 for OLAP (online analytical processing).

[0104] H: heat decay factor, H = e -λ , λ is the decay rate (usually 2-4).

[0105] S: System load rate, S=current capacity / maximum capacity.

[0106] ε: A very small constant (such as e -5 ) to avoid division by zero errors.

[0107] α: Empirical coefficient (usually 0.2-1).

[0108] When R < threshold, increasing the α value increases the cache weight.

[0109] R: Cache hit rate, R = number of hits / (number of hits + number of misses).

[0110] In order to better illustrate this embodiment, the following examples are given:

[0111] Assume that the database model is an OLTP model, the access heat decay factor is 2, the cache hit rate is 40%, α is 0.5, the overall system load rate is 40%, the DRAM access times in the first 5 minutes are 20,000, and the current access times are 30,000. Then the access times F can be obtained as F=0.7*20,000+0.3*30,000=23,000. The total access times Fsum is 50,000. Therefore, the DRAM Wdram is , substituting into the formula we can get ≈0.43. Thus, under this model, Wdram is 0.43. Similarly, based on these assumptions, we can calculate the memory weights for each NUMA domain and each level.

[0112] The system calculates the database model to obtain the weight of each NUMA domain. When reading and writing database IO data, the execution program allocates memory resources according to the weight. At the same time, relevant parameter values ​​are collected during the execution process to facilitate dynamic adjustment.

[0113] For example:

[0114] Assuming there are only two NUMA nodes, according to the above formula, the weight of NUMA node 1 (N1) is 0.43, and the weight of NUMA node 2 (N2) is 0.57. When the database initiates a query that requires reading data stored in the data store, the database's I / O executor (e.g., the thread responsible for processing the query) needs to allocate a memory buffer for the data to be read from the data store. The executor queries the currently effective weighting policy and finds that the read weights for the OLTP model are: N1: 0.43, N2: 0.57.

[0115] Suppose this read requires a 100MB contiguous memory buffer. The executor will prioritize allocating 43MB (100*0.43) of memory on N1. If N1 successfully allocates 43MB, it attempts to allocate the remaining 57MB (100*0.57) on N2. If N1 cannot allocate enough contiguous memory (for example, only 30MB of contiguous space remains), the executor will attempt to allocate as much as possible (30MB) on N1 and then allocate the remaining 13MB on N2. Alternatively, the executor will recalculate the weights and fall back to allocating more on N2, but prioritize proximity to N1. Ultimately, the 100MB of data is read and populated into this weighted memory area.

[0116] During the execution process, the database data parameters are collected and the performance impact on the database side is observed. If there are changes in performance or business models, dynamic calculations and adjustments are performed in a timely manner to achieve optimal performance.

[0117] The collection parameters include the following:

[0118] While I / O operations (read or write) are being executed, and as the data is subsequently processed by the database (e.g., SQL calculations, index builds), the system continuously collects key metrics related to performance and NUMA efficiency:

[0119] I / O latency: The time from the initiation of a read / write request to its completion.

[0120] Memory Access Latency: The latency of the CPU accessing the memory involved in this I / O operation. Pay particular attention to the ratio and specific duration of different NUMA nodes (CPU access to the memory of its own NUMA node and CPU access to the memory of another NUMA node across nodes).

[0121] CPU Utilization: CPU utilization for processing the I / O and related calculations, calculated by NUMA node.

[0122] Cache hit rate: The hit rate of CPU cache at each level.

[0123] Bandwidth utilization: The bandwidth usage of NUMA inter-node interconnections (such as QPI / UPI).

[0124] Memory allocation success rate / failure rate: indicates the ratio and size of successful and failed memory allocation attempts on different NUMA nodes based on weights.

[0125] Operation completion time: The execution time of the entire query or transaction.

[0126] Step 5: Performance feedback and dynamic correction;

[0127] This application dynamically adjusts the weight distribution of memory at each level by designing the weights of each NUMA domain. To ensure the effective implementation of this application, this application also designs a performance feedback mechanism and a dynamic correction mechanism. Using system monitoring tools, this application monitors system-level load conditions such as CPU utilization, memory utilization, and disk I / O load in real time. This allows the application to understand the overall resource usage of the system, allowing for comprehensive consideration of the global allocation of system resources when optimizing database performance. Based on the performance evaluation results, the weight distribution of each NUMA domain and different memory levels is adjusted.

[0128] For example, based on the above example, during another time period, the system load was very high, N1's memory was severely fragmented, and sufficient N1 memory could not be allocated according to the weight (43%). As a result, more data was forced to be allocated to N2 (for example, an average of 70%). Monitoring found that the threads processing these queries often ran on Socket 0, but the remote latency for accessing the 70% of data on N2 was very high, resulting in a large number of CPU cycles waiting for memory, worsening the overall query time, and tightening the QPI bandwidth.

[0129] Step 6: Performance optimization;

[0130] After analysis, the weight calculation module decided to adjust: the system observed that placing data as much as possible on N1 still offered significant benefits (lower latency), but current N1 memory pressure (fragmentation and high utilization) made allocating 43% difficult and costly (allocation failures and potential memory reclaim). After calculation, the system decided to slightly reduce N1's weight and increase N2's weight (for example, N1: 0.3; N2: 0.7). This reduced the probability of allocation failures and wait times on N1. Although single-access latency was slightly higher, it avoided the increased latency and QPI contention caused by insufficient N1 memory. The new weights of 0.3 / 0.7 were pushed to the executor, and subsequent new I / O operations for this business model would allocate memory according to these new weights.

[0131] Step 7. End.

[0132] The embodiment of the present application proposes a database optimization method based on CXL, which can achieve performance optimization of the database in different business scenarios, and can solve the problems of insufficient existing databases and inability to fully utilize CXL, uneven resource allocation, and thus poor database performance. The present application realizes optimal database performance in the case of large CXL memory by redistributing memory resources, and dynamically allocating memory resources through reasonable testing and calculation. Through this method, the performance of the database can be improved, the availability of the database and CXL memory can be increased, and the difficulty of configuration can be reduced. The performance feedback and dynamic correction mechanism are added, the breadth of database application is increased, and the overall capacity of the database system is increased. This application solves the current database-related performance and database availability problems. At the same time, the present application can be applied to other related equipment to improve database performance and improve work efficiency.

[0133] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0134] The embodiment of the present application also provides a device for responding to read and write requests. Figure 5 is a response device for a read / write request according to an embodiment of the present application, such as Figure 5 As shown, the device includes:

[0135] a partitioning module 52 configured to partition the plurality of memory modules into a plurality of memory access domains according to partitioning parameters corresponding to the plurality of memory modules, wherein the plurality of memory modules are located in a computer system and the plurality of memory access domains have different access delays;

[0136] an acquisition module 54 for acquiring first performance parameters of the plurality of memory access domains within a historical time period upon receiving a read or write request to a database of the computer system;

[0137] a calculation module 56, configured to calculate weight parameters of a plurality of memory access domains of the database according to the first performance parameter;

[0138] The response module 58 is configured to determine a target memory buffer in the plurality of memory access domains according to the plurality of weight parameters, and respond to the read / write request according to the target memory buffer.

[0139] This device, through the acquisition of real-time performance parameters, calculation of multi-dimensional weight parameters, and dynamic selection of target memory buffers, ensures intelligent and efficient resource allocation, significantly improving the operational efficiency and user experience of database systems based on the CXL memory architecture. This addresses the prior art challenge of how database systems in CXL memory architecture-based server environments can effectively allocate memory resources to respond to read and write requests based on real-time performance parameters and memory access patterns, thereby maximizing performance and resource utilization.

[0140] Optionally, the above-mentioned division module 52 is also used to determine the physical distance between the multiple memory modules and the central processing unit of the computer system based on the hardware layout information and hardware specification information of the computer system, wherein the division parameters include the physical distance; and divide the multiple memory modules into the multiple memory access domains according to the physical distance intervals in which the multiple physical distances are located, wherein the multiple memory access domains correspond to multiple different physical distance intervals respectively.

[0141] Optionally, the above-mentioned division module 52 is also used to determine the interface type of the connection interface corresponding to the multiple memory modules, wherein the connection interface is used to connect the memory module to the computer system, and the division parameter includes the interface type; according to the interface type, the multiple memory modules are divided into the multiple memory access domains, wherein the multiple memory access domains correspond one-to-one to the multiple interface types.

[0142] Optionally, the response module 58 is also used to collect in real time the second performance parameter of the database in response to the read and write request in the process of responding to the read and write request, wherein the first performance parameter and the second performance parameter are of different categories; dynamically update the multiple weight parameters according to the second performance parameter, and redetermine the target memory buffer based on the multiple weight parameters after dynamic update.

[0143] Optionally, the calculation module 56 is used to calculate Calculate the weight parameters W of the multiple memory access domains i , where W i is the weight parameter of the i-th memory access domain, F i is the data access frequency of the i-th memory access domain, L i is the delay sensitivity coefficient of the i-th memory access domain, H i is the heat decay factor of the i-th memory access domain, S i is the system load rate of the i-th memory access domain, R i is the cache hit rate of the i-th memory access domain, is the sum of the data access frequencies of the multiple memory access domains, ε is a very small constant, α is an empirical coefficient, i is a positive integer, and the first performance parameter includes: the data access frequency, the delay sensitivity coefficient, the heat attenuation factor, the system load rate, and the cache hit rate.

[0144] Optionally, the acquisition module 54 is further configured to calculate the data access frequency through an operating system tool, a hardware performance counter, a CXL management interface, and a database built-in monitor, wherein the operating system tool is configured to obtain overall IO data of the database operating system, the hardware performance counter is configured to count the number of read and write operations performed on the multiple memory access domains, the CXL management interface is configured to count the access frequency of the CXL memory through the CXL switch, and the database built-in monitor is configured to collect overall IO data of the database, wherein the multiple memory access domains include the CXL memory.

[0145] Optionally, the acquisition module 54 is further configured to acquire multidimensional access data of the database in real time through the operating system tool, the hardware performance counter, the CXL management interface, and the database built-in monitoring; and calculate the multidimensional access data using an exponentially weighted moving average method within a sliding window to obtain the data access frequency.

[0146] Optionally, the acquisition module 54 is further configured to determine a service type corresponding to the read / write request, wherein the service type includes: online transaction processing and online analytical processing; and to determine the delay sensitivity coefficient according to the service type.

[0147] Optionally, the above-mentioned response module 58 is also used to determine the first memory size corresponding to the read and write request; allocate the first memory size according to multiple weight parameters to obtain multiple second memory sizes, wherein the multiple second memory sizes correspond one-to-one to the multiple memory access domains; and determine the target memory buffer in the multiple memory access domains according to the multiple second memory sizes.

[0148] Optionally, the above-mentioned response module 58 is also used to allocate sub-memory buffers in the multiple memory access domains according to the multiple second memory sizes, wherein the sub-memory buffers are continuous memory spaces; when the multiple memory access domains have successfully allocated the sub-memory buffers, the multiple sub-memory buffers are determined as the target memory buffers.

[0149] Optionally, the above-mentioned response module 58 is also used to determine the fourth memory size of the maximum continuous memory space in the mth memory access domain when a sub-memory buffer of the third memory size is not allocated in the mth memory access domain, and allocate the maximum continuous memory space as a sub-memory buffer of the mth memory access domain, where m is a positive integer; and allocate a sub-memory buffer of the fifth memory size in the m+1th memory access domain, where the fifth memory size = the third memory size - the fourth memory size + the sixth memory size, and the sixth memory size is the second memory size corresponding to the m+1th memory access domain, and the access delay of the mth memory access domain is lower than the access delay of the m+1th memory access domain.

[0150] Optionally, the above-mentioned response module 58 is also used to allocate a discrete memory buffer of the fifth memory size in the m+1th memory access domain when the sub-memory buffer of the fifth memory size is not allocated in the m+1th memory access domain, or to recalculate the weight coefficients of the multiple memory access domains and re-determine the target memory buffer in the multiple memory access domains based on the updated weight coefficients.

[0151] Optionally, the second performance parameter includes: the IO delay of the read / write request, the memory access delay corresponding to the read / write request, the CPU utilization corresponding to the read / write request, the cache hit rate corresponding to the read / write request, the bandwidth utilization corresponding to the read / write request, the memory allocation success rate corresponding to the read / write request, and the operation completion time of the read / write request.

[0152] For the description of the features in the embodiment corresponding to the device for responding to read and write requests, please refer to the relevant description of the embodiment corresponding to the method for responding to read and write requests, which will not be repeated here.

[0153] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the method for responding to read and write requests.

[0154] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned read / write request response method embodiments when running.

[0155] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0156] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned read and write request response method embodiments are implemented.

[0157] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned read and write request response method embodiments.

[0158] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0159] The above is a detailed introduction to a method and device for responding to a read / write request, an electronic device, and a computer program product provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for responding to a read / write request, characterized in that: include: Dividing the plurality of memory modules into a plurality of memory access domains according to division parameters corresponding to the plurality of memory modules, wherein the plurality of memory modules are located in a computer system and the plurality of memory access domains have different access delays; Upon receiving a read or write request to a database of the computer system, obtaining first performance parameters of the plurality of memory access domains within a historical time period; Calculating weight parameters of multiple memory access domains of the database according to the first performance parameter; determining a target memory buffer in the plurality of memory access domains according to the plurality of weight parameters, and responding to the read / write request according to the target memory buffer; Calculating weight parameters of multiple memory access domains of the database according to the first performance parameter includes: By formula Calculate the weight parameters W of the multiple memory access domains i , where W i is the weight parameter of the i-th memory access domain, F i is the data access frequency of the i-th memory access domain, L i is the delay sensitivity coefficient of the i-th memory access domain, H i is the heat decay factor of the i-th memory access domain, S i is the system load rate of the i-th memory access domain, R i is the cache hit rate of the i-th memory access domain, is the sum of the data access frequencies of the multiple memory access domains, ε is a very small constant, α is an empirical coefficient, i is a positive integer, and the first performance parameter includes: the data access frequency, the delay sensitivity coefficient, the heat attenuation factor, the system load rate, and the cache hit rate.

2. The method for responding to a read / write request according to claim 1, wherein: Dividing the multiple memory modules into multiple memory access domains according to division parameters corresponding to the multiple memory modules includes: determining a physical distance between the plurality of memory modules and a central processing unit of the computer system according to hardware layout information and hardware specification information of the computer system, wherein the partitioning parameter includes the physical distance; The multiple memory modules are divided into the multiple memory access domains according to the physical distance intervals in which the multiple physical distances are located, wherein the multiple memory access domains correspond to multiple different physical distance intervals respectively.

3. The method for responding to a read / write request according to claim 1, wherein: Dividing the multiple memory modules into multiple memory access domains according to division parameters corresponding to the multiple memory modules includes: Determining interface types of connection interfaces corresponding to the multiple memory modules, wherein the connection interfaces are used to connect the memory modules to the computer system, and the partitioning parameters include the interface types; The multiple memory modules are divided into the multiple memory access domains according to the interface types, wherein the multiple memory access domains correspond one-to-one to the multiple interface types.

4. The method for responding to a read / write request according to claim 1, wherein: Responding to the read / write request according to the target memory buffer includes: In the process of responding to the read / write request, collecting in real time a second performance parameter of the database responding to the read / write request, wherein the first performance parameter and the second performance parameter are of different categories; The plurality of weight parameters are dynamically updated according to the second performance parameter, and the target memory buffer is re-determined according to the plurality of dynamically updated weight parameters.

5. The method for responding to a read / write request according to claim 1, wherein: Determining a target memory buffer in the plurality of memory access domains according to the plurality of weight parameters includes: Determining a first memory size corresponding to the read / write request; Allocating the first memory sizes according to the multiple weight parameters to obtain multiple second memory sizes, wherein the multiple second memory sizes correspond one-to-one to the multiple memory access domains; The target memory buffer is determined in the plurality of memory access domains according to the plurality of second memory sizes.

6. The method for responding to a read / write request according to claim 5, wherein: Determining the target memory buffer in the plurality of memory access domains according to the plurality of second memory sizes includes: Allocating sub-memory buffers in the multiple memory access domains according to the multiple second memory sizes, wherein the sub-memory buffers are continuous memory spaces; In a case where the multiple memory access domains all successfully allocate the sub-memory buffers, the multiple sub-memory buffers are determined as the target memory buffers.

7. The method for responding to a read / write request according to claim 6, wherein: After allocating sub-memory buffers in the multiple memory access domains according to the multiple second memory sizes, the method further includes: If no sub-memory buffer of the third memory size is allocated in the m-th memory access domain, determining a fourth memory size of a maximum continuous memory space in the m-th memory access domain, and allocating the maximum continuous memory space as the sub-memory buffer of the m-th memory access domain, where m is a positive integer; A sub-memory buffer of a fifth memory size is allocated in the m+1th memory access domain, wherein the fifth memory size = the third memory size - the fourth memory size + the sixth memory size, the sixth memory size is the second memory size corresponding to the m+1th memory access domain, and the access delay of the mth memory access domain is lower than the access delay of the m+1th memory access domain.

8. The method for responding to a read / write request according to claim 7, wherein: After allocating the maximum continuous memory space as a sub-memory buffer of the m-th memory access domain, the method further includes: In a case where the sub-memory buffer of the fifth memory size is not allocated in the m+1th memory access domain, a discrete memory buffer of the fifth memory size is allocated in the m+1th memory access domain, or, The weight coefficients of the multiple memory access domains are recalculated, and the target memory buffer is re-determined in the multiple memory access domains according to the updated weight coefficients.

9. The method for responding to a read / write request according to claim 1, wherein: Obtaining a first performance parameter of the database within a first time period includes: The data access frequency is calculated using operating system tools, hardware performance counters, a CXL management interface, and database built-in monitoring. The operating system tools are used to obtain overall I / O data of the database operating system. The hardware performance counters are used to count the number of read and write operations performed on the multiple memory access domains. The CXL management interface is used to count the access frequency of the CXL memory through the CXL switch. The database built-in monitoring is used to collect overall I / O data of the database, where the multiple memory access domains include the CXL memory.

10. The method for responding to a read / write request according to claim 9, wherein: The data access frequency is calculated using operating system tools, hardware performance counters, CXL management interfaces, and database built-in monitoring, including: Acquiring multi-dimensional access data of the database in real time through the operating system tools, the hardware performance counters, the CXL management interface, and the database built-in monitoring; The multi-dimensional access data is calculated by using an exponentially weighted moving average method within a sliding window to obtain the data access frequency.

11. A device for responding to read and write requests, characterized in that: include: a partitioning module, configured to partition the plurality of memory modules into a plurality of memory access domains according to partitioning parameters corresponding to the plurality of memory modules, wherein the plurality of memory modules are located in a computer system, and the plurality of memory access domains have different access delays; an acquisition module, configured to acquire first performance parameters of the plurality of memory access domains within a historical time period upon receiving a read or write request to a database of the computer system; a calculation module, configured to calculate weight parameters of a plurality of memory access domains of the database according to the first performance parameter; a response module, configured to determine a target memory buffer in the plurality of memory access domains according to the plurality of weight parameters, and respond to the read / write request according to the target memory buffer; The calculation module is also used to calculate the Calculate the weight parameters W of the multiple memory access domains i , where W i is the weight parameter of the i-th memory access domain, F i is the data access frequency of the i-th memory access domain, L i is the delay sensitivity coefficient of the i-th memory access domain, H i is the heat decay factor of the i-th memory access domain, S i is the system load rate of the i-th memory access domain, R i is the cache hit rate of the i-th memory access domain, is the sum of the data access frequencies of the multiple memory access domains, ε is a very small constant, α is an empirical coefficient, i is a positive integer, and the first performance parameter includes: the data access frequency, the delay sensitivity coefficient, the heat attenuation factor, the system load rate, and the cache hit rate.

12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method according to any one of claims 1 to 10 when executing the computer program.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method according to any one of claims 1 to 10 when executed by a processor.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Memory access method, system and device

    CN118276773A