Data processing method and apparatus, and computing device

By automatically spilling cached data to external memory in a big data processing system, performance bottlenecks and OOM failures caused by insufficient memory resources are solved, and the stability and performance of the system are improved.

WO2025118665A1PCT designated stage expired Publication Date: 2025-06-12HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Patent Information

Application Number
PCT/CN2024/110363
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-19
Filing Date
2024-08-07
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Under the memory resource limitation, the big data processing system leads to high CPU vacancy rate, insufficient task concurrency, performance bottlenecks, and prone to memory overflow failures, affecting system stability.

Method used

Provide a data processing method, which automatically overflows pre-stored cached data to external memory in a server or server cluster deployed by the data processing system, avoids frequent reporting of OOM errors and improves system stability.

Benefits of technology

It effectively avoids the frequent reporting of OOM errors by data processing systems due to insufficient memory, improves the stability and performance of the system, and ensures the full utilization of task concurrency and computing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024110363_12062025_PF_FP_ABST
    Figure CN2024110363_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A data processing method and apparatus, and a computing device. The method is executed by a data processing system, after obtaining cache data, when determining that the cache data volume of the cache data is greater than the available memory space capacity of the memory of a locally deployed server cluster, the data processing system can automatically spill out part of the cache data into an external memory, so as to prevent the data processing system from frequently reporting OOM errors, thereby improving the stability of the data processing system.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method, device and computing equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 5, 2023, with application number 202311677930.6 and application name “A data processing system for internal and external memory for big data”, the Chinese patent application filed with the State Intellectual Property Office of China on February 7, 2024, with application number 202410175069.1 and application name “A data processing method, device and computing device”, and the Chinese patent application filed with the State Intellectual Property Office of China on April 19, 2024, with application number 202410496617.0 and application name “A data processing method, device and computing device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present invention relates to the field of data processing technology, and in particular to a data processing method, apparatus and computing device. Background Art

[0003] The memory wall refers to the speed mismatch between a computer system's processor and memory. As processor speeds continue to increase, processors can execute multiple instructions within a single clock cycle. However, reading data from memory requires multiple clock cycles, causing the processor's speed to exceed the memory's read speed. This prevents the processor from fully utilizing its computing power, limiting its performance.

[0004] On the one hand, due to memory resource constraints, distributed clusters in big data processing systems often experience high central processing unit (CPU) vacancy rates, resulting in insufficient task concurrency and performance bottlenecks. On the other hand, big data processing systems often experience out-of-memory (OOM) failures during computational tasks, causing task execution failures and severely impacting the stability of production operations. For example, in the Spark system, it is difficult to accurately estimate the memory usage of Spark jobs, necessitating multiple reruns to determine the required memory usage. However, the cost of Spark job failures or retries is becoming increasingly unbearable. As the volume of data in Spark production environments continues to grow, estimated memory usage quickly becomes out-of-memory, necessitating re-evaluation. Furthermore, the increasing complexity of Structured Query Language (SQL) statements in Spark production environments results in multiple execution phases and prolonged execution times, making multiple reruns increasingly expensive.

[0005] Summary of the Invention

[0006] To address the aforementioned issues, embodiments of the present application provide a data processing method. When memory resources on a server or server cluster deployed by a data processing system are insufficient, pre-stored cached data can be automatically overflowed to an external memory device. This prevents the data processing system from frequently reporting Out-of-Memory (OOM) errors, thereby improving the stability of the data processing system. Furthermore, the present application also provides a data processing device and computing equipment corresponding to the data processing method.

[0007] To this end, the following technical solutions are adopted in the embodiments of the present application:

[0008] In a first aspect, an embodiment of the present application provides a data processing method, in which a data processing system is deployed in a server cluster, and the server cluster includes multiple computing devices and at least one memory. The method is executed by any one or more computing devices among the multiple computing devices, and includes: obtaining cache data; the cache data refers to data that needs to be temporarily stored during the operation of the data processing system; when the cache data amount of the cache data is greater than the available memory space capacity of the local memory, sending part of the cache data in the cache data to the at least one memory, so that the at least one memory stores part of the cache data in the cache data.

[0009] In this embodiment, after obtaining the cached data, if the data processing system determines that the cached data volume is greater than the available memory space capacity of the memory of the locally deployed server cluster, part of the cached data can be automatically overflowed to the external memory, thereby avoiding the data processing system from frequently reporting OOM errors, thereby improving the stability of the data processing system.

[0010] In another possible embodiment, sending part of the cached data to the at least one memory specifically includes: obtaining the code of a memory allocation interface, replacing the memory allocation interface with a custom memory allocation interface; the custom memory allocation interface is used to allocate storage space in the local memory and / or the at least one memory; and storing part of the cached data in the allocated storage space in the at least one memory through the custom memory allocation interface.

[0011] In this embodiment, the data processing system can obtain code related to the memory allocation interface in the source code, replace the code related to the memory allocation interface in the source code with a custom memory allocation interface, and use a static call method to replace the memory allocation interface in the source code. The custom memory allocation interface can allocate storage space in local memory or external storage, allowing the CPU to access data in the memory space, effectively controlling the interface interception range, and avoiding the impact on other unrelated components in the data processing system.

[0012] In another possible embodiment, when the amount of cached data in the cached data is greater than the available memory space capacity of the local memory, before sending part of the cached data to the at least one memory, the method further includes: configuring a transparent memory for each executor; the computing device includes at least one executor, and the computing device uses the executor to store the cached data in the transparent memory configured by the executor, and the transparent memory includes a first set size of the available memory space capacity of the local memory and a second set size of the storage space of the at least one memory.

[0013] In this embodiment, the data processing system can configure a transparent memory including storage space of an external memory for each executor, so that when the local memory of the data processing system is insufficient, the executor can store part of the cached data in the external memory in the transparent memory, thereby automatically overflowing the cached data to the external memory.

[0014] In another possible embodiment, the storing of part of the cache data in the storage space allocated in the at least one memory through the custom memory allocation interface specifically includes: sending a memory allocation request to the at least one memory when the amount of cache data of the cache data is greater than the available memory space capacity of the local memory; the memory allocation request is used to allocate a target storage space for storing the cache data; judging through the custom memory allocation interface whether the target storage space requested to be allocated by the memory allocation request is greater than a set memory threshold; when the target storage space requested to be allocated by the memory allocation request is greater than the set memory threshold, allocating storage space in the transparent memory configured by the executor, and storing part of the cache data in the target storage space.

[0015] In another possible embodiment, the method further includes: sending a first request instruction to the at least one memory through the custom memory allocation interface; the first request instruction is used to request the at least one memory to allocate free storage space behind the target storage space.

[0016] In this embodiment, the data processing system can send a first request instruction to the external memory through a customized memory allocation interface, so that after the external memory allocates storage space according to the memory allocation request, it can allocate free storage space after the allocated storage space according to the first request instruction, so that when the data processing system allocates storage space in the external memory for the second time, it can directly use the free storage space allocated based on the first request instruction, without the need to copy the data of the memory block allocated last time, thereby realizing zero copy of cached data.

[0017] In another possible embodiment, when the amount of cached data in the cached data is greater than the available memory space capacity of the local memory, part of the cached data in the cached data is sent to the at least one external memory, specifically including: detecting the number of times each data in the cached data is accessed and / or stored; sending first cached data to the at least one external memory; the first cached data refers to data in the cached data whose number of accesses and / or storage times is less than the set number.

[0018] In this embodiment, the data processing system can detect the number of times cached data is accessed and stored, and determine the hotness or coldness of each data item in the cache. The data processing system can then place hot data in the main memory and cold data in external storage, thereby avoiding degradation of the data processing system's memory access performance.

[0019] In another possible embodiment, the method also includes: obtaining a memory access request instruction; the memory access request instruction is used to read or write specified data, and the memory access request instruction includes the virtual memory address of the specified data; querying the locally stored address mapping table to determine the physical storage address of the specified data; the address mapping table records the mapping relationship between the virtual memory address and the physical storage address; the physical storage address refers to the physical memory address of the local memory and the physical address of the at least one memory; detecting whether the physical storage address of the specified data is the physical memory address of the local memory; if the physical storage address of the specified data is the physical memory address of the local memory, reading the data stored in the physical memory address of the specified data, or writing the specified data to the corresponding physical memory address.

[0020] In this embodiment, when accessing specified data, the data processing system can search the address mapping table for the physical storage address mapped to the virtual memory address carried in the memory access request instruction. The data processing system determines that the physical storage address of the specified data is the physical memory address of the local memory and directly reads the data from the corresponding physical memory address in the local memory, thereby enabling the data processing system to successfully read data stored in the local memory.

[0021] In another possible embodiment, the method further includes: when the physical storage address of the specified data is not the physical memory address of the local memory, allocating a specified storage space in the local memory; forwarding the memory access request instruction to the at least one memory; the memory access request instruction is used to request the at least one memory to read the data stored at the physical storage address corresponding to the virtual memory address of the specified data; receiving the specified data sent by the at least one memory, and storing the specified data in the specified storage space.

[0022] In this embodiment, the data processing system determines that the physical storage address of the specified data is not the physical memory address of the local memory, and can call a cache replacement algorithm to select a cold data block from the physical memory space. The data processing system can send an IO request to the external memory to write the cold data block to the external memory, or read the accessed data to the selected cold data block location. After the data processing system determines that the IO operation is complete, it can return the physical memory address corresponding to the accessed data and record the mapping relationship between the virtual memory address and the physical memory address of the accessed data, as well as the mapping relationship between the virtual address of the cold data block and the physical address of the external memory in the address mapping table, so that the data processing system can successfully access the data stored in the external memory.

[0023] In another possible embodiment, before obtaining the memory access request instruction, the method further includes: obtaining the code of the external memory access interface, replacing the external memory access interface with a custom external memory access interface; the custom external memory access interface is used to read the data stored in the at least one memory.

[0024] In this embodiment, the data processing system can obtain code related to the external memory access interface in the source code, replace the code related to the external memory access interface in the source code with a customized external memory access interface, and replace the external memory access interface in the source code using a static call method. The customized external memory access interface is used to obtain data stored in the external memory, allowing the CPU to access the data in the storage space of the external memory.

[0025] In the second aspect, an embodiment of the present application provides a data processing device, including: a transceiver unit for obtaining cache data; the cache data refers to data that needs to be temporarily stored during the operation of the data processing system; a processing unit for sending part of the cache data to the at least one memory when the amount of cache data of the cache data is greater than the available memory space capacity of the local memory, so that the at least one memory stores part of the cache data.

[0026] In another possible embodiment, the transceiver unit is further used to obtain the code of the memory allocation interface; the processing unit is further used to replace the memory allocation interface with a custom memory allocation interface; the custom memory allocation interface is used to allocate storage space in the local memory and / or the at least one memory; and part of the cache data is stored in the storage space allocated in the at least one memory through the custom memory allocation interface.

[0027] In another possible embodiment, the processing unit is further used to configure a transparent memory for each executor; the computing device includes at least one executor, and the computing device uses the executor to store the cache data in the transparent memory configured by the executor, and the transparent memory includes the available memory space capacity of the local memory of a first set size and the storage space of the at least one memory of a second set size.

[0028] In another possible embodiment, the processing unit is specifically used to send a memory allocation request to the at least one memory when the cached data amount of the cached data is greater than the available memory space capacity of the local memory; the memory allocation request is used to allocate a target storage space for storing the cached data; determine through the customized memory allocation interface whether the target storage space requested to be allocated by the memory allocation request is greater than the set memory threshold; when the target storage space requested to be allocated by the memory allocation request is greater than the set memory threshold, allocate storage space in the transparent memory configured by the executor, and store part of the cached data in the target storage space.

[0029] In another possible embodiment, the processing unit is further used to send a first request instruction to the at least one memory through the custom memory allocation interface; the first request instruction is used to request the at least one memory to allocate free storage space behind the target storage space.

[0030] In another possible embodiment, the processing unit is specifically used to detect the number of times each data in the cache data is accessed and / or stored; and send the first cache data to the at least one memory; the first cache data refers to data in the cache data whose number of accesses and / or storage times is less than the set number.

[0031] In another possible embodiment, the processing unit is further used to obtain a memory access request instruction; the memory access request instruction is used to read or write specified data, and the memory access request instruction includes the virtual memory address of the specified data; query the locally stored address mapping table to determine the physical storage address mapped to the virtual memory address of the specified data; the address mapping table records the mapping relationship between the virtual memory address and the physical storage address; the physical storage address refers to the physical memory address of the local memory and the physical address of the at least one memory; detect whether the physical storage address of the specified data is the physical memory address of the local memory; if the physical storage address of the specified data is the physical memory address of the local memory, read the data stored in the physical memory address of the specified data, or write the specified data to the corresponding physical memory address.

[0032] In another possible embodiment, the processing unit is further used to allocate a designated storage space in the local memory when the physical storage address of the designated data is not the physical memory address of the local memory; forward the memory access request instruction to the at least one memory; the memory access request instruction is used to request the at least one memory to read the data stored at the physical storage address corresponding to the virtual memory address of the designated data; receive the designated data sent by the at least one memory, and store the designated data in the designated storage space.

[0033] In another possible embodiment, the processing unit is further used to obtain the code of the external memory access interface and replace the external memory access interface with a customized external memory access interface; the customized external memory access interface is used to read the data stored in the at least one memory.

[0034] In a third aspect, an embodiment of the present application provides a computing device, comprising: at least one memory; and at least one processor, the processor being configured to execute instructions stored in the memory so that the computing device executes various possible implementations of the first aspect.

[0035] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, comprising computer program instructions. When the computer program instructions are executed by the computing device, the computing device executes the various possible implementations of the first aspect.

[0036] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions, characterized in that the computer program product stores instructions, which, when executed by the computing device, enable the computing device to implement various possible implementation embodiments of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The following is a brief introduction to the drawings required for describing the embodiments or prior art.

[0038] FIG1 is a schematic diagram of the hardware architecture of a server cluster deployed in a data processing system provided in an embodiment of the present application;

[0039] FIG2 is a schematic diagram of the software architecture of a data processing system provided in an embodiment of the present application;

[0040] FIG3 is a flow chart of a process in which a scheduling module selects a memory allocator according to an embodiment of the present application;

[0041] FIG4 (a) is a schematic diagram of a process of extending a data block using an mremap interface in the related art;

[0042] FIG4( b) is a schematic diagram of the process of extending a data block using the mremap interface provided in an embodiment of the present application;

[0043] FIG5 is a schematic diagram of the memory of the cache layer management spark system provided in an embodiment of the present application;

[0044] FIG6 is a schematic diagram of a process for modifying the source code of ClickHouse by the interface layer of the scheduling module provided in an embodiment of the present application;

[0045] FIG7 is a schematic diagram of a process for memory allocation in a cache layer provided in an embodiment of the present application;

[0046] FIG8 is a schematic diagram of a process of querying the physical memory address of access data at the cache layer according to an embodiment of the present application;

[0047] FIG9 is a schematic diagram of a process of monitoring memory access behavior at the cache layer provided in an embodiment of the present application;

[0048] FIG10 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;

[0049] FIG11 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0050] FIG12 is a schematic diagram of the architecture of a computing device cluster provided in an embodiment of the present application;

[0051] FIG13 is a schematic diagram of the architecture of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0053] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.

[0054] The terms "first" and "second" in this specification and claims are used to distinguish different objects rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish different response messages rather than to describe a specific order of response messages.

[0055] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0056] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.

[0057] Before introducing the technical solution protected by this application, several terms that appear in the technical solution protected by this application are introduced, namely:

[0058] Spark, also known as Apache Spark, is an open-source big data processing framework. Spark is primarily used for distributed computing on large datasets, typically running in a cluster environment. Spark supports a variety of data sources, such as the Hadoop Distributed File System (HDFS), Apache Cassandra, and Apache HBase. Spark utilizes in-memory computing technology, enabling data to be loaded into memory for processing, significantly accelerating data processing.

[0059] A Spark job is the unit of execution for a Spark application, used to describe and execute data processing and analysis tasks. Spark jobs leverage Spark's distributed computing capabilities and optimization strategies to achieve high-performance, high-concurrency data processing.

[0060] Off-heap memory is a way of allocating memory outside the Java heap. In Java applications, the Java heap can be used to allocate memory space for objects, while off-heap memory is allocated in an unmanaged memory area outside the Java heap.

[0061] The natural language acceleration engine (NLE) vectorized acceleration system refers to a system that uses the NLE and vectorization technology to improve the performance of intensive applications. The NLE utilizes a specific vectorized instruction set and optimization algorithms to break down computing tasks into multiple subtasks and use vector registers to process multiple subtasks simultaneously.

[0062] A data processing system refers to a software system or collection of tools used to process massive amounts of data. Data processing systems efficiently store, process, analyze, and manage large amounts of structured, semi-structured, and unstructured data to extract valuable information and insights. Common data processing systems include Apache Hadoop, Apache Spark, and Apache Flink, providing core components such as distributed storage, distributed computing frameworks, and data processing engines to support large-scale data processing and analysis. Data processing systems can be deployed on large-scale server clusters to complete data processing tasks, providing high-performance, high-reliability, and high-scalability data processing capabilities.

[0063] When processing large amounts of data, data processing systems need to utilize the memory resources of server clusters to store cached data. As the amount of cached data increases, the memory resources required also increase. If the data processing system cannot effectively manage and process this large amount of data, memory resources will become strained. If the memory resources cannot accommodate all the cached data, the data processing system will frequently report Out-of-Memory (OOM) errors, reducing system stability.

[0064] In order to solve the defects existing in the related technology, an embodiment of the present application provides a data processing system. When the memory resources of the server or server cluster deployed by the data processing system are insufficient, the pre-stored cache data can be automatically overflowed to the external memory, which can avoid the data processing system from frequently reporting OOM errors and improve the stability of the data processing system.

[0065] Figure 1 is a schematic diagram of the hardware architecture of a server cluster deployed in a data processing system according to an embodiment of the present application. As shown in Figure 1, server cluster 100 may include multiple computing devices 110, memory 120, and a communication bus 130. Multiple computing devices 110 may include computing device 110-1, computing device 110-2, computing device 110-3, and so on. Multiple computing devices 110 and memory 120 may establish communication connections via communication bus 130 for data transmission.

[0066] Computing device 110 may be a server, computer, smartphone, or other device. Multiple computing devices 110 may be a single device or a combination of multiple devices. Multiple computing devices 110 may be deployed with applications that run a data processing system to perform data processing system functions, such as data cleansing, data storage and management, data conversion and integration, data analysis and mining, real-time data processing, and security and privacy protection.

[0067] The computing device 110 may be provided with a storage unit, which may be a cache, dynamic random access memory (DRAM), static random access memory (SRAM), or other memory for temporarily storing data, and is used to store data that needs to be temporarily stored during the operation of the data processing system for access or operation by the processor or other hardware. In an embodiment of the present application, the storage unit may store data corresponding to computing tasks to be executed by the data processing system, data corresponding to computing tasks that have not been executed in interrupted computing tasks, intermediate data during the processing process, processed data, and other cached data.

[0068] The computing device 110 may be provided with a communication interface that may be connected to the communication bus 130 to enable data transmission with other computing devices, storage devices, and other devices. The communication interface may be a wired transmission interface, such as a Compute Express Link (CXL) interface, a Peripheral Component Interconnect Express (PCIe) interface, a Universal Serial Bus (USB) interface, etc. The communication interface may be a wireless transmission interface, such as a Bluetooth (BT) module, a Wireless Fidelity (WI-FI) module, a wireless communication module, etc.

[0069] The memory 120 can be an independent storage device or a memory of the computing device 110. The memory 120 can be a hard disk drive (HDD), a solid state drive (SSD), a NAND flash memory, a magnetic disk, etc., and is used to store cache data overflowed when multiple computing devices 110 run a data processing system, as well as other data, such as programs for running the data processing system. It should be noted that the connection between the memory 120 and the computing device 110-3 in FIG1 is not limited to the memory 120 establishing a communication connection with the computing device 110-3, but also means that the memory 120 establishes a communication connection with the data processing system deployed by the computing device 110-3, so that the memory 120 can establish a communication connection with any computing device 110.

[0070] The memory 120 may be provided with a communication interface that can be connected to the communication bus 130 to enable data transmission with multiple computing devices 110 and other devices. The communication interface may be a wired transmission interface, such as a CXL interface, a PCIe interface, a USB interface, etc. The communication interface may also be a wireless transmission interface, such as a BT module, a Wi-Fi module, a wireless communication module, etc.

[0071] The communication bus 130 may refer to communication hardware such as cables, optical fibers, routers, base stations, etc., which are respectively connected to the communication interface of the computing device 110 and the communication interface of the memory 120, so that a communication connection is established between multiple computing devices 110 and the memory 120 for data forwarding.

[0072] Figure 2 is a schematic diagram of the software architecture of a data processing system provided in an embodiment of the present application. As shown in Figure 2, the data processing system 200 can be divided into a front-end 210, a back-end 220 and an operating system (OS) 230 from a software perspective.

[0073] The front-end 210 is responsible for the development of user interfaces and applications. In the embodiment of the present application, in the data processing system 200, the front-end 210 can be written in a development language such as Java or Scala. The front-end 210 can input data, query and display results, as well as perform scheduling and monitoring functions.

[0074] The back end 220 is the core part of the data processing system 200, which is responsible for the processing, calculation and analysis of data. The back end 220 can be written in natural language (native code) (such as C language, C++ language, assembly language or other similar languages). The back end 220 is written in Java language, which can ensure the ease of use of the data processing system 200. The back end 220 is written in C++ language, which can improve the efficiency of executing instructions of the data processing system 200. Natural language generally refers to C language, C++ language, assembly language or other similar languages, which can interact with computer hardware and OS230 more directly. In an embodiment of the present application, the back end 220 can compile the C programming language (C programming language, C) library and / or C++ library into a shared library and load it together with libjvm.so to call the memory allocation interface in the C library to allocate storage space of the external memory. The memory allocation interface may include a memory allocation (malloc) interface, a clear allocation (calloc) interface, a reallocation (reallocate) interface, a memory mapped file (mmap) interface, and other interfaces. The external memory read / write interface may include a read interface, a write interface, and the like.

[0075] OS 230 is the underlying foundation of data processing system 200 and is responsible for managing and controlling hardware resources. OS 230 can provide basic functions such as process scheduling, memory management, file system and network connections, so that front-end 210 and back-end 220 can run correctly and work together on the hardware.

[0076] In the embodiment of the present application, the data processing system 200 can add a scheduling module 240 between the OS 230 and the front end 210 and the back end 220. From a software perspective, the scheduling module 240 can be divided into an interface layer 241, a cache layer 242 and a driver layer 243.

[0077] The interface layer 241 is mainly responsible for interfacing with the C library interface called by the underlying native code of the backend 220. In the embodiment of the present application, after the interface layer 241 is connected to the interface of the backend 220, a custom memory allocation interface 2411, a custom external memory access interface 2412 and a custom memory access interface 2413 are generated.

[0078] Typically, the data processing system 200 has a memory allocation interface, which may be a malloc interface, a calloc interface, a realloc interface, a new interface, a new[] interface, an mmap interface, a memory remap (mremap) interface, a delete interface, a delete[] interface, a free interface, and the like.

[0079] The Malloc interface is used to dynamically allocate a memory block of a specified size and return a pointer to the allocated memory. The Calloc interface is used to dynamically allocate a memory block of a specified number and size and initialize the allocated memory block to zero. The Realloc interface is used to reallocate the size of an allocated memory block, either increasing or decreasing it. The New interface is used to dynamically allocate memory for a single object and call its constructor for initialization. The New interface returns a pointer to the newly created object. The New[] interface, also used in C++, is used to dynamically allocate memory for an array of objects and call each object's constructor for initialization. The Mmap interface is a system-level interface used to map files or allocate anonymous memory areas in memory. The Mremap interface is used to remap an allocated memory area in memory, adjusting its size and position. The Delete interface is used to free the memory of a single object created using the new operator and call the object's destructor for cleanup. The Delete[] interface is used to free the memory of an array of objects created using the new[] operator and call each object's destructor for cleanup. The Free interface is used to free a memory block allocated using the malloc, calloc, or realloc interfaces.

[0080] In an embodiment of the present application, the interface layer 241 can obtain the code about the memory allocation interface in the source code, replace the code about the memory allocation interface in the source code with a custom memory allocation interface 2411, and use a static call method to replace the memory allocation interface in the source code. The custom memory allocation interface 2411 is used to allocate storage space in local memory or external memory so that the CPU can access the data in the memory space. For example, the malloc interface is replaced with a transparent memory allocation (transparent_malloc) interface. The interface layer 241 uses a static call method and does not use a dynamic interception method for replacement, which can effectively manage and control the interface interception range and avoid affecting other unrelated components in the data processing system 200. The source code refers to the code corresponding to the application program running the data processing system 200.

[0081] The interface layer 241 uses low-level virtual machine (LLVM) plug-in technology to replace all memory access instructions (such as load instructions, store instructions, memmove instructions, etc.) in the source code during the code compilation stage to achieve application transparency.

[0082] Typically, most memory allocation requests in the source code are for allocating temporary data. Temporary data occupies a relatively small amount of memory and is released very quickly after use. If temporary data is overflowed to the external memory, the performance of the data processing system 200 will be lost. In addition, a memory allocator is integrated into the shared library to intercept the memory allocation interface in the program code that calls the shared library. Since the memory access instructions in the program code that calls the shared library have not been replaced, errors will occur in the application program. Therefore, the interface layer 241 cannot be compatible with the Java virtual machine (JVM) plug-in mechanism and cannot run simultaneously with other memory allocators in the data processing system 200.

[0083] To address the flaws in the interface layer 241's use of LLVM plug-in technology, the interface layer 241 can identify memory allocation interfaces in the source code that are prone to OOMs and memory access instructions in related code, referred to as "target memory allocation interfaces" and "target memory access instructions." Target memory allocation interfaces typically allocate large blocks of memory that are used for long periods of time. The interface layer 241 can then use automated analysis tools to identify target memory allocation interfaces and target memory access instructions in the Spark system source code that are prone to OOMs. The target memory allocation interfaces are then replaced with custom memory allocation interfaces 2411. The interface layer 241 can use LLVM plug-in technology to only replace target memory access requests in related source code snippets. The interface layer 241 can overflow memory allocation requests for large blocks of memory to external storage, preventing performance losses in the data processing system 200. The interface layer 241 uses a static call algorithm to replace memory allocation interfaces that request large blocks of memory, without intercepting calls to memory allocation interfaces of shared library programs and causing errors in the running of related applications.

[0084] Typically, the data processing system 200 has an external memory access interface, which may be a read interface, a write interface, and the like.

[0085] In an embodiment of the present application, the interface layer 241 can obtain code related to the external memory access interface in the source code, replace the code related to the external memory access interface in the source code with a custom external memory access interface 2412, and use a static call method to replace the external memory access interface in the source code. The custom external memory access interface 2412 is used to obtain data stored in the external memory, allowing the CPU to access the data in the storage space of the external memory. For example, the read interface is replaced with a transparent read (transparent_read) interface.

[0086] Typically, the data processing system 200 has a memory access interface, which may be a load interface, a store interface, a memory copy (memcpy) interface, a memory move (memove) interface, a string copy (strcpy) interface, a string concatenation (strcat) interface, etc.

[0087] The Load interface is used to read data from memory into a specified register or variable, typically implemented by instructions or functions similar to ld or load. The Store interface is used to write data from a register or variable to a specified location in memory, typically implemented by instructions or functions similar to st or store. The Memcpy interface is used to copy data from one block of memory to another. It is typically used to copy data between memory blocks, providing efficient and reliable memory copy functionality. The Memmove interface is similar to the Memcpy interface, used to copy data between memory blocks, but is more secure when handling overlapping memory, and can handle cases where the source and destination memory blocks partially or completely overlap. The Strcpy interface is used to copy a string from source memory to destination memory until a null character '\0' is encountered. To use the strcpy interface, it is generally necessary to ensure that the destination memory has sufficient space to accommodate the copied string. The Strcat interface is used to append one string to the end of another, and it is also necessary to ensure that the destination string has sufficient space to accommodate the appended result.

[0088] In an embodiment of the present application, the interface layer 241 can obtain the code related to the memory access interface in the source code, compile the code related to the memory access interface in the source code into target code, and replace the memory access interface in the source code with a static call method to obtain a customized memory access interface 2413. The customized memory access interface 2413 is used to access data in local memory and external storage. For example, the load interface is replaced with a transparent load (transparent_load) interface.

[0089] Traditional memory allocation interfaces (such as malloc and mmap) generally allocate a contiguous block of virtual memory addresses from the user's virtual memory space. The CPU then uses memory access instructions (such as load and store) to access the data in the virtual memory addresses. However, under the control and management of the OS and hardware, virtual memory addresses are automatically converted to physical memory addresses through address mapping in the translation lookaside buffer (TLB) or page table entry (PTE), allowing the CPU to retrieve cached data from physical memory.

[0090] In the embodiment of the present application, the custom memory access interface 2413 can access data in local memory and data in external memory. The custom memory access interface 2413 can map virtual memory addresses to physical memory and to external memory, so that the CPU can use the custom memory access interface 2413 to obtain data from local memory and obtain data from external memory.

[0091] During the compilation phase, the interface layer 241 can use LLVM static calls to replace the memory allocation interface 2411. Then, the interface layer 241 uses automated analysis tools to accurately control the replacement range of memory allocation requests (such as load instructions, store instructions, etc.). The interface layer 241 detects whether the upstream memory allocation interface has been replaced. After the upstream memory allocation interface has been replaced, it can detect the memory access instructions related to the replaced memory allocation interface and only replace these related memory access instructions with the custom memory access interface 2413. Irrelevant memory access instructions are not replaced, thereby avoiding serious performance losses.

[0092] Because the memory allocator in the related art may have problems such as slow allocation speed, many fragments, and memory leaks, the interface layer 241 can use lightweight multi-memory allocator management technology and be deployed locally. After the customized memory allocation interface 2411 receives the memory allocation request, the interface layer 241 can first utilize the cache management algorithm to process it, and then choose to use a user-specified memory allocator to perform memory allocation. Current mainstream memory allocators include thread-specific memory allocators (per thread memory allocator, PTMALLOC), Jemma memory allocator (Jason Evans memory allocator, JEMALLOC), Microsoft memory allocator (Microsoft allocator, MIMALLOC), etc. Different types of memory allocators have their own advantages in different scenarios, so the interface layer 241 can flexibly select different types of memory allocators according to demand.

[0093] For example, FIG3 is a flow chart of the process of selecting a memory allocator by the scheduling module provided in an embodiment of the present application. As shown in FIG3 , the process of selecting a memory allocator is performed by the interface layer 241, and the specific implementation process is as follows:

[0094] In step S301 , the interface layer 241 calls the user-defined memory allocation interface 2411 and receives a memory allocation request through the user-defined memory allocation interface 2411 .

[0095] In step S302 , the interface layer 241 imports “determination conditions” and “selected memory allocator” from the configuration parameters before the application is started.

[0096] In step S303, the interface layer 241 determines whether the memory allocation request meets the determination condition. In one case, the interface layer 241 determines that the memory allocation request does not meet the determination condition, and the process proceeds to step S304. In another case, the interface layer 241 determines that the memory allocation request meets the determination condition, and the process proceeds to step S305. The determination condition may be that the memory block requested for allocation by the memory allocation request is larger than the set memory size.

[0097] In step S304 , the interface layer 241 puts the memory allocation request into the OS memory management mechanism for management.

[0098] In step S305 , the interface layer 241 places the memory allocation request into the transparent memory for management.

[0099] Transparent memory refers to an unused virtual memory space reserved by the operating system. Taking the Spark system as an example, each computing device 110 can run at least one executor, and computing device 110 can store data in the transparent memory configured for the executor. Transparent memory includes the available memory space of local memory of a first set size and the storage space of external memory of a second set size.

[0100] For example, using the Spark data processing system 200 as an example, when the Spark system is running, the interface layer 241 can set a determination condition that the memory block requested by the memory allocation request is larger than 1MB. If the interface layer 241 determines that the memory block requested by the memory allocation request is less than or equal to 1MB, the memory allocation request can be placed in the OS memory management mechanism for management. If the interface layer 241 determines that the memory block requested by the memory allocation request is larger than 1MB, the memory allocation request can be placed in transparent memory for management.

[0101] In step S306, the interface layer 241 selects a corresponding memory allocator. The memory allocators selected by the interface layer 241 may include PTMALLOC, JEMALLOC, MIMALLOC, etc. For example, if the interface layer 241 determines that the memory block requested by the memory allocation request is larger than 1MB, JEMALLOC may be selected as the memory allocator.

[0102] Normally, data processing system 200 carries out mremap and realloc operation to virtual memory space, can cause the copy of buffered data.Therefore, interface layer 241 can static call self-defining memory allocation interface 2411, as transparent_mremap interface, transparent_realloc interface etc., and self-defining memory allocation interface 2411 is deployed in this locality.Interface layer 241 can utilize self-defining memory allocation interface 2411, when distributing for the first time, reserve enough virtual memory spaces, be used for future expansion.When interface layer 241 distributes again at subsequent moment, after receiving mremap instruction, can expand in reserved virtual memory space, the memory block that distributed last time is expanded into the memory block of larger internal memory, do not need the data of the memory block that distributed last time are copied, realize the zero copy of buffered data.Zero copy is not needed the data in the original memory block to be copied when being meant the extended memory block.

[0103] In one embodiment, shown in Fig. 4 (a), the mremap interface of the correlation technology carries out the mmap operation at the T1 moment, and in the continuous distribution multiple memory blocks of the virtual memory space, the size of each memory block is 16KB.At the T2 moment, when the mremap interface of the correlation technology carries out the mremap operation, a plurality of memory blocks need be expanded into the memory blocks of 2 times of sizes.At this moment, the mremap interface of the correlation technology distributes a 32KB memory block in the virtual memory space behind a plurality of memory blocks, and then copies the data of the memory block inside of first 16KB, and the data after the copying are stored in the memory block of 32KB.

[0104] As shown in Figure 4 (b), the application's customized mremap interface carries out the mmap operation at the T1 moment, can allocate multiple memory blocks in the virtual memory space, and reserve the free storage space of setting size between each memory block.The reserved free storage space is greater than 16KB.At the T2 moment, when the application's customized mremap interface carries out the mremap operation, it is necessary that multiple memory blocks be expanded into the memory blocks of 2 times of size.At this moment, the application's customized mremap interface can allocate a 16KB storage space in the reserved storage space between two memory blocks, and merge this storage space with first memory block, realize that the size of first memory block is expanded into 32KB by 16KB.

[0105] Optionally, the customized memory allocation interface 2413 may send a first request instruction to the external memory. After allocating storage space according to the memory allocation request, the external memory may allocate free storage space after the allocated storage space according to the first request instruction.

[0106] The cache layer 242 can manage local memory and external memory. Taking the data processing system 200 as a spark system as an example, as shown in Figure 5, when the CPU executes a spark job, it specifies a certain amount of off-heap memory space for each executor, and then allocates a portion of memory space from the off-heap memory space. Each executor can combine the allocated local memory storage space and the external memory storage space to form a transparent memory. Among them, each computing device 110 can run at least one executor, and the computing device 110 can store data in the transparent memory configured by the executor. Local memory refers to the storage unit of each computing device 110, and external memory refers to the storage unit of the memory 120. Transparent memory can record the memory access frequency of data in the cache and manage hot and cold data in the data.

[0107] The cache layer 242 can store some data in the external storage within the transparent memory to avoid frequent OOM errors when the memory resources of the multiple computing devices 110 deployed in the data processing system 200 are insufficient. Preferably, the cache layer 242 can detect the hotness of the cached data and place hot data in the local memory and cold data in the external storage to avoid reducing the memory access performance of the data processing system 200.

[0108] The cache layer 242 can manage the mapping of virtual memory addresses to physical memory addresses. The cache layer 242 can store an address mapping table locally, and the address mapping table records the mapping relationship between the virtual memory address and the physical storage address. The address mapping table in the related art can only map the virtual memory address to the physical memory address, and cannot be mapped to the external storage. The address mapping table protected by this application can not only map the virtual memory address to the local physical memory address, but also map it to the physical address of the external memory. The physical storage address refers to the physical memory address of the local memory and the physical address of the external memory.

[0109] Taking the Spark system as an example, as shown in Figure 5 , when the CPU executes a Spark job, it can call the custom memory allocation interface 2411 provided by the interface layer 241 to allocate a virtual memory address from the storage space of the transparent memory. When the data storage system 200 executes an application, the cache layer 242 can use the custom memory access interface 2413 to access data in the transparent memory. For example, when the CPU uses a load instruction or a store instruction to access data in the memory, it can use the custom memory access interface 2413 instead of the load instruction or the store instruction.

[0110] After receiving the memory access request instruction, the custom memory access interface 2413 can parse out the virtual memory address carried by the memory access request instruction. The memory access request instruction is used to read the specified data, and the virtual memory address of the memory access request instruction refers to the virtual address of the accessed data. The custom memory access interface 2413 queries the address mapping table of the local storage to obtain the physical storage address mapped by the virtual memory address. The custom memory access interface 2413 can detect whether the physical storage address of the specified data is the physical memory address of the local memory. In one case, the custom memory access interface 2413 determines that the physical storage address of the specified data is the physical memory address of the local memory, and returns the physical memory address to the cache layer 242.

[0111] In another case, the custom memory access interface 2413 determines that the physical storage address of the specified data is not the physical memory address of the local memory, and can call the cache replacement algorithm to allocate a storage space of a set size in the local memory, and forward the memory access request instruction to the external memory. After the external memory obtains the virtual memory address of the memory access request instruction, it sends the data stored in the physical storage address corresponding to the virtual memory address of the memory access request instruction to a certain computing device 110, allowing the computing device 110 to cache the read data in the local memory allocated to the storage space. After the custom memory access interface 2413 determines that the external memory caches the accessed data in the local memory allocated to the storage space, it returns the physical memory address cached in the local memory to the cache layer 242. The custom memory access interface 2413 can record the mapping relationship between the virtual memory address of the accessed data and the physical memory address corresponding to the storage space allocated to the local memory in the address mapping table.

[0112] The cache layer 242 can monitor memory access behavior. The cache layer 242 can set a counter in the custom memory access interface 2413 to record the number of times each memory block is accessed during application execution, classifying different memory blocks as hot and cold, which serves as a basis for cache replacement. The cache layer 242 can also record the maximum amount of storage space overflowed into external memory. By adding the maximum external memory storage space to the off-heap local memory, the cache layer 242 can accurately estimate the memory usage required for task execution. The cache layer 242 utilizes the custom memory allocation interface 2411 in the interface layer 241 to easily transmit software information. Combined with monitoring the memory access behavior of the custom memory access interface 2413, the cache layer 242 can track the object memory access behavior of Spark jobs at the software level. Using analysis algorithms, the cache layer 242 can easily summarize the memory access characteristics of Spark jobs, such as streaming access, cyclic access, and random access. Using these summarized access characteristics, the cache layer 242 can accurately predict the future and adaptively adjust the cache management algorithm. The cache layer 242 can obtain an optimal cache hit rate to improve the performance of the entire cache layer 242.

[0113] The access behavior of an application program at each stage is constantly changing, so there are random access, streaming access, and cyclic access. If the access behavior of an application program changes and the cache management algorithm is not adjusted accordingly, the cache hit rate will decrease, and the performance of the data processing system 200 will decrease. The cache layer 242 tracks the memory access behavior of the application program and can observe and predict changes in the memory access behavior of the application program. For example, if the cache layer 242 determines that the memory access behavior of the application program has changed from random access to streaming access, it can adjust the cache management algorithm to ensure that the cache hit rate always remains at a high level, thereby avoiding loss of memory access performance of the data processing system 200. The data processing system of the related art can only try to increase the memory usage after the spark job encounters OOM failure by rerunning it multiple times. And after the spark job of the related art succeeds, it can try to reduce the memory usage. After repeated retries, the data processing system of the related art can estimate the accurate memory usage for a spark job. The cache layer 242 in the embodiment of the present application records the maximum value of the storage space of the data overflow to the external memory, so that the accurate memory usage can be estimated with only one successful run, reducing the time cost of manual tuning.

[0114] Driver layer 243 can deploy OS 230's user-mode input / output (IO) drivers, such as asynchronous IO (AIO), IO ring (IOUring), and storage performance development kit (SPDK). After OS 230's user-mode completes most IO operations through driver layer 243, the number of user-mode and kernel-mode switches in a traditional IO stack can be reduced, avoiding the performance overhead caused by frequent context switching.

[0115] The following takes the data processing system 200 being a spark system as an example to introduce the implementation process of the technical solution protected by this application.

[0116] The technical solution protected by this application can be applied to the spark system, and can be applied to other systems with insufficient stability due to the presence of OOM. When the Spark system executes a spark job, when memory resources are insufficient, the data cached in the memory can be automatically overflowed to the external memory. Compared with independent operation of the memory, the data cached in the memory automatically overflows to the external memory, and the performance of the data processing system 200 is reduced. When the data processing system 200 needs to use the external memory, it is necessary to configure the parameters of the external memory in advance, such as specifying the storage directory of the external memory. The cache layer 242 in the big data storage system 200 counts some advanced memory optimization features, such as lightweight multi-memory allocator management, custom mremap interface, cache replacement algorithm, data prefetching algorithm, etc., and it is necessary to set the configuration parameters in advance.

[0117] The technical solution protected by this application can be deployed in a natural language acceleration engine vectorization acceleration system for the spark system. The natural language acceleration engine vectorization acceleration system can be an open source tool used in combination with gluten and clickhouse. Gluten is a Java-based middleware for streaming and processing data. Clickhouse is a column-oriented distributed database management system with high scalability and high performance, which can better handle massive amounts of data. The natural language acceleration engine vectorization acceleration system combines gluten and clickhouse. Data can be loaded into memory through gluten, and then clickhouse can be used for efficient data query and processing, which can improve response time, processing efficiency, and reduce processing costs and energy consumption.

[0118] Compared to the related art Spark system, which manages memory space within the JVM, the Natural Language Accelerator vectorized acceleration system utilizes a Spark plugin mechanism. Gluten intercepts Spark query plans in the middle layer, and sends them to the lower-level vectorized engine, ClickHouse, which then executes the query tasks, bypassing the inefficient execution path of the native Spark system. Gluten calls shared libraries via the Java Native Interface (JNI), directly calling natural language (native) code within the Spark executor task thread, eliminating the need for a complex threading model.

[0119] For example, Figure 6 is a flow chart of how the interface layer of the scheduling module provided in an embodiment of the present application modifies the source code of ClickHouse. As shown in Figure 6, this process is performed by the interface layer 241 of the scheduling module 240, and the specific implementation process is as follows:

[0120] Step S601, the interface layer 241 obtains the source code of clickhouse.

[0121] In step S602 , the interface layer 241 uses the user-defined memory allocation interface 2411 to replace the memory allocation interface in the C library.

[0122] In step S603 , the interface layer 241 determines upstream and downstream dependent libraries of the customized memory allocation interface 2411 .

[0123] In step S604 , the interface layer 241 replaces the memory access interface in the affected dependent library during the compilation phase.

[0124] Step S605 : The interface layer 241 generates a shared library libch.so.

[0125] In an embodiment of the present application, the interface layer 241 of the scheduling module 240 is connected to the vectorization engine clickhouse, and can be used as a component of clickhouse to participate in compilation. The interface layer 241 can obtain the source code of clickhouse to call the interface of the C library. The interface layer 241 can use the customized memory allocation interface 2411 to replace the memory allocation interface in the C library, and replace it with the malloc interface, calloc interface, realloc interface, mmap interface, mremap interface, new interface, new[] interface, delete interface, delete[] interface, free interface or other interfaces. The interface layer 241 uses automated analysis tools to clarify the upstream and downstream libraries of the customized memory allocation interface 2411. When the interface layer 241 uses the LLVM plug-in for the compilation stage, it can replace the memory allocation requests in the dependent libraries affected by the compilation, such as load, store, memcpy, memset and other instructions. After compilation, the interface layer 241 generates a shared libch.so and provides it to gluten for loading and use.

[0126] For example, FIG7 is a flow chart of memory allocation performed by the cache layer according to an embodiment of the present application. As shown in FIG7 , the process is performed by the cache layer 242 of the scheduling module 240 , and the specific implementation process is as follows:

[0127] In step S701 , the cache layer 242 uses the JNI of the JVM to call a shared database.

[0128] In step S702 , the cache layer 242 obtains the natural language code in the shared database.

[0129] In step S703 , the cache layer 242 uses the customized memory allocation interface 2411 to obtain the memory corresponding to the memory allocation request.

[0130] In step S704, the cache layer 242 detects whether the memory corresponding to the obtained memory allocation request is greater than the determination condition. In one case, the cache layer 242 determines that the memory corresponding to the obtained memory allocation request is greater than the determination condition, and then executes step S705. In another case, the cache layer 242 determines that the memory corresponding to the obtained memory allocation request is less than or equal to the determination condition, and then executes step S706.

[0131] Step S705 : The cache layer 242 obtains the user's virtual memory space.

[0132] In step S706 , the cache layer 242 obtains the virtual memory space of the transparent memory.

[0133] In the embodiment of the present application, the cache layer 242 is responsible for processing the allocation request issued by the interface layer 241. Before executing the spark job, the CPU will set a judgment condition. The judgment condition can be set to place the memory allocation request exceeding 1MB into the transparent memory address space for management. After receiving the allocation request issued by the interface layer 241, the cache layer 242 will make a selection based on the judgment condition. In one case, the cache layer 242 determines that the allocation request does not meet the judgment condition and can be managed in the traditional user's virtual memory space. In another case, the cache layer 242 determines that the allocation request meets the judgment condition and can be placed in the virtual memory space of the transparent memory for management.

[0134] For example, FIG8 is a flow chart of the cache layer querying the physical memory address of accessed data provided by an embodiment of the present application. As shown in FIG8 , this process is executed by the cache layer 242 of the scheduling module 240 , and the specific implementation process is as follows:

[0135] In step S801 , the cache layer 242 receives a memory access request instruction.

[0136] In step S802 , the cache layer 242 queries the address mapping table after determining that the customized memory access interface 2413 receives a memory access request instruction.

[0137] In step S803, the cache layer 242 determines whether the physical storage address is a physical memory address of the local memory. In one case, the cache layer 242 determines that the physical storage address is not a physical memory address of the local memory, and the process proceeds to step S804. In another case, the cache layer 242 determines that the physical storage address is a physical memory address of the local memory, and the process proceeds to step S805.

[0138] In step S804 , the cache layer 242 uses a cache replacement algorithm to swap the data in the external memory into the local memory through the driver layer 243 .

[0139] Step S805: The cache layer 242 returns the physical memory address.

[0140] In an embodiment of the present application, the cache layer 242 is responsible for processing the memory access request instruction issued by the interface layer 220. After the cache layer 242 determines that the custom memory access interface 2413 receives the memory access request instruction, it will query the address mapping table to obtain the physical storage address mapped to the virtual memory address. The custom memory access interface 2413 can detect whether the physical storage address of the specified data is the physical memory address of the local memory. In one case, the custom memory access interface 2413 determines that the physical storage address of the accessed data is the physical memory address of the local memory, and returns the physical memory address to the cache layer 242.

[0141] In another case, if the custom memory access interface 2413 determines that the physical storage address of the specified data is not the physical memory address of the local memory, it can call the cache replacement algorithm to select a cold data block from the physical memory space. The cache layer 242 can send an IO request to the external memory to write the cold data block to the external memory, or read the accessed data to the selected cold data block location. After the cache layer 242 determines that the IO operation is complete, it can return the physical memory address corresponding to the accessed data and record the mapping relationship between the virtual memory address and the physical memory address of the accessed data, as well as the mapping relationship between the virtual address of the cold data block and the physical address of the external memory in the address mapping table, so that the data processing system can successfully access the data stored in the external memory.

[0142] For example, FIG9 is a flow chart of the cache layer monitoring memory access behavior provided by an embodiment of the present application. As shown in FIG9 , the process is performed by the cache layer 242 of the scheduling module 240, and the specific implementation process is as follows:

[0143] In step S901 , the cache layer 242 creates an object information table.

[0144] In step S902 , the cache layer 242 analyzes memory access characteristics in real time according to the object information table.

[0145] In step S903 , the cache layer 242 selects an appropriate management strategy based on the memory access characteristics.

[0146] In an embodiment of the present application, the cache layer 242 is responsible for monitoring the memory access behavior of the spark job. After the cache layer 242 determines that the custom memory allocation interface 2411 receives a memory allocation request, it can create a record in the object information table. The record includes information such as the size of the allocated memory, the physical memory address of the allocated memory, and the memory access frequency. When the cache layer 242 detects that the custom memory access interface 2413 is called, it can update the memory access frequency information in the object information table in real time. The cache layer 242 can use a local built-in analysis algorithm to analyze memory access characteristics in real time, such as streaming access, sequential access, random access, etc. The cache layer 242 can select appropriate cache replacement algorithms and data prefetching algorithms based on the memory access characteristics.

[0147] FIG10 is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. As shown in FIG10 , the data processing device 1000 can be divided into a transceiver unit 1010 and a processing unit 1020 according to the execution function. The specific implementation process of the data processing device 1000 is as follows:

[0148] The transceiver unit 1010 is configured to obtain cached data. Cache data refers to data that needs to be temporarily stored during the operation of the data processing system. The processing unit 1020 is configured to send a portion of the cached data to at least one memory, if the cached data volume exceeds the available memory space capacity of the local memory, so that the at least one memory stores the portion of the cached data.

[0149] In another possible embodiment, the transceiver unit 1010 is further configured to obtain a memory allocation interface code. The processing unit 1020 is further configured to replace the memory allocation interface with a custom memory allocation interface. The custom memory allocation interface is configured to allocate storage space in the local memory and / or at least one memory. The processing unit 1020 is further configured to store a portion of the cached data in the allocated storage space in the at least one memory using the custom memory allocation interface.

[0150] In another possible embodiment, the processing unit 1020 is further configured to configure a transparent memory for each executor. The computing device includes at least one executor. The computing device uses the executor to store cache data in the transparent memory configured for the executor. The transparent memory includes available memory space capacity of a local memory of a first set size and storage space of at least one memory of a second set size.

[0151] In another possible embodiment, the processing unit 1020 is specifically configured to send a memory allocation request to at least one memory when the amount of cached data is greater than the available memory space capacity of the local memory. The memory allocation request is used to allocate a target storage space for storing the cached data. The processing unit 1020 is specifically configured to determine, through a custom memory allocation interface, whether the target storage space requested to be allocated by the memory allocation request is greater than a set memory threshold. The processing unit 1020 is specifically configured to allocate storage space in the transparent memory configured by the executor when the target storage space requested to be allocated by the memory allocation request is greater than the set memory threshold, and store part of the cached data in the target storage space.

[0152] In another possible embodiment, the processing unit 1020 is further configured to send a first request instruction to the at least one memory through a custom memory allocation interface. The first request instruction is configured to request the at least one memory to allocate free storage space after the target storage space.

[0153] In another possible embodiment, processing unit 1020 is specifically configured to detect the number of times each data item in the cached data has been accessed and / or stored. Processing unit 1020 is specifically configured to send first cached data to at least one memory. The first cached data refers to data in the cached data that has been accessed and / or stored less than a set number of times.

[0154] In another possible embodiment, the processing unit 1020 is also used to obtain a memory access request instruction. The memory access request instruction is used to read or write specified data, and the memory access request instruction includes the virtual memory address of the specified data. The processing unit 1020 is also used to query the address mapping table stored locally to determine the physical storage address mapped to the virtual memory address of the specified data. The address mapping table records the mapping relationship between the virtual memory address and the physical storage address. The physical storage address refers to the physical memory address of the local memory and the physical address of at least one memory. The processing unit 1020 is also used to detect whether the physical storage address of the specified data is the physical memory address of the local memory. The processing unit 1020 is also used to read the data stored in the physical memory address of the specified data, or write the specified data to the corresponding physical memory address when the physical storage address of the specified data is the physical memory address of the local memory.

[0155] In another possible embodiment, the processing unit 1020 is further configured to allocate a designated storage space in the local memory when the physical storage address of the designated data is not a physical memory address of the local memory. The processing unit 1020 is further configured to forward the memory access request instruction to the at least one memory. The memory access request instruction is configured to request the at least one memory to read data stored at the physical memory address of the designated data. The processing unit 1020 is further configured to receive the designated data sent by the at least one memory and store the designated data in the designated storage space.

[0156] In another possible embodiment, the transceiver unit 1010 is further configured to obtain a code for an external memory access interface. The processing unit 1020 is further configured to replace the external memory access interface with a custom external memory access interface. The custom external memory access interface is used to read data stored in at least one memory.

[0157] Figure 11 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. As shown in Figure 11, computing device 1100 includes a bus 1110, a processor 1120, a memory 1130, and a communication interface 1140. Processor 1120, memory 1130, and communication interface 1140 communicate with each other via bus 1110. Computing device 1100 can be a server, a computer, a portable notebook, etc. It should be understood that this application does not limit the number of processors and memories in computing device 1100.

[0158] Bus 1110 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG11 shows a single line, but this does not imply a single bus or type of bus. Bus 1110 may include a path for transmitting information between various components of computing device 1100 (e.g., processor 1120, memory 1130, and communication interface 1140).

[0159] The processor 1120 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0160] The memory 1130 may include a volatile memory, such as a random access memory (RAM). The memory 1130 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0161] Memory 1130 stores executable program code, which processor 1120 executes to implement the functions of the aforementioned modules, such as interface layer 241, cache layer 242, and driver layer 243, thereby implementing the data processing method. In other words, memory 1130 stores instructions for executing the data processing method.

[0162] Alternatively, the memory 1130 stores executable codes, and the processor 1120 executes the executable codes to respectively implement the functions of the aforementioned modules, thereby implementing the data processing method. In other words, the memory 1130 stores instructions for executing the data processing method.

[0163] The communication interface 1140 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.

[0164] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0165] As shown in Figure 12, the computing device cluster includes at least one computing device 1100. The memory 1130 of one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the data processing method.

[0166] In some possible implementations, the memory 1130 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for executing the data processing method.

[0167] It should be noted that the memory 1130 in different computing devices 1100 in the computing device cluster can store different instructions, each for executing part of the functions of the above-mentioned multiple modules. In other words, the instructions stored in the memory 1130 in different computing devices 1100 can implement the functions of one or more modules in the above-mentioned multiple modules.

[0168] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network or a local area network, etc. FIG13 shows a possible implementation. As shown in FIG13 , two computing devices, namely computing device 1100A and computing device 1100B, are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1130 in the computing device 1100A stores instructions for executing the functions of some modules among the above-mentioned multiple modules. At the same time, the memory 1130 in the computing device 1100B stores instructions for executing the functions of another part of the modules among the above-mentioned multiple modules.

[0169] The connection method between the computing device clusters shown in Figure 13 can be considered to be that the data processing method provided by this application requires a large amount of data storage, so the functions implemented by another part of the above-mentioned multiple modules are considered to be handed over to the computing device 1100B for execution.

[0170] It should be understood that the functionality of the computing device 1100A shown in FIG13 may also be implemented by multiple computing devices 1100. Similarly, the functionality of the computing device 1100B may also be implemented by multiple computing devices 1100.

[0171] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection method of the computing device cluster described in Figures 11 and 12. However, the memory 1130 in one or more computing devices 1100 in this computing device cluster can store the same instructions for executing the data processing method.

[0172] In some possible implementations, the memory 1130 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more computing devices 1100 can jointly execute the instructions for executing the data processing method.

[0173] It should be noted that the memory 1130 in different computing devices 1100 in the computing device cluster may store different instructions for executing part of the functions of the computing device 1100. In other words, the instructions stored in the memory 1130 in different computing devices 1100 may implement the functions of one or more of the multiple modules described above.

[0174] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the data processing method.

[0175] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the data processing method.

[0176] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A data processing method, characterized in that: The data processing system is deployed in a server cluster, the server cluster includes multiple computing devices and at least one memory, and the method is executed by any one or more computing devices among the multiple computing devices, including: Acquire cache data; the cache data refers to data that needs to be temporarily stored during the operation of the data processing system; When the cached data amount of the cached data is greater than the available memory space capacity of the local memory, part of the cached data is sent to the at least one memory so that the at least one memory stores part of the cached data.

2. The method according to claim 1, characterized in that The sending part of the cached data to the at least one memory specifically includes: Obtaining a code of a memory allocation interface, and replacing the memory allocation interface with a custom memory allocation interface; the custom memory allocation interface is used to allocate storage space in the local memory and / or the at least one memory; Part of the cache data is stored in the storage space allocated in the at least one memory through the customized memory allocation interface.

3. The method according to claim 1 or 2, characterized in that: In a case where the cached data amount of the cached data is greater than the available memory space capacity of the local memory, before sending part of the cached data to the at least one memory, the method further includes: A transparent memory is configured for each executor; the computing device includes at least one executor, and the computing device uses the executor to store the cache data in the transparent memory configured by the executor, wherein the transparent memory includes the available memory space capacity of the local memory of a first set size and the storage space of the at least one memory of a second set size.

4. The method according to claim 3, characterized in that The storing part of the cache data in the storage space allocated in the at least one memory through the customized memory allocation interface specifically includes: When the cached data amount of the cached data is greater than the available memory space capacity of the local memory, a memory allocation request is sent to the at least one memory; the memory allocation request is used to allocate a target storage space for storing the cached data; Determining, through the custom memory allocation interface, whether the target storage space allocated by the memory allocation request is greater than a set memory threshold; When the target storage space allocated by the memory allocation request is greater than the set memory threshold, storage space is allocated in the transparent memory configured by the executor, and part of the cache data is stored in the target storage space.

5. The method according to any one of claims 2 to 4, characterized in that: The method further comprises: A first request instruction is sent to the at least one memory through the custom memory allocation interface; the first request instruction is used to request the at least one memory to allocate free storage space after the target storage space.

6. The method according to any one of claims 1 to 5, characterized in that: When the amount of cached data of the cached data is greater than the available memory space capacity of the local memory, sending part of the cached data to the at least one memory specifically includes: Detecting the number of times each data in the cache data is accessed and / or stored; Sending first cache data to the at least one memory; the first cache data refers to data in the cache data whose number of accesses and / or storages is less than a set number of times.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Obtaining a memory access request instruction; the memory access request instruction is used to read or write specified data, and the memory access request instruction includes a virtual memory address of the specified data; Querying a locally stored address mapping table to determine a physical storage address of a virtual memory address mapping of the specified data; the address mapping table records a mapping relationship between a virtual memory address and a physical storage address; the physical storage address refers to a physical memory address of the local memory and a physical address of the at least one memory; Detecting whether the physical storage address of the specified data is the physical memory address of the local memory; In the case where the physical storage address of the designated data is the physical memory address of the local memory, the data stored in the physical memory address of the designated data is read, or the designated data is written to the corresponding physical memory address.

8. The method according to claim 7, characterized in that The method further comprises: In a case where the physical storage address of the designated data is not the physical memory address of the local memory, allocating designated storage space in the local memory; forwarding the memory access request instruction to the at least one memory; the memory access request instruction is used to request the at least one memory to read the data stored at the physical storage address of the specified data; The designated data is received from the at least one memory, and the designated data is stored in the designated storage space.

9. The method according to claim 7 or 8, characterized in that: Before obtaining the memory access request instruction, the method further includes: Obtain the code of the external memory access interface, and replace the external memory access interface with a custom external memory access interface; the custom external memory access interface is used to read the data stored in the at least one memory.

10. A data processing device, characterized in that: include: A transceiver unit, used for obtaining cache data; The cache data refers to data that needs to be temporarily stored during the operation of the data processing system; The processing unit is used to send part of the cached data to the at least one memory when the cached data amount of the cached data is greater than the available memory space capacity of the local memory, so that the at least one memory stores part of the cached data.

11. The device according to claim 10, characterized in that The transceiver unit is also used to obtain the code of the memory allocation interface; The processing unit is further used to replace the memory allocation interface with a custom memory allocation interface; the custom memory allocation interface is used to allocate storage space in the local memory and / or the at least one memory; Part of the cache data is stored in the storage space allocated in the at least one memory through the customized memory allocation interface.

12. The device according to claim 10 or 11, characterized in that The processing unit is also used to configure a transparent memory for each executor; the computing device includes at least one executor, and the computing device uses the executor to store the cache data in the transparent memory configured by the executor, and the transparent memory includes an available memory space capacity of the local memory of a first set size and a storage space of the at least one memory of a second set size.

13. The device according to any one of claims 10 to 12, characterized in that: The processing unit is specifically configured to send a memory allocation request to the at least one memory when the cached data amount of the cached data is greater than the available memory space capacity of the local memory; The memory allocation request is used to allocate a target storage space for storing the cache data; Determining, through the custom memory allocation interface, whether the target storage space allocated by the memory allocation request is greater than a set memory threshold; When the target storage space allocated by the memory allocation request is greater than the set memory threshold, storage space is allocated in the transparent memory configured by the executor, and part of the cache data is stored in the target storage space.

14. The device according to any one of claims 11 to 13, characterized in that: The processing unit is further used to send a first request instruction to the at least one memory through the custom memory allocation interface; the first request instruction is used to request the at least one memory to allocate free storage space after the target storage space.

15. The device according to any one of claims 10 to 14, characterized in that: The processing unit is specifically used to detect the number of times each data in the cache data is accessed and / or stored; Sending first cache data to the at least one memory; the first cache data refers to data in the cache data whose number of accesses and / or storages is less than a set number of times.

16. The device according to any one of claims 10 to 15, characterized in that: The processing unit is further used to obtain a memory access request instruction; the memory access request instruction is used to read or write specified data, and the memory access request instruction includes a virtual memory address of the specified data; Querying a locally stored address mapping table to determine a physical storage address of a virtual memory address mapping of the specified data; the address mapping table records a mapping relationship between a virtual memory address and a physical storage address; the physical storage address refers to a physical memory address of the local memory and a physical address of the at least one memory; Detecting whether the physical storage address of the specified data is the physical memory address of the local memory; In the case where the physical storage address of the designated data is the physical memory address of the local memory, the data stored in the physical memory address of the designated data is read, or the designated data is written to the corresponding physical memory address.

17. The device according to claim 16, characterized in that The processing unit is further configured to allocate a designated storage space in the local memory when the physical storage address of the designated data is not a physical memory address of the local memory; forwarding the memory access request instruction to the at least one memory; The memory access request instruction is used to request the at least one memory to read the data stored in the physical storage address of the specified data; The designated data is received from the at least one memory, and the designated data is stored in the designated storage space.

18. The device according to claim 16 or 17, characterized in that The transceiver unit is also used to obtain the code of the external memory access interface; The processing unit is further used to replace the external memory access interface with a custom external memory access interface; the custom external memory access interface is used to read the data stored in the at least one memory.

19. A computing device, characterized in that include: at least one memory; At least one processor, wherein the processor is configured to execute instructions stored in the memory so that the computing device executes the method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a computing device, the computing device performs the method according to any one of claims 1 to 9.

21. A computer program product comprising instructions, characterized in that The computer program product stores instructions, which, when executed by a computing device, enable the computing device to implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • A method and a device for preventing memory overflow during data synchronization

    CN109189577A

  • Data warehouse processing method and device, equipment and medium

    CN115712692A

  • Data processing method and related equipment

    CN115905042A

  • Memory access popularity statistical method, related apparatus and device

    WO2023227004A1

Cited By

  • Distributed tracking data processing method and system based on OTel

    CN120821729A

  • Data processing method, electronic equipment, system, storage medium and program product

    CN121166043A