High-performance data lake data service technology based on kernel bypass
By adopting kernel bypass technology and spatial topology optimization in high-performance data lakes, data transmission latency and memory footprint issues are solved, and more efficient data access and response capabilities are achieved.
Patent Information
- Application Number
- CN202411983367.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art faces the problems of data transmission latency and memory space overhead in high-performance data lakes, especially in the process of multi-node data transmission.
The kernel bypass technology is adopted to directly interact with the network driver through the user layer to reduce the delay caused by operating system-level switching, and combine the optimization of data transmission paths based on spatial topology to select efficient and low-latency channels.
It significantly reduces the data transmission delay and memory usage of data management nodes, and improves the response capability and data access efficiency of the query system.
Smart Images

Figure CN119996291A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a high-performance data lake client technology based on kernel bypass, which reduces the response time and memory space occupancy overhead caused by data transmission of data management nodes through operating systems and network drivers. Background Art
[0002] Digital twins require high-performance data lakes and data management capabilities. For data queries, especially high-frequency, large-capacity, and distributed data queries for visualization and data analysis, high-performance communication support is required. In particular, for data queries and analysis of spatiotemporal information, the spatiotemporal data lake platform needs to have the ability to support high frequencies and different data capacities, and queries initiated from different computing nodes must obtain efficient responses. Therefore, kernel bypass technology is used to reduce the response time and memory space overhead caused by data transmission of data management nodes through the operating system and network drivers. Through direct interaction between the user layer and the network driver, the delay caused by operating system-level switching is reduced. At the same time, combined with the optimization of data transmission paths based on spatial topology, efficient and low-latency channels are selected in the data transmission process across multiple storage nodes to improve the overall response capability of the query system.
[0003] Kernel bypass technology has been proven to have significant advantages over traditional TCP network driver processing in terms of reducing the number of system calls, lowering latency, and improving data processing throughput. In some high-performance database systems, high-performance data interaction is achieved by using RDMA technology. In the data lake scenario, multiple data types need to be processed, including structured data based on relational data representation, unstructured data, spatiotemporal data, file data (such as HDFS (Hadoop Distributed File System)), streaming data environment, and RESTful API based on HTTP protocol.
[0004] This invention proposes an efficient data query response technology to meet the requirements of high-performance, distributed data lake service quality improvement, and improves the query system response capability from two levels. At the end-to-end transmission path level, the kernel bypass technology is used to reduce the data transmission delay and memory usage of the data management node. At the multi-node transmission level, the transmission delay is reduced through spatial topology optimization (i.e., transmission path optimization). Summary of the invention
[0005] The present invention provides the following technical solutions: a high-performance data lake data service technology based on kernel bypass, which avoids the transmission delay and memory resource occupation caused by context switching at the operating system level during the process of large-capacity data transmission in the data lake through kernel bypass technology; The implementation methods of this technology specifically include: S1. Data interaction process design, including kernel bypass and routing optimized data interface connection; S2. Resource planning and management, including data connection diagram G DC The data storage and query clients and data relay transmission nodes are managed in a unified manner; S3. Data routing service.
[0006] Preferably, the data interaction process design is based on not modifying the existing data query and interaction interface, and uses DPDK technology to implement kernel bypass at the network transmission level. The query process is as follows: Step 1: Establish a connection between the data consumption end and the data production end; Step 2: Search for the best route based on the locations of both ends, and establish a data channel based on the best route; Step 3: Complete data transmission and release transmission path resources.
[0007] Preferably, the connection between the data production end and the consumption end is realized through a standard TCP connection, and the connection and interactive handshake between the two ends are the same as the standard data lake client query request process.
[0008] Preferably, in the step of searching for the best route, the route search is implemented through link resource management.
[0009] Preferably, in the step of establishing the data channel, the data channel is established according to the best route and the channel resources are locked. If the locking fails, the best route search will be repeatedly performed.
[0010] Preferably, the data connection graph G DC Includes node type set V DC and node connection type E DC .
[0011] Preferably, wherein said V DC Including data storage nodes, client nodes and relay transmission nodes, E DC Four connection types are included.
[0012] Preferably, the data routing service includes an optimal routing search algorithm and a resource locking process.
[0013] Preferably, the optimal route search starts with a random search to find at most K more optimized paths, where K is a user-configurable value.
[0014] Preferably, the resource locking process stores the corresponding path of each node on the optimized path in the connection pool P on the node. S and P F中 Lock the resource. If the lock fails, restart the search until the optimized resource is successfully locked.
[0015] Compared with the prior art, the present invention provides a high-performance data lake data service technology based on kernel bypass, which has the following beneficial effects: This technology implements a data access plug-in based on kernel bypass technology on the client side and provides corresponding kernel bypass support on the data server side, thus realizing data access services for multiple interfaces, multiple data types, and high concurrency. It directly interacts with the network driver through the user layer to reduce the delay caused by operating system-level switching, and optimizes the data transmission path based on spatial topology to improve the response capability of the query system. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 The data query / management interface workflow based on kernel bypass of the present invention; Figure 2 A data channel pool for bypassing the kernel of the present invention; Figure 3 Optimizing the data routing of the present invention; Figure 4 The data interaction flow chart of the present invention; Figure 5 This is a specific flow chart of the optimal data routing of the present invention. DETAILED DESCRIPTION
[0017] The specific implementation of this invention includes three parts: data interaction process design, resource planning and management, and data routing. The data interaction process design shows how to combine the kernel bypass and routing optimization proposed by this invention on the basis of the existing data interface; resource planning management uses a data channel pool to achieve efficient management and reuse of kernel bypass channels to avoid resource loss and response delays caused by multiple connection initiations, while providing services for channel selection during data transmission between multiple nodes; data routing, as a transmission path optimization, provides optimized paths and communication methods during data transmission across multiple nodes. The specific implementation methods are described as follows: Data interaction process: like Figure 1 As shown, data interaction is based on not modifying the existing data query and interaction interface. Therefore, the present invention directly includes the following three steps in the data interaction process when initiating and ending a query: 1. Establish a connection between the data consumption end and the data production end; 2. Search for the best route based on the locations of both ends and establish a data channel based on the best route; 3. Complete data transmission and release transmission path resources.
[0018] In the first step, the connection between the data production end and the data consumption end does not need to consider the capacity, so it is implemented through a standard TCP connection. The connection and interactive handshake between the two ends are the same as the standard data lake client query request process, and the route search is implemented through link resource management; The second step is to search for the best route to find the available transmission path with the smallest transmission delay at both the production and consumption ends, establish a data channel based on the best route and lock the channel resources. If the lock fails, the search will be repeated; The last step completes the data transfer and releases the resources of the transmission path.
[0019] The overall process is attached Figure 4 .
[0020] Resource Management: In this invention, resource management is to manage the data transmission channel between the entire distributed data lake and the query client in a unified manner. Therefore, the data storage and query client and data relay transmission nodes are represented in the form of a data connection graph GDC, which is specifically represented as follows: G DC = {V DC , E DC},V DC = {v dc0 , v dc1 , … v dcm}, v dci Represents 3 node types: 1) Data storage node; 2) Client node; 3) Relay transmission node, E DC = {e dc0 , e dc1 , … e dck}, e dcj Indicates 4 types of connections: 1) <data storage node, client node>; 2) <Data storage node, relay transmission node>; 3) <Relay transmission node, client node>; 4) <relay transmission node, relay transmission node>.
[0021] like Figure 2 As shown, each edge manages two connection pools: TCP connection pool: P S and kernel bypass connection pool P FThe kernel bypass at the network transmission level uses Intel's high-performance data processing package framework (data plane development kit DPDK) technology. Each routing query is provided by each node to the query process. S and P F The status of the connection, including the number of available connections and connection lock services. And resource allocation priority allocation P F If P F If the connection cannot be provided, S Apply for connection. Each routing query resource locking process will F or P S The selected link in the P is set to an unavailable state. The resource release after each data transmission is completed will F or P S The corresponding data link in is restored to an available state.
[0022] Data routing: like Figure 3 As shown in Figure 1, data routing includes the optimal route search algorithm and resource locking process. The optimal route search starts with a random search to find at most K more optimal paths (lower transmission delay), where K is a user-configurable value. The delay calculation method for the optimized path is:
[0023] get_connect represents the connection resources allocated by each edge for this transmission; The cost represents the transmission delay for a given connection resource.
[0024] The resource locking process will be the corresponding path of each node on the optimized path to be found in the connection resource pool P on the node S and P F If the lock fails, the search is restarted until the optimized resource is successfully locked.
[0025] The overall process input of optimal route search and resource locking is the data connection graph G DC , and data production node v dcp , Data consumption phase v dcc , the output is a list of edges in the data connection graph and the associated connections (P S or P F中 ).
[0026] The overall process of data routing is as follows Figure 5 shown.
[0027] Therefore, by directly interacting between the user layer and the network driver, the delay caused by operating system-level switching can be reduced, and the data transmission path can be optimized based on spatial topology to improve the responsiveness of the query system.
Claims
1. A high-performance data lake data service technology based on kernel bypass, characterized by: Kernel bypass technology is used to avoid transmission delays and memory resource usage caused by context switching at the operating system level during large-capacity data transmission in the data lake; The implementation methods of this technology specifically include: S1. Data interaction process design, including kernel bypass and routing optimized data interface connection; S2. Resource planning and management, including data connection diagram G DC The data storage and query clients and data relay transmission nodes are managed in a unified manner; S3. Data routing service.
2. The high-performance data lake data service technology based on kernel bypass according to claim 1 is characterized in that: The data interaction process design is based on not modifying the existing data query and interaction interface, and uses DPDK technology to implement kernel bypass at the network transmission level. The query process is as follows: Step 1: Establish a connection between the data consumption end and the data production end; Step 2: Search for the best route based on the locations of both ends, and establish a data channel based on the best route; Step 3: Complete data transmission and release transmission path resources.
3. The high-performance data lake data service technology based on kernel bypass according to claim 2 is characterized in that: The connection between the data production end and the data consumption end is realized through a standard TCP connection, and the connection and interactive handshake between the two ends are the same as the standard data lake client query request process.
4. The high-performance data lake data service technology based on kernel bypass according to claim 2 is characterized in that: In the step of searching for the best route, the route search is implemented through link resource management.
5. The high-performance data lake data service technology based on kernel bypass according to claim 2 is characterized in that: In the step of establishing a data channel, a data channel is established according to the best route and channel resources are locked. If the locking fails, the best route search will be repeatedly performed.
6. The high-performance data lake data service technology based on kernel bypass according to claim 1 is characterized in that: The data connection graph G DC Includes node type set V DC and node connection type E DC .
7. The high-performance data lake data service technology based on kernel bypass according to claim 6 is characterized in that: Wherein V DC Including data storage nodes, client nodes and relay transmission nodes, E DC Four connection types are included.
8. The high-performance data lake data service technology based on kernel bypass according to claim 1 is characterized in that: The data routing service includes an optimal routing search algorithm and a resource locking process.
9. The high-performance data lake data service technology based on kernel bypass according to claim 8 is characterized in that: The optimal route search starts with a random search, looking for at most K more optimal paths, where K is a user-configurable value.
10. The high-performance data lake data service technology based on kernel bypass according to claim 8, characterized in that: The resource locking process stores the corresponding path of each node on the optimized path in the connection pool P on the node. S and P F中 Lock the resource. If the lock fails, restart the search until the optimized resource is successfully locked.