Multi-modal distributed storage RDMA primitive optimization method and system and related equipment
Through multimodal data analysis and machine learning, the combination of RDMA primitives is solved, and the performance and reliability of distributed storage systems are improved.
Patent Information
- Application Number
- CN202510375538.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The existing RDMA primitives are difficult to take into account high throughput and low latency in streaming data scenarios, and the heterogeneous hardware resource utilization is low.
Receive access requests through the multimodal storage interface, analyze data modal tags, QoS indicators and environmental parameters, match the initial RDMA primitive combination using a pre-constructed mapping relationship rule library, and dynamically optimize it in combination with the machine learning model to generate target RDMA primitive combinations, and perform RDMA operations in pipeline order.
It significantly improves the performance and reliability of RDMA operations in distributed storage systems, can cope with real-time network status and hardware resources changes, and achieve dynamic switching with high throughput and low latency.
Smart Images

Figure CN120429253A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of RDMA primitive optimization, and particularly to an RDMA primitive optimization method, system and related devices for multi-modal distributed storage. Background Art
[0002] With the increasing demand of distributed storage systems for low latency and high throughput, Remote Direct Memory Access (RDMA) technology has gradually become a key means to optimize data access performance. Currently, RDMA primitives usually statically classify data access requests based on a single dimension. For example, small messages are assigned to SEND / RECV primitives, and large chunks of data use WRITE / READ primitives, and predefined primitive combinations are matched based on fixed rules. However, for some streaming data, the sudden high throughput demand may exceed the processing capacity of the fixed primitive combination. In addition, the existing primitive combinations are difficult to cope with real-time network status and hardware resource changes. For example, in the event of sudden network congestion, fixedly using the high-throughput READ primitive may lead to a sharp increase in packet loss retransmission, and it is impossible to dynamically switch to atomic operation primitives that support congestion control. That is, it is difficult to balance the goals of high throughput and low latency in the streaming data scenario, and the utilization rate of heterogeneous hardware resources is low.
[0003] In view of this, there is a need for an RDMA primitive optimization method, system and related devices for multi-modal distributed storage. Summary of the Invention
[0004] Embodiments of the present application provide an RDMA primitive optimization method, system and related devices for multi-modal distributed storage, which are used to solve the problem that it is difficult to balance the goals of high throughput and low latency in the streaming data scenario and the low utilization rate of heterogeneous hardware resources.
[0005] The first aspect of the embodiments of the present application provides an RDMA primitive optimization method for multi-modal distributed storage, including:
[0006] Receiving an access request through a multi-modal storage interface;
[0007] Analyzing the data modality label, QoS metrics and environmental parameters in the access request, where the data modality label includes structured data, unstructured data and streaming data, the QoS metrics include latency requirements, throughput requirements and consistency levels, and the environmental parameters include network latency, packet loss rate, network card queue depth and hardware load;
[0008] Matching an initial RDMA primitive combination corresponding to the data modality label and QoS metrics based on a pre-constructed mapping relationship rule library;
[0009] Combine the real-time collected network status data and the environmental parameters, and dynamically optimize the initial RDMA primitive combination through a machine learning model to generate a target RDMA primitive combination;
[0010] Send the target RDMA primitive combination to heterogeneous hardware resources and execute RDMA operations in a pipeline order, where the pipeline order is scheduled based on operation dependencies and hardware resource parallelism.
[0011] Furthermore, parsing the data modality label, QoS metrics, and environmental parameters in the access request, where the data modality label includes structured data, unstructured data, and streaming data, the QoS metrics include latency requirements, throughput requirements, and consistency levels, and the environmental parameters include network latency, packet loss rate, network card queue depth, and hardware load, includes:
[0012] Parse the transaction type and associated table of structured data, metadata of unstructured data, BLOB fingerprint data, and stream ID and time series index of streaming data according to the data type respectively to generate a modality identifier;
[0013] Extract the latency threshold, throughput requirements, and consistency level from the access request header fields, map the transaction isolation level for structured data, and dynamically calculate the throughput and set the default consistency for streaming data;
[0014] Fuse historical data and real-time data by measuring network latency, counting packet loss rate, monitoring network card queue depth, and collecting hardware load;
[0015] Package the parsed and fused multi-dimensional parameters into a unified context object.
[0016] Furthermore, the matching of the initial RDMA primitive combination corresponding to the data modality label and QoS metrics based on a pre-constructed mapping relationship rule library includes:
[0017] Preprocess the multi-dimensional features in the historical performance data, where the preprocessing includes encoding the data modality label as an enumeration value, normalizing the QoS metrics, and discretizing the environmental parameters into low, medium, and high level intervals;
[0018] Select splitting features from the multi-dimensional features based on information gain;
[0019] Recursively split to generate a decision tree until the number of node samples is less than the threshold. Each leaf node in the decision tree stores the RDMA primitive combination strategy with the highest comprehensive score, where the comprehensive score is the weighted sum of latency, throughput, and error rate.
[0020] Furthermore, the selection of splitting features from the multi-dimensional features based on information gain includes:
[0021]
[0022] Where: S is the data set, A is the feature to be evaluated, S v For the subset of feature A whose value is v, C is the number of RDMA primitive combination categories, p i is the sample ratio of the i-th type of RDMA primitive combination in the dataset.
[0023] Furthermore, the matching of the initial RDMA primitive combination corresponding to the data modality label and the QoS indicator based on the pre-built mapping relationship rule base further includes:
[0024] When hardware resource changes or new RDMA primitive types are detected, new historical performance data is collected and the decision tree model is retrained;
[0025] Conflict detection is performed on the original leaf node policy. If the QoS indicator difference between the new and old policies exceeds the threshold, the new policy with higher confidence is replaced;
[0026] The updated decision tree model is serialized into a new mapping relationship rule base, and the old rule base at runtime is replaced through the hot loading mechanism.
[0027] Furthermore, the method combines the real-time collected network status data and the environmental parameters to dynamically optimize the initial RDMA primitive combination through a machine learning model to generate a target RDMA primitive combination, including:
[0028] Collect network bandwidth, packet loss rate, and hardware load data in real time, and construct a normalized feature vector containing data modality, QoS indicators, and initial primitive encoding;
[0029] Inputting the normalized feature vector into a pre-trained deep Q-network model, and selecting the optimal adjustment action by maximizing the objective function;
[0030] A target RDMA primitive combination is generated according to the model output action, wherein the target RDMA primitive combination includes primitive replacement, parameter tuning, or mixed mode allocation, and the target RDMA primitive combination is verified in a sandbox environment and then submitted to a production environment.
[0031] Furthermore, generating a target RDMA primitive combination according to the model output action, wherein the target RDMA primitive combination includes primitive replacement, parameter tuning, or mixed mode allocation, and submitting the target RDMA primitive combination to the production environment after verification in the sandbox environment, includes:
[0032] Compile the compute-intensive operations in the target RDMA primitive combination into CUDA kernels and allocate them to the streaming multiprocessors of the GPU for parallel execution;
[0033] Convert the data transfer primitive into a DMA descriptor chain of the DPU through the NVIDIA DOCA SDK and register it to the network card queue;
[0034] Construct a DAG scheduling graph based on the operation dependencies and split the tasks of the independent nodes according to the hardware parallelism.
[0035] Furthermore, the target RDMA primitive combination is sent to heterogeneous hardware resources, and the RDMA operations are executed in a pipeline order. The pipeline order is scheduled based on operation dependencies and hardware resource parallelism, including:
[0036] Construct a directed acyclic graph based on the operation semantics of the target RDMA primitive combination, where the nodes represent RDMA operations and the edges represent the data dependencies between operations;
[0037] Use the critical path first scheduling algorithm to calculate the total latency of each path in the directed acyclic graph, and preferentially allocate the operations on the longest path to low-latency hardware resources;
[0038] Dynamically split the tasks of independent nodes according to the hardware resource parallelism;
[0039] Allocate the split tasks to the multiple queues of the GPU, DPU, and intelligent network card, and synchronously insert hardware barrier instructions.
[0040] The second aspect of the embodiments of the present application provides an RDMA primitive optimization system for multi-modal distributed storage, including:
[0041] An access request receiving unit, configured to receive an access request through a multi-modal storage interface;
[0042] A data parameter parsing unit, configured to parse the data modality label, QoS metric, and environmental parameter in the access request. The data modality label includes structured data, unstructured data, and streaming data. The QoS metric includes latency requirements, throughput requirements, and consistency levels. The environmental parameter includes network latency, packet loss rate, network card queue depth, and hardware load;
[0043] An initial RDMA primitive combination matching unit, configured to match an initial RDMA primitive combination corresponding to the data modality label and QoS metric based on a pre-constructed mapping relationship rule library;
[0044] A target RDMA primitive combination generation unit, which is used to combine the network status data collected in real time and the environmental parameters, and dynamically optimize the initial RDMA primitive combination through a machine learning model to generate a target RDMA primitive combination;
[0045] An RDMA operation execution unit, which is used to send the target RDMA primitive combination to heterogeneous hardware resources and execute RDMA operations in a pipeline order, and the pipeline order is scheduled based on operation dependencies and hardware resource parallelism.
[0046] In the third aspect of the embodiments of the present application, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the RDMA primitive optimization method for multi-modal distributed storage as described in any one of the above are implemented.
[0047] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0048] Based on the multi-dimensional feature analysis of data modality, QoS, and environmental parameters, the present invention determines the mapping relationship between multi-modal data and primitives based on a pre-constructed mapping relationship rule library, and matches the optimal primitive combination for scenarios such as strong consistency of structured data and high throughput of streaming data from multiple dimensions; then, through a reinforcement learning model, the network status and hardware load are fused in real time, and the primitive combination and parameters are dynamically adjusted, so as to automatically switch to a degradation strategy in case of sudden congestion or hardware failure. Through the dynamic classification of multi-modal data, the collaborative optimization of the rule library and machine learning, and the heterogeneous hardware pipeline scheduling, the present invention significantly improves the performance and reliability of RDMA operations in a distributed storage system. Description of the Drawings
[0049] Figure 1 It is a schematic flowchart of an embodiment of a method for optimizing RDMA primitives for multi-modal distributed storage in the present invention. Detailed Embodiments
[0050] In the description, claims and the above-mentioned drawings of the present invention, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0051] Embodiment 1
[0052] In this embodiment, the implementation method can be implemented in a system, in a server, or in a terminal, and no specific limitation is made. Next, from the perspective of system implementation, the RDMA primitive optimization method for multimodal distributed storage in the present application will be introduced. Please refer to Figure 1 , the method provided by the embodiment of the present application includes the following steps:
[0053] S11. Receive an access request through a multimodal storage interface;
[0054] In this embodiment, the access request includes structured data, unstructured data, and streaming data. Receiving the access request through the multimodal storage interface lies in multi-protocol adaptation, dynamic dispatching, and metadata enhancement. Specifically, the interface supports protocols such as HTTP and RDMA-native Verbs based on a unified communication framework, automatically analyzes structured data, unstructured data, and streaming data, and ensures the integrity and security of the request through CRC check and RBAC authentication. Subsequently, it is dispatched to the corresponding processing queues according to the data type, and the queues include a high-consistency transaction queue, a large-block data buffer, and a low-latency circular buffer; at the same time, metadata such as data modality tags and timestamps are injected into the request header. Here, by being compatible with the characteristics of multimodal data, ensuring secure access, and pre-classifying to adapt to subsequent processing processes, high-throughput and low-latency request reception, precise resource pre-allocation, and end-to-end traceability capabilities are finally achieved, providing a data basis for subsequent steps.
[0055] S12. Analyze the data modality tags, QoS metrics, and environmental parameters in the access request. The data modality tags include structured data, unstructured data, and streaming data. The QoS metrics include latency requirements, throughput requirements, and consistency levels. The environmental parameters include network latency, packet loss rate, network card queue depth, and hardware load;
[0056] Step S12 includes the following:
[0057] S121. Parse the transaction type and associated tables of structured data, the metadata of unstructured data, the BLOB fingerprint data, and the stream ID and time series index of stream data according to the data type, and generate a modality identifier. Here, the BLOB fingerprint data is a unique identifier generated for a binary large object (BLOB) through a hashing algorithm, which is used to characterize the integrity and uniqueness of the BLOB content. Its essence is a hash value with a fixed length, which can uniquely map the original data content. Even if the data changes slightly, the fingerprint will be significantly different;
[0058] Extract key features according to the data type to generate a unified modality identifier, providing a classification basis for subsequent RDMA primitive matching.
[0059] Structured data parsing includes: parsing the operation_type field from the Protobuf / JSON payload, such as SELECT, UPDATE, and determining the transaction type by combining SQL syntax tree analysis; extracting the table names and foreign key relationships in the SQL statement to generate a Schema hash value for identifying data correlation; finally, outputting the identifier.
[0060] Unstructured data parsing includes: reading Content-Type, Content-Length, and custom tags from the object storage request header; calculating the SHA-256 hash for the binary large object (BLOB) to generate a unique fingerprint; finally, outputting the identifier.
[0061] Stream data parsing includes: extracting the Stream ID (such as stream_123), sequence number, and timestamp from the packet header, and generating a time series index according to a time window (such as 1 second); reading the TCP window size or Kafka partition ID to determine the stream priority; finally, outputting the identifier.
[0062] S122. Extract the latency threshold, throughput requirement, and consistency level from the access request header fields, map the transaction isolation level for structured data, and dynamically calculate the throughput and set the default consistency for stream data;
[0063] This step combines the explicitly declared QoS requirements with implicit rules, converts them into quantifiable parameters, and guides the optimization of primitive combinations. Specifically, the latency requirement parsing includes: reading the maximum allowed latency from the X-QoS-Latency field in the request header; if not explicitly specified, setting default values according to the data modality; for structured data, the defaults are for strong consistency transactions and weak consistency queries, and for streaming data, the defaults are for real-time processing and batch processing. Calculate the throughput requirements based on the operation type of unstructured data, and estimate the throughput requirements based on the sampling rate and single-sample size of streaming data. The consistency level is mapped according to the transaction isolation level of structured data, including weak consistency and strong consistency; for streaming data, weak consistency is the default, and it is upgraded to strong consistency if the request header contains X-Consistency-Level.
[0064] S123. By measuring network latency, statistically analyzing packet loss rate, monitoring the network card queue depth, and collecting hardware load, fuse historical data with real-time data;
[0065] This step predicts short-term environmental changes by collecting network and hardware status in real time and combining historical data, improving the robustness of the optimization strategy. The fusion includes aligning historical sliding window data with real-time collected values according to timestamps and removing transient noise.
[0066] S124. Package the parsed and fused multi-dimensional parameters into a unified context object.
[0067] Finally, integrate the parsing results into a unified data structure, write the context object into the shared memory or DMA mapped area, avoid multiple memory copies, and cache the context object for high-frequency requests for 5 seconds to reduce the overhead of repeated parsing.
[0068] S13. Based on a pre-constructed mapping relationship rule library, match the initial RDMA primitive combination corresponding to the data modality label and QoS metrics;
[0069] Step S13 includes the following:
[0070] S131. Preprocess the multi-dimensional features in the historical performance data. The preprocessing includes encoding the data modality label as an enumeration value, normalizing the QoS metrics, and discretizing the environmental parameters into low, medium, and high level intervals;
[0071] The encoding is structured data → enumeration value 0, unstructured data → 1, streaming data → 2; the normalization process is to perform linear mapping and logarithmic scaling; convert the original heterogeneous data into a normalized format suitable for decision tree training, which can eliminate the dimension difference and improve the model convergence efficiency.
[0072] S132. Select splitting features from the multi-dimensional features based on information gain;
[0073]
[0074] Where: S is the data set, A is the feature to be evaluated, and S v is the subset where the value of feature A is v, C is the number of categories of RDMA primitive combination classes, and p i is the proportion of samples of the i-th type of RDMA primitive combination in the data set.
[0075] Sort the calculated features in descending order of gain. For example, data modality is greater than latency, and latency is greater than throughput.
[0076] S133. Recursively split to generate a decision tree until the number of samples in the node is less than the threshold. Each leaf node in the decision tree stores the RDMA primitive combination strategy with the highest comprehensive score, and the comprehensive score is the weighted sum of latency, throughput, and error rate.
[0077] Select the feature with the highest gain and the splitting point as the splitting condition. The stopping condition is set to that the number of samples in the node is less than 5% of the total samples or the tree depth is greater than or equal to 5 levels. Calculate the score of the leaf node, and select the primitive combination with the highest score as the initial strategy.
[0078] Specifically, the dynamic update method of the rule library includes the following:
[0079] 1. When detecting changes in hardware resources or new types of RDMA primitives, collect the new historical performance data and retrain the decision tree model;
[0080] 2. Detect conflicts in the original leaf node strategies. If the QoS index difference between the new and old strategies exceeds the threshold, replace it with the new strategy with higher confidence;
[0081] 3. Serialize the updated decision tree model into a new mapping relationship rule library, and replace the old rule library at runtime through the hot loading mechanism.
[0082] Detect new hardware or new primitives, record the performance data in the new scenario, including features and labels, and perform incremental training or full retraining on the original decision tree. Judge conflicts based on the QoS difference threshold between the new and old strategies. Only replace the old strategy when the confidence of the new strategy is ≥ 90%. Load the new rule library into the memory standby area, atomically switch the pointer to point to the new library, and asynchronously release the resources of the old library.
[0083] S14. Combine the real-time collected network status data and environmental parameters, and dynamically optimize the initial RDMA primitive combination through a machine learning model to generate the target RDMA primitive combination;
[0084] Step S14 includes the following:
[0085] S141. Collect network bandwidth, packet loss rate, and hardware load data in real time, and construct a normalized feature vector containing data modality, QoS metrics, and initial primitive encoding;
[0086] S142. Input the normalized feature vector into a pre-trained deep Q-network model, and select the optimal adjustment action by maximizing the objective function;
[0087] Concatenate the normalized data modality, QoS metrics, and initial primitive encoding into a unified feature vector, and input the feature vector into the pre-trained DQN model. The model architecture includes an input layer, a hidden layer, and an output layer. The objective function is: Q(s,a) = α·T + β·(1 / L) ― γ·E, where T is the throughput gain, L is the latency penalty, E is the error rate, and α, β, and γ are the corresponding weight coefficients. Finally, adopt the ε-greedy strategy to select the action with the highest Q value with a 90% probability and randomly explore new strategies with a 10% probability.
[0088] S143. Generate a target RDMA primitive combination according to the model output action. The target RDMA primitive combination includes primitive replacement, parameter tuning, or hybrid mode allocation, and submit the target RDMA primitive combination to the production environment after verification in the sandbox environment.
[0089] Simulate the target combination in an independent RDMA channel for the policy corresponding to the generated target RDMA primitive combination, write the verified target combination into the RDMA queue, and notify the client to switch to the new policy. Specifically, it includes the following:
[0090] 1. Compile the compute-intensive operations in the target RDMA primitive combination into CUDA kernels and allocate them to the streaming multiprocessors of the GPU for parallel execution;
[0091] 2. Convert the data transfer primitive to a DPU DMA descriptor chain through the NVIDIA DOCA SDK and register it to the network card queue;
[0092] 3. Construct a DAG scheduling graph based on the operation dependency relationship and split tasks for non-dependent nodes according to the hardware parallelism.
[0093] Seamlessly issue the target combination through the hot swap mechanism, and at the same time compile the compute-intensive operations into CUDA kernels, DPU descriptor chains, and DAG scheduling tasks to achieve end-to-end low latency and high throughput execution, and monitor the performance in real time and feedback it to the model for iterative optimization.
[0094] S15. Issue the target RDMA primitive combination to heterogeneous hardware resources and execute RDMA operations in a pipeline order. The pipeline order is scheduled based on operation dependencies and hardware resource parallelism.
[0095] S151. Construct a directed acyclic graph based on the operation semantics of the target RDMA primitive combination, where nodes represent RDMA operations and edges represent data dependency relationships between operations;
[0096] S152. Use the critical path first scheduling algorithm to calculate the total latency of each path in the directed acyclic graph, and preferentially allocate operations on the longest path to low-latency hardware resources;
[0097] S153. Dynamically split the task of independent nodes according to the parallelism of hardware resources;
[0098] S154. Allocate the split tasks to multiple queues of GPUs, DPUs, and intelligent network cards, and synchronously insert hardware barrier instructions.
[0099] The above steps are specifically as follows: Parse based on the target RDMA primitive combination, and mark the maximum parallelism for each node. Preferentially schedule operations on the longest path in the DAG to minimize the overall latency; allocate compute-intensive operations to GPUs / DPUs and high-throughput continuous transmissions to the multi-queues of intelligent network cards; dynamically migrate tasks to low-load devices according to the real-time hardware load. Generate the operation execution order. For each operation node, select hardware from the resource pool that meets the supported operation type and the current load is less than the maximum capacity threshold. Use the greedy algorithm to allocate time windows for operations to ensure maximum parallelization of independent operations. Convert the RDMA primitive into the underlying instructions of the target hardware, generate a DMA descriptor chain through the DOCA SDK, and compile it into a CUDA kernel; directly map the data buffer address to the hardware access space through GPUDirect RDMA to avoid CPU memory copying. Ensure the serialization of results through a global memory lock. If any operation fails, trigger a rollback in reverse along the DAG.
[0100] Through multi-modal data dynamic classification, collaborative optimization of the rule base and machine learning, and heterogeneous hardware pipeline scheduling, the selected primitive combination in the above embodiments can cope with real-time network status and hardware resource changes, effectively improving the performance and reliability of RDMA operations in the distributed storage system.
[0101] Embodiment 2
[0102] An embodiment of the RDMA primitive optimization system for multi-modal distributed storage in the present invention includes the following steps:
[0103] An access request receiving unit, configured to receive an access request through a multi-modal storage interface;
[0104] A data parameter parsing unit for parsing the data modality tag, QoS metrics, and environmental parameters in the access request, where the data modality tag includes structured data, unstructured data, and streaming data, the QoS metrics include latency requirements, throughput requirements, and consistency levels, and the environmental parameters include network latency, packet loss rate, network card queue depth, and hardware load;
[0105] An initial RDMA primitive combination matching unit for matching an initial RDMA primitive combination corresponding to the data modality tag and QoS metrics based on a pre-constructed mapping relationship rule library;
[0106] A target RDMA primitive combination generation unit for dynamically optimizing the initial RDMA primitive combination through a machine learning model by combining real-time collected network status data and the environmental parameters to generate a target RDMA primitive combination;
[0107] An RDMA operation execution unit for sending the target RDMA primitive combination to heterogeneous hardware resources and executing RDMA operations in a pipeline order, where the pipeline order is scheduled based on operation dependencies and hardware resource parallelism.
[0108] Embodiment III
[0109] The present invention provides a computer device, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the above method are implemented.
[0110] Those of ordinary skill in the art can realize that the units of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0111] In the embodiments provided by the present invention, it should be understood that the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units can be combined into one unit, one unit can be split into multiple units, or some features can be ignored. In addition, the functional units in each embodiment of the present invention can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0112] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (such as a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0113] It can be understood that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.
Claims
1. A RDMA primitive optimization method for multimodal distributed storage, characterized in that: include: receiving an access request via a multimodal storage interface; Parsing the data modality tag, QoS indicator, and environmental parameters in the access request, where the data modality tag includes structured data, unstructured data, and streaming data; the QoS indicator includes delay requirement, throughput requirement, and consistency level; and the environmental parameters include network delay, packet loss rate, network card queue depth, and hardware load; Matching an initial RDMA primitive combination corresponding to the data modality label and QoS indicator based on a pre-built mapping relationship rule base; In combination with the real-time collected network status data and the environmental parameters, the initial RDMA primitive combination is dynamically optimized through a machine learning model to generate a target RDMA primitive combination; The target RDMA primitive combination is sent to heterogeneous hardware resources, and RDMA operations are performed in a pipeline sequence, where the pipeline sequence is scheduled based on operation dependencies and hardware resource parallelism.
2. The RDMA primitive optimization method for multimodal distributed storage according to claim 1, characterized in that: The parsing of the data modality tag, QoS indicator, and environmental parameters in the access request, wherein the data modality tag includes structured data, unstructured data, and streaming data; the QoS indicator includes delay requirement, throughput requirement, and consistency level; and the environmental parameters include network delay, packet loss rate, network card queue depth, and hardware load, includes: Based on the data type, the transaction type and association table of structured data, metadata and BLOB fingerprint data of unstructured data, and stream ID and time series index of stream data are analyzed to generate modality identifiers. Extract latency thresholds, throughput requirements, and consistency levels from access request header fields, map structured data to transaction isolation levels, dynamically calculate throughput for streaming data, and set default consistency. By measuring network latency, calculating packet loss rate, monitoring network card queue depth, and collecting hardware load, historical data is integrated with real-time data. Encapsulates the parsed and fused multi-dimensional parameters into a unified context object.
3. The RDMA primitive optimization method for multimodal distributed storage according to claim 1, characterized in that: The matching of the initial RDMA primitive combination corresponding to the data modality label and the QoS indicator based on the pre-built mapping relationship rule base includes: Preprocess the multidimensional features in historical performance data, including encoding data modality labels into enumerated values, normalizing QoS indicators, and discretizing environmental parameters into low, medium, and high level intervals. selecting a split feature from the multidimensional features based on information gain; A decision tree is generated by recursive splitting until the number of node samples is less than a threshold. Each leaf node in the decision tree stores the RDMA primitive combination strategy with the highest comprehensive score, which is a weighted sum of latency, throughput, and error rate.
4. The RDMA primitive optimization method for multimodal distributed storage according to claim 3, characterized in that: The selecting a split feature from the multidimensional feature based on information gain includes: Where: S is the data set, A is the feature to be evaluated, S v For the subset of feature A whose value is v, C is the number of RDMA primitive combination categories, p i is the sample ratio of the i-th type of RDMA primitive combination in the dataset.
5. The RDMA primitive optimization method for multimodal distributed storage according to claim 4, characterized in that: The matching of the initial RDMA primitive combination corresponding to the data modality label and the QoS indicator based on the pre-built mapping relationship rule base also includes: When hardware resource changes or new RDMA primitive types are detected, new historical performance data is collected and the decision tree model is retrained; Conflict detection is performed on the original leaf node policy. If the QoS indicator difference between the new and old policies exceeds the threshold, the new policy with higher confidence is replaced; The updated decision tree model is serialized into a new mapping relationship rule base, and the old rule base at runtime is replaced through the hot loading mechanism.
6. The RDMA primitive optimization method for multimodal distributed storage according to claim 1, characterized in that: The method of dynamically optimizing the initial RDMA primitive combination by combining the real-time collected network status data and the environmental parameters through a machine learning model to generate a target RDMA primitive combination includes: Collect network bandwidth, packet loss rate, and hardware load data in real time, and construct a normalized feature vector containing data modality, QoS indicators, and initial primitive encoding; Inputting the normalized feature vector into a pre-trained deep Q-network model, and selecting the optimal adjustment action by maximizing the objective function; A target RDMA primitive combination is generated according to the model output action, wherein the target RDMA primitive combination includes primitive replacement, parameter tuning, or mixed mode allocation, and the target RDMA primitive combination is verified in a sandbox environment and then submitted to a production environment.
7. The RDMA primitive optimization method for multimodal distributed storage according to claim 6, characterized in that: Generating a target RDMA primitive combination according to the model output action, wherein the target RDMA primitive combination includes primitive replacement, parameter tuning, or mixed mode allocation, and submitting the target RDMA primitive combination to the production environment after verification in the sandbox environment, includes: Compile the computationally intensive operations in the target RDMA primitive combination into CUDA kernels and distribute them to the streaming multiprocessors of the GPU for parallel execution; Convert the data transfer primitive into the DPU's DMA descriptor chain and register it to the network card queue; Build a DAG scheduling graph based on operation dependencies and split tasks for non-dependent nodes according to hardware parallelism.
8. The RDMA primitive optimization method for multimodal distributed storage according to claim 1, characterized in that: The target RDMA primitive combination is sent to heterogeneous hardware resources, and RDMA operations are performed in a pipeline sequence, wherein the pipeline sequence is scheduled based on operation dependencies and hardware resource parallelism, including: Constructing a directed acyclic graph based on the operational semantics of the target RDMA primitive combination, where nodes represent RDMA operations and edges represent data dependencies between operations; The critical path first scheduling algorithm is used to calculate the total delay of each path in the directed acyclic graph, and operations on the longest path are preferentially allocated to low-latency hardware resources; Dynamically split non-dependent node tasks based on hardware resource parallelism; The split tasks are assigned to multiple queues of the GPU, DPU, and SmartNIC, and hardware barrier instructions are inserted synchronously.
9. A multimodal distributed storage RDMA primitive optimization system, characterized in that: The RDMA primitive optimization method for multimodal distributed storage according to any one of claims 1 to 8 comprises: an access request receiving unit, configured to receive an access request via the multimodal storage interface; a data parameter parsing unit, configured to parse a data modality tag, QoS indicators, and environmental parameters in the access request, wherein the data modality tag includes structured data, unstructured data, and streaming data; the QoS indicators include delay requirements, throughput requirements, and consistency levels; and the environmental parameters include network delay, packet loss rate, network card queue depth, and hardware load; An initial RDMA primitive combination matching unit, configured to match an initial RDMA primitive combination corresponding to the data modality label and the QoS indicator based on a pre-built mapping relationship rule base; a target RDMA primitive combination generating unit, configured to dynamically optimize the initial RDMA primitive combination by using a machine learning model based on the real-time collected network status data and the environmental parameters to generate a target RDMA primitive combination; The RDMA operation execution unit is used to send the target RDMA primitive combination to the heterogeneous hardware resources and execute the RDMA operation in a pipeline sequence, wherein the pipeline sequence is scheduled based on the operation dependency and the parallelism of the hardware resources.
10. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein: When the processor executes the computer-readable instructions, the processor implements the steps of the RDMA primitive optimization method for multimodal distributed storage according to any one of claims 1 to 8.
Citation Information
Patent Citations
An RDMA application transmission parameter adaptive selection method in a data center
CN109831321A
RDMA communication method and device, electronic equipment and storage medium
CN117354153A
Data transmission method and device for remote memory access, equipment and storage medium
CN118093499A
Machine learning techniques for implementing tree-based network congestion control
US20240007403A1
Method and apparatus for optimizing switching-resource allocation in multi-modal network, and medium
WO2024093041A1
Cited By
DMA performance optimization system and method, electronic equipment and storage medium
CN122220272A