Distributed system linear consistency detection method based on programmable switch

Through the method of real-time pre-verification and global collaborative verification in programmable switches, the linear consistency detection problem of distributed systems in high concurrency and high throughput environments is solved, and fast and accurate system consistency detection is achieved.

CN120263687APending Publication Date: 2025-07-04FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510418730.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

It is difficult for existing distributed systems to achieve fast and accurate linear consistency verification in high concurrency and high throughput environments. Traditional server-side stress testing methods have limitations in traffic generation and verification delays.

Method used

By utilizing the circular forwarding, replication and lightweight local linearized verification strategies of template packets in a programmable switch, real-time pre-verification is performed, and the data packets that fail local verification are fed back to the control server for global verification. Combined with the fault injection module to simulate exception scenarios, high-throughput data generation and fast consistency detection are achieved.

Benefits of technology

It significantly improves the efficiency and accuracy of distributed system consistency verification, reduces verification delay and resource consumption, and can quickly detect potential errors in high concurrency environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263687A_ABST
    Figure CN120263687A_ABST
Patent Text Reader

Abstract

The invention relates to a distributed system linear consistency detection method based on a programmable switch, and the method is characterized in that the automatic configuration of a test task is realized through a high-level test primitive, and the linear consistency of a distributed system is detected through the cyclic forwarding, copying and lightweight local linearization verification strategy of a template data packet in the programmable switch. According to the embodiment of the invention, real-time pre-verification is performed on the distributed system, and meanwhile, the data packet failed in local verification is fed back to the control server for global verification, so that the verification delay is reduced, the resource consumption is reduced, and the system consistency verification efficiency is improved. According to the method, the consistency of the distributed system can be rapidly and accurately detected in a high-concurrency and high-throughput network environment, and the limitation of a traditional server pressure testing method in the aspects of flow generation and time delay verification is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed systems, and in particular to a method for detecting linear consistency of a distributed system based on a programmable switch. Background Art

[0002] When ensuring data consistency in existing distributed systems, it usually relies on the server side for stress testing: triggering potential errors by simulating high-concurrency read and write operations and fault scenarios, and recording the node operation history to verify linearizability consistency. However, as the processing scale and throughput of distributed systems continue to expand, it is difficult for the server side to balance large-scale data sending and efficient analysis under high traffic, resulting in some errors that can only be exposed under high-concurrency conditions being missed. At the same time, linearizability verification often requires replaying and analyzing a large number of operation histories, with high computational costs and difficulty in meeting the requirements for real-time and accuracy.

[0003] Programmable switches have a data processing capacity of Tbps level, and can directly generate, forward, and process a large number of data packets in the network, providing a new idea for distributed system verification. However, due to limitations in the programming model and hardware resources (such as registers, storage space, instruction logic) of switches, it is difficult to independently undertake the complete linearizability verification task. Therefore, there is an urgent need for a verification scheme that combines programmable switches and server-side collaborative processing to simultaneously achieve high-throughput test data generation and fast consistency verification, thereby improving the verification efficiency and coverage of distributed systems in a high-concurrency environment. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for detecting linear consistency of a distributed system based on a programmable switch, which can quickly and accurately detect the consistency of a distributed system in a high-concurrency and high-throughput network environment, and overcome the limitations of traditional server-side stress testing methods in terms of traffic generation and verification latency.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is: a method for detecting linear consistency of a distributed system based on a programmable switch, characterized in that it realizes automatic configuration of test tasks through high-level test primitives, and uses circular forwarding, replication, and lightweight local linearizability verification strategies of template data packets inside the programmable switch to perform real-time pre-verification on the distributed system, and at the same time feeds back the data packets with local verification failures to the control server for global verification, thereby reducing verification latency, reducing resource consumption, and improving the efficiency of system consistency verification.

[0006] Furthermore, the method includes the following steps: S1) High-level test primitive construction: Define multiple high-level test primitives to form a set of high-level test primitives; convert the high-level test primitives into data plane configuration instructions applicable to programmable switches and task scheduling instructions for control servers, so as to realize the automated description and modular configuration of test tasks; S2) Traffic generation in programmable switches: In programmable switches, generate high-throughput test data streams at the Tbps level through the loop forwarding and replication mechanism of constructed template packets to simulate high-concurrency access scenarios for large-scale distributed systems; S3) Local linearization verification: During the data stream transmission process, use the preset storage registers in the programmable switch to record the operations and timestamp information of the most recent several packets, and quickly judge the read-write consistency in combination with the latest status registers; when potential conflicts are found, extract the local operation set and perform simplified linearization verification; S4) Global verification and fault injection: Submit the operation sequences that cannot be determined by local verification to the control server for in-depth playback and global linearization analysis; at the same time, simulate abnormal scenarios including network partitions and node outages through the fault injection module to improve the verification coverage rate; S5) Result feedback and dynamic tuning: Dynamically adjust the traffic scale, fault strategy, and primitive configuration according to the results of the collaborative verification between the programmable switch and the server to maintain high verification efficiency and accuracy in different scale and load environments.

[0007] Furthermore, the high-level test primitives include: Set_protocol: Set the communication protocol; Set_nodes: Configure distributed nodes; Set_template_queries: Define template packets; Set_rate: Set the data packet transmission rate; Set_ratio: Set the read-write request ratio; Set_faults_type and Set_next_fault: Configure the fault injection type and trigger interval; Thus, the automated description and modular configuration of test tasks are realized.

[0008] Furthermore, the lightweight local linearization verification is implemented in the programmable switch through the following steps: 1) Set storage registers in the programmable switch to store the operation records and timestamps of a preset number of the most recent data packets; during the data stream transmission process, use the storage registers to record the operations and timestamp information of the most recent several packets; 2) Use the latest status register to update the status of the received write operation, and compare the read operation with the current status in real time to preliminarily judge the consistency of the data packet operation sequence; 3) When detecting a mismatch between read and write operations, extract part of the operation records from the storage register to form an operation set, generate all possible operation sequences, and perform local linearization verification on them; 4) Adopt a pre-judgment and pruning strategy during the local verification process to reduce hardware resource consumption and ensure the efficiency of verification.

[0009] Furthermore, the control server and the programmable switch adopt a collaborative verification strategy, that is, while completing real-time pre-verification inside the programmable switch, for data packets that do not meet the consistency requirements in local verification, automatically forward them to the control server for global detailed verification, thereby optimizing the overall efficiency of system verification.

[0010] Furthermore, the programmable switch discovers potential consistency conflicts in real time by adding lightweight local verification logic and completing high-speed data packet scheduling in the data plane; the control server is responsible for global linearization verification of large-scale operation histories and provides a basis for subsequent test parameter adjustment.

[0011] Furthermore, the template data packets in the programmable switch generate high-throughput data streams through a loop forwarding and replication mechanism, and embed operation events and timestamp information in the data packets to provide sufficient data basis for local linearization verification.

[0012] Furthermore, the local linearization verification further includes: when detecting inconsistent read and write operations, generate all possible permutations and combinations of the operation records extracted from the storage register, and combine chronological verification to verify whether the operation set can meet the global linearization requirements, thereby ensuring that the final verification result has high accuracy and consistency.

[0013] Compared with the prior art, the present invention has the following beneficial effects: 1) High-throughput data generation: The programmable switch directly constructs and replicates template data packets in the data plane to achieve traffic output at the Tbps level, significantly improving the simulation ability for high-concurrency scenarios; 2) Hierarchical verification collaboration: The built-in local verification in the switch reduces the burden on the server for massive data playback and analysis, reduces verification latency and improves real-time performance; 3) Flexible fault injection: Combining the programmable switch and the control server, faults can be injected from multiple perspectives at the network layer and application layer to enhance the coverage of potential errors; 4) Strong adaptability: Through test primitives and dynamic tuning mechanisms, it can be flexibly configured according to the scale and scenario requirements of different distributed systems to achieve a balance between verification efficiency and accuracy.

[0014] Therefore, the present invention provides an effective technical solution for the consistency verification of distributed systems in high-concurrency and high-throughput environments, laying a solid foundation for quickly locating system defects and improving overall reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is the overall framework diagram of the method according to an embodiment of the present invention.

[0016] Figure 2 is an example diagram of the data packet acceleration method in an embodiment of the present invention.

[0017] Figure 3 is an example diagram of the linear consistency verification within a programmable switch in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0019] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0020] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0021] This embodiment provides a method for linear consistency detection of a distributed system based on a programmable switch, which realizes the automatic configuration of test tasks through high-level test primitives, and uses the loop forwarding, replication, and lightweight local linearization verification strategy of template data packets inside the programmable switch to perform real-time pre-verification on the distributed system. At the same time, the data packets with local verification failures are fed back to the control server for global verification, thereby reducing verification latency, reducing resource consumption, and improving the efficiency of system consistency verification. This method is specifically implemented through the following steps.

[0022] S1) Construction of high-level test primitives Define a plurality of high-level test primitives to form a set of high-level test primitives. The high-level test primitives include: Set_protocol: Set the communication protocol; Set_nodes: Configure distributed nodes; Set_template_queries: Define the template data packet; Set_rate: Set the data packet transmission rate; Set_ratio: Set the read / write request ratio; Set_faults_type and Set_next_fault: Configure the fault injection type and trigger interval.

[0023] Convert high-level test primitives into data plane configuration instructions applicable to programmable switches and task scheduling instructions for control servers, so as to realize the automated description and modular configuration of test tasks.

[0024] S2) Traffic generation in the programmable switch In the programmable switch, generate high-throughput test data streams at the Tbps level by constructing a loop forwarding and replication mechanism for the template data packet, and simulate high-concurrency access scenarios for large-scale distributed systems.

[0025] S3) Local linearization verification During the data stream transmission process, use the preset storage registers in the programmable switch to record the operations and timestamp information of the most recent several data packets, and quickly judge the read / write consistency in combination with the latest status register; when potential conflicts are found, extract the local operation set and perform simplified linearization verification.

[0026] The lightweight local linearization verification is implemented in the programmable switch through the following steps: 1) Set storage registers in the programmable switch to store the operation records and timestamps of a preset number of the most recent data packets; during the data stream transmission process, use the storage registers to record the operations and timestamp information of the most recent several data packets; 2) Use the latest status register to update the status of the received write operations, and compare the read operations with the current status in real time to preliminarily judge the consistency of the data packet operation sequence; 3) When it is detected that the read / write operations do not match, extract some operation records from the storage registers to form an operation set, and generate all possible operation sequences, and perform local linearization verification on them; 4) Adopt a pre-judgment and pruning strategy during the local verification process to reduce the consumption of hardware resources and ensure the efficiency of verification.

[0027] The local linearization verification further includes: when it is detected that the read / write operations are inconsistent, generate all possible permutations and combinations of the operation records extracted from the storage registers, and verify whether the operation set can meet the global linearization requirements in combination with the time sequence check, so as to ensure that the final verification result has high accuracy and consistency.

[0028] S4) Global Verification and Fault Injection Submit the operation sequences that cannot be determined by local verification to the control server for in-depth replay and global linearization analysis; at the same time, simulate abnormal scenarios including network partitioning and node downtime through the fault injection module to improve the verification coverage rate.

[0029] S5) Result Feedback and Dynamic Tuning Dynamically adjust the traffic scale, fault strategy, and primitive configuration according to the results of the collaborative verification between the programmable switch and the server, so as to maintain high verification efficiency and accuracy in different scale and load environments.

[0030] In this method, the control server and the programmable switch adopt a collaborative verification strategy, that is, while completing real-time pre-verification inside the programmable switch, for the data packets that do not meet the consistency requirements in local verification, automatically forward them to the control server for global detailed verification, thereby optimizing the overall efficiency of system verification. The programmable switch discovers potential consistency conflicts in real time by adding lightweight local verification logic and completing high-speed data packet scheduling in the data plane; the control server is responsible for global linearization verification of large-scale operation histories and provides a basis for subsequent test parameter adjustment.

[0031] The following further details this method in combination with the accompanying drawings and specific embodiments.

[0032] 1. Overall Overview As Figure 1 shown, the overall framework of this method includes the collaborative work of the Monica control end and the programmable switch. The control end is mainly composed of a Primitive Handler and a Consistency Analyzer, which are used to receive user-defined high-level test primitives (such as Set_protocol, Set_nodes, etc.), parse them into data plane instructions executable by the programmable switch, and perform global consistency analysis on the local verification results from the switch in subsequent stages. Inside the programmable switch, multiple key modules are integrated, including an Accelerator, a Replicator, an Editor, an Invoke Generator, and a Nemesis module, etc., to complete the generation of large-scale data streams and the simulation of fault scenarios; at the same time, a Verification unit for read-write comparison and local linearization verification is also deployed in the switch. Multiple distributed nodes (Host 1, Host2, Host 3, Host 4) are connected to the programmable switch through the network to form a complete distributed system test environment.

[0033] In this overall design, the Monica control end and the programmable switch cooperate through the P4 program and primitive configuration: the control end is responsible for mapping the test intent into the switch's configuration process and collecting and analyzing the verification information sent back by the switch during the test; the switch directly sends and processes data packets at the Tbps rate in the data plane and performs quick consistency verification on local suspicious operations.

[0034] 2. Traffic Generation within the Programmable Switch As Figure 2 shown, the programmable switch efficiently generates test data streams through the method of "template data packets + loop forwarding and replication". The core process includes the following steps: i) Primitive Handler: After the Monica control end parses the user-defined test primitives, it sends the key parameters required for generating the data stream (such as protocol type, read / write ratio, fault type, etc.) to the switch.

[0035] ii) Accelerator: Realize the loop forwarding of the template data packets within the switch data plane; these template data packets are regarded as the "source" of high-throughput data and can flow repeatedly in the switch pipeline.

[0036] iii) Replicator: Copy the template data packets entering this module and adjust the sending rate as needed to form a high-traffic data stream with multi-port concurrent output, easily achieving traffic from 10 Mbps to hundreds of Gbps or even Tbps level.

[0037] iv) Editor: Perform header and payload modification operations on the real data packets formed after replication, including setting the read / write operation type, injecting random fields, etc., so as to flexibly simulate various real or extreme request scenarios in the distributed system.

[0038] Through the above mechanism, the programmable switch can directly generate a large number of read / write requests in the data plane and distribute them to each distributed node (Host), effectively simulating the system operation state in a high-concurrency and high-throughput environment.

[0039] 3. Local Linearization Verification within the Switch As Figure 3 shown, the present invention further deploys lightweight local linearization verification logic inside the switch to discover potential consistency problems in real time during the high-speed data forwarding process, mainly including the following functional modules: i) Storage Register (SR): Used to cache the operation information of several recent data packets (such as timestamp, operation type, operation value, etc.), and the storage quantity is usually within dozens of packets, which is convenient for quick comparison under the hardware limitations of the switch.

[0040] ii) Latest Status Register (LSR): It is used to record the reference status updated by the most recent write operation. When a new read operation arrives, it is quickly compared with the value in the LSR. If they match, it is considered to temporarily meet the consistency; if not, further analysis is required.

[0041] iii) Local Linearization Verification: When there is a mismatch between read and write operations, a small number of recent operations are extracted from the SR to form an operation set, and all possible execution orders are generated and simple linearization tests are performed; a pruning strategy is adopted in the test to reduce unnecessary combination traversal, so as to efficiently complete local verification under the limited computing resources of the switch.

[0042] iv) Result Processing: If the local verification cannot draw a conclusion of consistency or determine the existence of potential conflicts, the corresponding data packet and operation sequence are sent back to the Monica control end, and the Consistency Analyzer performs a more comprehensive global linearization analysis.

[0043] Through the above local verification logic, the present invention effectively reduces the pressure on the control end: only when the local verification cannot judge or a suspected error occurs, is it necessary to perform a more expensive global analysis, thereby greatly reducing the verification delay and improving the system test coverage.

[0044] 4. Cooperative Verification and Fault Injection In the above basic process, the present invention also introduces the Nemesis module and the fault injection mechanism, allowing various abnormal scenarios such as network latency, partitioning, and node downtime to be flexibly simulated at the switch end. Specifically, the control end can set the fault type and trigger time through primitive methods (Set_faults_type, Set_next_fault, etc.), and the switch inserts or simulates these faults in real time during the data flow generation and forwarding process. Thus, the race conditions and errors that may occur in the distributed system in the face of high load and fault environment are more easily exposed, further improving the accuracy and completeness of verification.

[0045] After the fault is triggered, if the local verification logic in the switch captures a potential conflict, the corresponding data will be sent back to the control end; at this time, the Consistency Analyzer will combine the complete operation history and the fault sequence to globally determine the consistency of the distributed system, so as to achieve an organic combination of high-speed processing by the switch and in-depth analysis by the server.

[0046] In summary, through the collaborative processing method of the switch and the control end, the present invention provides an effective technical solution for quickly detecting consistency problems in distributed systems in a high-throughput and high-concurrency environment, which can significantly reduce the verification delay and resource consumption while ensuring relatively high accuracy.

[0047] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0048] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0049] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0050] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0051] As mentioned above, it is only the preferred embodiments of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for detecting linear consistency of a distributed system based on a programmable switch, characterized in that Automatically configure test tasks through high-level test primitives, and use the loop forwarding, replication, and lightweight local linearization verification strategy of template packets inside the programmable switch to perform real-time pre-verification on the distributed system. At the same time, feedback the packets with local verification failures to the control server for global verification, thereby reducing verification latency, reducing resource consumption, and improving the efficiency of system consistency verification.

2. The linear consistency detection method for a distributed system based on a programmable switch according to claim 1, wherein Including the following steps: S1) High-level test primitive construction: Define multiple high-level test primitives to form a high-level test primitive set; convert the high-level test primitives into data plane configuration instructions applicable to the programmable switch and task scheduling instructions for the control server, so as to realize the automated description and modular configuration of test tasks. S2) Traffic generation in the programmable switch: In the programmable switch, generate a high-throughput test data stream at the Tbps level through the loop forwarding and replication mechanism of the constructed template packets to simulate a high-concurrency access scenario for large-scale distributed systems. S3) Local linearization verification: During the data stream transmission process, use the preset storage registers in the programmable switch to record the operations and timestamp information of the most recent several packets, and quickly judge the read-write consistency in combination with the latest status register; when potential conflicts are found, extract the local operation set and perform simplified linearization verification. S4) Global verification and fault injection: Submit the operation sequences that cannot be determined by local verification to the control server for in-depth playback and global linearization analysis; at the same time, simulate abnormal scenarios including network partitions and node outages through the fault injection module to improve the verification coverage rate. S5) Result feedback and dynamic tuning: Dynamically adjust the traffic scale, fault strategy, and primitive configuration according to the results of the collaborative verification between the programmable switch and the server to maintain high verification efficiency and accuracy in different scale and load environments.

3. The linear consistency detection method for a distributed system based on a programmable switch according to claim 2, characterized in that The high-level test primitives include: Set_protocol: Set the communication protocol; Set_nodes: Configure distributed nodes; Set_template_queries: Define template packets; Set_rate: Set the packet transmission rate; Set_ratio: Set the read-write request ratio; Set_faults_type and Set_next_fault: Configure the fault injection type and trigger interval; Thus, realize the automated description and modular configuration of test tasks.

4. A method for detecting linear consistency of a distributed system based on a programmable switch according to claim 2, characterized in that, The lightweight local linearization verification is implemented in the programmable switch through the following steps: 1) Set storage registers in the programmable switch to store the operation records and timestamps of a preset number of the most recent packets; during the data stream transmission process, use the storage registers to record the operations and timestamp information of the most recent several packets. 2) Use the latest status register to update the status of the received write operations, and compare the read operations with the current status in real time to initially judge the consistency of the packet operation sequence. 3) When detecting a mismatch between read and write operations, extract part of the operation records from the storage register to form an operation set, generate all possible operation sequences, and perform local linearization verification on them. 4) During the local verification process, a pre-judgment and pruning strategy is adopted to reduce hardware resource consumption and ensure the efficiency of verification.

5. A method for detecting linear consistency of a distributed system based on a programmable switch according to claim 2, characterized in that, The control server and the programmable switch adopt a cooperative verification strategy, that is, while the real-time pre-verification is completed inside the programmable switch, for the data packets that do not meet the consistency requirements in the local verification, they are automatically forwarded to the control server for global detailed verification, thereby optimizing the overall efficiency of the system verification.

6. A method for detecting linear consistency of a distributed system based on a programmable switch according to claim 2, characterized in that, The programmable switch discovers potential consistency conflicts in real time by adding lightweight local verification logic and completing high-speed data packet scheduling in the data plane; the control server is responsible for the global linearization verification of the large-scale operation history and provides a basis for subsequent test parameter adjustment.

7. A method for detecting linear consistency of a distributed system based on a programmable switch according to claim 2, characterized in that, The template data packets in the programmable switch generate high-throughput data streams through a loop forwarding and replication mechanism, and embed operation events and timestamp information in the data packets to provide sufficient data basis for local linearization verification.

8. A linear consistency detection method for a distributed system based on a programmable switch according to claim 4, characterized in that The local linearization verification further includes: when detecting inconsistent read and write operations, all possible permutations and combinations are generated from the operation records extracted from the storage register, and combined with the time sequence check to verify whether the operation set can meet the global linearization requirements, thereby ensuring that the final verification result has high accuracy and consistency.