Multi-variant recovery method and device based on record playback
By introducing monitors and system call coordinators into the multivariate execution system, recording and replaying key information of system calls, the problem of state inconsistency of the multivariate execution system after attack is solved, and efficient state recovery and interruption-free operation is achieved.
Patent Information
- Application Number
- CN202510316197.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-18
AI Technical Summary
After detecting an attack exception, the multivariate execution system lacks an effective method to restore the system state, resulting in the problem of inconsistent state before and after recovery.
The system call coordinator of variants is monitored and intercepted through the monitor, the system call coordinator is used to record key information, and the state recovery is restored using an asynchronous execution mechanism in exceptional situations. The system call coordinator module is designed to ensure consistency of the state before and after recovery.
The state consistency in the state recovery process of the multi-variable execution system is achieved, the recovery efficiency is improved, and the interruption is achieved.
Smart Images

Figure CN120337208A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and more particularly, to a multi-variant recovery method and device based on record replay. Background Art
[0002] In today's digital age, the rapid development of information technology has brought great convenience and efficiency improvement to all walks of life, but at the same time, it has also brought severe security challenges. With the continuous increase in the number of programs and codes, various exploitable vulnerabilities have emerged in an endless stream. Attackers can use the vulnerabilities and defects of the system to carry out various malicious attack behaviors, resulting in the risk of being invaded and attacked in the actual supply, production and other links at all times. Among them, software homogenization is the main reason for the spread and diffusion of vulnerabilities. A large number of the same or similar open-source frameworks and code libraries used in system development reduce the cost of system attacks while improving production efficiency, enabling attackers to quickly damage similar systems using the same attack techniques, further increasing the security threats faced by the cyberspace.
[0003] In order to reverse the passive situation faced by cyberspace security, many scholars have begun to study new defense means. Among them, active defense and endogenous security have become the mainstream research directions, and progress has been made in recent years in directions such as intrusion tolerance, honeypot, Moving Target Defence (MTD), and Mimic Defence (MD). However, technical means such as intrusion tolerance and detection need to rely on certain prior knowledge and will bring latency, and will appear powerless when dealing with unknown attacks based on unknown vulnerabilities. In recent years, some researchers have applied MultiVariant Execution (MVX) to the field of network security, using the ideas of diverse isomers and redundant execution therein to address the security threats brought by software homogenization and to discover security threats based on unknown backdoors.
[0004] However, there is a problem of variant state recovery in the current multi-variant execution system: after detecting an attack anomaly and cleaning the current execution body set, the new execution body set should be synchronized to the state before cleaning in a friendly manner. However, most multi-variant execution frameworks do not specifically give a method for system state recovery after being attacked. Due to the heterogeneity and redundancy of the multi-variant execution system, directly adopting traditional fault recovery methods such as rollback recovery and record replay will result in problems such as incompatible checkpoint files and inconsistent variant states. Summary of the Invention
[0005] In view of this, the present invention provides a multi-variant recovery method based on record replay to solve the problem of variant recovery after cleaning and the problem of inconsistent states of variants before and after recovery.
[0006] To solve the above technical problems, a first aspect of the present invention provides a multi-variant recovery method based on record replay, including:
[0007] Using a monitor to monitor the system calls of variants, wherein a child process is generated by a thread as a variant;
[0008] Intercepting and adjudicating the system calls of variants by the monitor according to the key information of the system calls, wherein the key information of the system calls includes the system call number, system call parameters, and return results. For the system calls adjudicated as legal, the key information of the system calls is recorded by the system call coordinator, and the system call coordinator is a component of the monitor;
[0009] When the system call of a certain variant is adjudicated as illegal, then the variant is an abnormal variant. When the abnormal variant is taken offline and then brought back online, an asynchronous execution mechanism is adopted for state recovery. Among them, when the variant in state recovery executes a system call, the system call coordinator determines an execution policy according to the category of the system call, and then replays the system call of the variant in state recovery according to the execution policy and the recorded key information of the legal system calls.
[0010] In one implementation, using a monitor to monitor the system calls of variants includes:
[0011] The monitor uses a preset tool to monitor the system calls of variants. Among them, the variant synchronizes at each system call and waits for the system to adjudicate it. The monitor is the core module in the variant parent process and is responsible for the monitoring and synchronization of system calls.
[0012] In one implementation, intercepting and adjudicating the system calls of variants by the monitor according to the key information of the system calls includes:
[0013] After the variant synchronizes at the system call, the key information of the system call is written into the first shared buffer, and the first shared buffer is used for the monitor to adjudicate the variant and allows the variant to access it;
[0014] The monitor makes a ruling on system calls based on the information in the first shared buffer. If the system call numbers, system call parameters, and return results of all variants are the same, the ruling result is legal; otherwise, the ruling result is illegal. Among them, for system calls ruled as legal, the system call coordinator records the key information in the second shared buffer. The second shared buffer is used for recording system calls and allows other variants to access it. When the second shared buffer is full, the system call coordinator batches the content in the second shared buffer and writes it to an external file for persistent storage, and then clears the second shared buffer.
[0015] In one implementation, when restarting an abnormal variant, an asynchronous execution mechanism is used for state recovery, including:
[0016] The system call coordinator is solely responsible for the state recovery of the abnormal variant without interfering with the monitor's interception and ruling of normal variants;
[0017] When the system call coordinator performs state recovery on the abnormal variant, it continues to record the key information of the system calls of the normal variants.
[0018] In one implementation, system call replay is performed according to the execution policy and the key information recorded by the system call coordinator, including:
[0019] The system call coordinator batches and reads the key information of system calls from the external file into the third shared buffer in the recorded order. The third shared buffer is used for system call replay. After all the information in the third shared buffer is used, all the content will be cleared and the key information of system calls will be read again from the external file;
[0020] The system call coordinator captures the system calls of the variant undergoing state recovery;
[0021] Determine the corresponding execution policy according to the category of the system call;
[0022] The system call coordinator performs deterministic replay of the system calls of the variant in state recovery according to the execution policy and the key information of the read system calls.
[0023] In one implementation, determining the corresponding execution policy according to the category of the system call includes:
[0024] For non-deterministic and sensitive system calls, skip execution during the recovery phase and send the recorded return value to the variant;
[0025] For external input type system calls, return the pre-recorded input data to the variant;
[0026] For process or thread creation system calls, use a virtual process identifier to replace the real process identifier and send it to the variant.
[0027] In one embodiment, the method further includes determining whether recovery has been completed according to the number of system calls for which replay has been completed.
[0028] Based on the same inventive concept, a second aspect of the present invention provides a multi-variant recovery device based on record replay, including:
[0029] A monitoring module, configured to monitor the system calls of the variant by using a monitor, wherein a child process is generated by a thread as the variant;
[0030] An interception and adjudication module, configured to intercept and adjudicate the system calls of the variant by the monitor according to the key information of the system call. The key information of the system call includes the system call number and the system call parameters. For a system call adjudged to be legal, the key information of the system call is recorded by a system call coordinator, and the system call coordinator is a component of the monitor;
[0031] A recovery module, configured to when the system call of a certain variant is adjudged to be illegal, then the variant is an abnormal variant, perform a take-offline process on the abnormal variant, and when the abnormal variant is brought back online, use an asynchronous execution mechanism to perform state recovery. When the variant in the state recovery executes a system call, the system call coordinator determines an execution policy according to the category of the system call, and then replays the system call of the variant in the state recovery according to the execution policy and the recorded key information of the legal system call. The category of the system call includes an indeterminate system call, a sensitive system call, an external input system call, and a process or thread creation system call.
[0032] In one embodiment, the monitoring module is specifically configured to:
[0033] The monitor uses a preset tool to monitor the system calls of the variant. The variant synchronizes at each system call and waits for the system to adjudicate it. The monitor is a core module in the variant parent process and is responsible for monitoring and synchronizing system calls.
[0034] In one embodiment, the interception and adjudication module is specifically configured to:
[0035] After the variant synchronizes at the system call, write the key information of the system call into a first shared buffer, and the first shared buffer is used for the monitor to adjudicate the variant and allows the variant to access it;
[0036] The monitor makes a ruling on system calls based on the information in the first shared buffer. If the system call numbers, system call parameters, and return results of each variant are all the same, the ruling result is legal; otherwise, the ruling result is illegal. Among them, for the system calls ruled as legal, the system call coordinator records the key information in the second shared buffer, which is used for recording system calls and allows other variants to access. When the second shared buffer is full, the system call coordinator batches the content in the second shared buffer and writes it to an external file for persistent storage, and then clears the second shared buffer.
[0037] In one implementation, the recovery module is further configured to:
[0038] The system call coordinator is solely responsible for the state recovery of the variant and does not interfere with the monitor's interception and ruling of normal variants;
[0039] When the system call coordinator performs state recovery on an abnormal variant, it continues to record the key information of the system calls of normal variants.
[0040] In one implementation, the recovery module is further configured to:
[0041] The system call coordinator batches and reads the key information of system calls from the external file into the third shared buffer in the recorded order. The third shared buffer is used for replaying system calls. After all the information in the third shared buffer is used, all the content will be cleared and the key information of system calls will be read from the external file again;
[0042] The system call coordinator captures the system calls of the variant whose state is being recovered;
[0043] Determine the corresponding execution policy according to the category of the system call;
[0044] The system call coordinator performs deterministic replay of the system call according to the execution policy and the key information of the read system call.
[0045] In one implementation, the recovery module is further configured to:
[0046] For non-deterministic and sensitive system calls, skip execution during the recovery phase and send the recorded return value to the variant;
[0047] For external input class system calls, return the pre-recorded input data to the variant;
[0048] For system calls related to process or thread creation, use a virtual process identifier to replace the real process identifier and send it to the variant.
[0049] In one embodiment, the method further includes a judgment module for judging whether the recovery has been completed according to the number of system calls that have been replayed.
[0050] Based on the same inventive concept, the third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for multi-variant recovery based on recording and replay described in the first aspect.
[0051] Based on the same inventive concept, the fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for multi-variant recovery based on recording and replay described in the first aspect.
[0052] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:
[0053] The present invention provides a method for multi-variant recovery based on recording and replay. Threads are introduced into the multi-variant execution system, and the child processes generated by the threads are variants. The monitor monitors the system calls of the variants; the system coordinator records the key information of the system calls and judges whether the system calls of the variants are legal according to the recorded key information of the system calls; when the system call adjudication of a certain variant is illegal, then the variant is an abnormal variant, and the abnormal variant is cleaned (offline processing), and the status of the variant that is re-onlined is restored. The present invention improves the monitor and designs a system call coordinator module for recording the key information of the system calls and accurately replaying the system calls of the variants in the status recovery, solving the problem of inconsistent status before and after recovery. Further, an asynchronous execution mechanism is designed in the variant recovery stage, improving the recovery efficiency and achieving interruption-free operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0055] Figure 1 It is the overall flowchart of the method for multi-variant recovery based on recording and replay in the embodiments of the present invention;
[0056] Figure 2 It is the detailed flowchart of the method for multi-variant recovery based on recording and replay in the embodiments of the present invention;
[0057] Figure 3This is the architecture diagram of the system call coordinator module in the embodiments of the present invention;
[0058] Figure 4 This is the execution flowchart for state recovery of a resurrected variant using an asynchronous execution mechanism in the embodiments of the present invention;
[0059] Figure 5 This is the uninterrupted operation diagram provided in the embodiments of the present invention. Detailed implementation manners
[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0061] Embodiment 1
[0062] This embodiment discloses a multi-variant recovery method based on record and replay. Please refer to Figure 1 , including:
[0063] S1: Use a monitor to monitor the system calls of a variant; wherein, a child process is generated through a thread as the variant;
[0064] S2: Intercept and adjudicate the system calls of the variant by the monitor according to the key information of the system calls. For the system calls adjudged to be legal, the key information of the system calls is recorded by the system call coordinator, and the system call coordinator is a component of the monitor. Among them, the key information of the system calls includes the system call number, system call parameters, and return results;
[0065] S3: When the system call of a certain variant is adjudged to be illegal, then the variant is an abnormal variant. When the abnormal variant is taken offline and then brought back online, an asynchronous execution mechanism is used for state recovery. Among them, when the variant in the state recovery executes a system call, the system call coordinator determines an execution policy according to the category of the system call, and then replays the system call according to the execution policy and the recorded key information of the legal system calls. Among them, the categories of the system calls include non-deterministic system calls, sensitive system calls, external input system calls, and system calls for process or thread creation.
[0066] Specifically, steps S1 and S2 mean that during the normal operation of the system, the monitor intercepts the system calls of the variant and makes a ruling. For the system calls ruled to be legal, the system call coordinator records their key information for use in the state recovery of the variant. Step S3 means that after the monitor intercepts and synchronizes at a certain system call, if it is found that there is an abnormality in the system call of a certain variant, at this time, the system call of this variant is ruled illegal, and the monitor will clean the abnormal variant, that is, take the abnormal variant offline, and then recover the state of the variant that is restarted online. During the state recovery process, according to the execution policy and the key information of the recorded legal system calls, the system calls of the variant during state recovery are replayed.
[0067] In one implementation, monitoring the system calls of the variant by using a monitor includes:
[0068] The monitor uses a preset tool to monitor the system calls of the variant. Among them, the variant synchronizes at each system call and waits for the system to make a ruling on it. The monitor is the core module in the variant parent process and is responsible for the monitoring and synchronization of system calls.
[0069] In the specific implementation process, threads are introduced into the multi-variant execution system. Among them, n child processes are generated through the threads to be used as variants, and n represents the redundancy in the multi-variant execution system. The variant synchronizes at each system call and waits for the system to make a ruling on it. The variant is a diversified program running by the child process, and the monitor is a module in the multi-variant execution system. In the specific implementation process, the ptrace can be used to monitor the system calls of the variant.
[0070] The monitor is responsible for monitoring the system calls of the variant. Specifically, the monitor intercepts the system calls of the variant and makes a consistency ruling on them. For the system calls ruled to be legal, the monitor records the key information of the system calls in the system call coordinator for use in the subsequent recovery of the variant. The interception and ruling are completed by the monitor, and its component, the system call coordinator, is responsible for recording the legal system calls for use in the state recovery of the variant.
[0071] In one implementation, intercepting and ruling the system calls of the variant by the monitor according to the key information of the system calls includes:
[0072] After the variant synchronizes at the system call, it writes the key information of the system call into the first shared buffer, and the first shared buffer is used for the monitor to make a ruling on the variant and allows the variant to access it;
[0073] The monitor makes a ruling on system calls based on the information in the first shared buffer. If the system call numbers, system call parameters, and return results of all variants are the same, the ruling result is legal; otherwise, the ruling result is illegal. Among them, for the system calls ruled as legal, the system call coordinator records the key information in the second shared buffer, which is used for recording system calls and allows other variants to access. When the second shared buffer is full, the system call coordinator batches the content in the second shared buffer and writes it to an external file for persistent storage, and then clears the second shared buffer.
[0074] Specifically, after intercepting a system call, the monitor writes the content into a shared buffer (the first shared buffer, shared buffer A, the memory of the monitor). If the ruling is legal, the key information of the system call will be saved in another buffer (the second shared buffer, shared buffer B). The system call coordinator will regularly write the key information of the system calls saved in shared buffer B to an external file for persistent storage.
[0075] In the specific implementation process, in addition to recording the key information of system calls, the system call coordinator is also used to classify system calls.
[0076] In one implementation, when restarting an abnormal variant, an asynchronous execution mechanism is adopted for state recovery, including:
[0077] The system call coordinator is solely responsible for the state recovery of the variant without interfering with the synchronization and voting of the monitor for normal variants;
[0078] When the system call coordinator performs state recovery on the variant, it continues to record the key information of the system calls of normal variants.
[0079] Specifically, the asynchronous execution mechanism allows the cleaned variant to independently recover during the recovery phase and join the execution after reaching the synchronization point of other variants. The synchronization point is set at the entry point where the variant executes the system call. The variant will synchronize at the synchronization point during normal execution, and the monitor makes a ruling on the variant at the synchronization point.
[0080] In one implementation, system call replay is performed according to the execution policy and the key information recorded by the system call coordinator, including:
[0081] The system call coordinator batches reads the key information of system calls from the external file in the recorded order into the third shared buffer (shared buffer C), and the third shared buffer is used for the replay of system calls. After all the information in the third shared buffer is used, all the content will be cleared and the key information of system calls will be read again from the external file;
[0082] The system call coordinator captures the system calls of the variant that are in the process of state recovery;
[0083] Determine the corresponding execution policy according to the category of system calls;
[0084] The system call coordinator performs deterministic replay on the system calls of the variant in state recovery according to the execution policy and the key information of the read system calls.
[0085] In the specific implementation process, when performing system call replay during the state recovery stage of the variant, the system call coordinator batches reads the key information recorded in step S2 from the external file into the third shared buffer for the replay of system calls.
[0086] The replay process of system calls is based on the category of system calls and the recorded key information, and adopts the corresponding execution policy to ensure that the system calls executed during the recovery stage are consistent with the system calls executed before recovery in terms of function and effect, so as to ensure the integrity and consistency of the variant recovery state.
[0087] In one implementation manner, determining the corresponding execution policy according to the category of system calls includes:
[0088] For non-deterministic and sensitive system calls, skip execution during the recovery stage and send the recorded return value to the variant;
[0089] For external input class system calls, return the pre-recorded input data to the variant;
[0090] For process or thread creation class system calls, use a virtual process identifier to replace the real process identifier and send it to the variant.
[0091] In one implementation manner, the method further includes judging whether the recovery has been completed according to the number of system calls that have completed replay.
[0092] Specifically, if the recovery is not completed, the variant continues to recover. When encountering the next system call, the variant will pass it to the system call coordinator. If the recovery is completed, the variant ends the asynchronous recovery and joins the synchronous execution of other variants.
[0093] The variant recovery method of the present invention will be described below through a specific process.
[0094] As Figure 2 shown, a variant recovery method based on record replay includes the following steps:
[0095] Step S101, introduce a thread in the multi-variant execution system and monitor the system calls of the variant. Specifically, the above step S101 includes:
[0096] Step a1, the thread generates n child processes to serve as variants. The n is the redundancy in the multi-variant execution system. The variant is a diversified program in which the child process runs.
[0097] Step a2, the variant will synchronize at each system call and wait for the system to make a judgment on it.
[0098] Specifically, the process creates n threads through create. Since ptrace is only applicable to the parent process to trace and intercept the child process, the process cannot monitor the threads it generates through ptrace. Therefore, the thread generates a child process through fork to start the variant and monitors it through ptrace.
[0099] Step S102, the monitor intercepts and judges the system calls executed by the variant, and records the key information of the system call through the system coordinator. Specifically, the above step S102 includes:
[0100] Step b1, after the variant synchronizes at the system call, it writes the current system call number and system call parameters into the shared buffer. The shared buffer is a memory area in the monitor, allowing the variant child process to access it.
[0101] Step b2, the monitor judges the system call. If the system call numbers, system call parameters, and return results are all the same, the judgment result is legal. Otherwise, the judgment result is illegal.
[0102] Step b3, the monitor records the key information of the system call in the system call coordinator. The key information is the parameters and return values of the current system call. The system call coordinator is a component in the monitor, responsible for recording the key information and classifying the system calls.
[0103] Specifically, let the number of variants be n, and SYN_count is a variable defined in the shared buffer, representing the number of variants that have been synchronized currently. For each thread, it is captured at the system call of the variant through waitpid, and at this time, the value of SYN_count is incremented by 1, that is, the number of synchronized variants is incremented by 1. When SYN_count is equal to the number of variants n, it means that all variants have been synchronized. During the system call process, the system call number and related parameters usually need to be passed through specific registers. Each thread obtains the key information of the system call of the corresponding variant through PTRACE_GETREGS (a command related to the ptrace debugging interface in the Linux system, used to obtain the content of the registers in the variant process), and writes it into the shared buffer to achieve the distribution of data between variants. If the key information of the system calls of each variant in the buffer is consistent, the vote passes; otherwise, the variant is cleaned and restored. After passing the vote, SYN_count is set to 0, and then the current system call is executed and recorded according to the classification of the system call.
[0104] Specifically, the implementation of the system call coordinator (Syscall_Coodinator) described in step b3 is as Figure 3 shown, including a classification module, an execution module, and a shared buffer, where:
[0105] Classification module Classifier, the system call coordinator classifies the system call (syscall) and adopts different strategies for execution according to different system calls.
[0106] Execution module Execution, responsible for executing the system call of the variant according to the corresponding execution strategy in the classification module.
[0107] Shared buffer Shared buffer B, including three parts: input, ret, and param. input records the content read from the external device. When the file descriptor of the read system call is 0, it represents the standard input, and the system call coordinator saves the input content in input in sequence. param is responsible for recording the parameters of the system call during the execution of the variant. For a system call with a legal verdict, the system call coordinator saves the relevant parameters of the system call in it in sequence. When the variant is restored, it compares the parameters of the system call with it to ensure legality. ret records the return value of the system call. The return values of system calls such as getpid, random, and time are random. The system call coordinator records the return values of such system calls in ret, and uses the recorded return values to replace the execution results of the system call when the variant is restored to ensure the consistency of the variant before and after restoration.
[0108] Step S103, when the system call adjudication is abnormal, clean and restore the variant.
[0109] Specifically, during the synchronization of this system call, if the system call number, system call parameters, and system call return value of a variant are different from those of other variants in the shared buffer, the system determines that the variant is abnormal, immediately takes it offline, and reselects a variant to go online.
[0110] Step S104, the re - online variant enters the recovery stage and uses an asynchronous execution mechanism to perform state recovery. The asynchronous execution mechanism allows the cleaned variant to independently recover during the recovery stage and join the synchronous execution after reaching the synchronization point of other variants, realizing the uninterrupted operation of the system.
[0111] Specifically, the flow chart of the asynchronous execution during the variant recovery stage is as Figure 4 shown. Syscall_count is defined in the shared buffer and is used to record the number of system calls that the variant has executed. Syscall_count_replay is defined in its respective thread space and is used to record the number of system calls that the variant has executed during the recovery stage. During the normal execution of the variant, the value of Syscall_count is incremented by 1 for each system call executed. When the variant is cleaned and enters the state recovery, Syscall_count remains unchanged, and the uncleaned variants are in a lock - step state at this time. Syscall_count_replay is set to 0. The value of Syscall_count_replay is incremented by 1 every time a system call is executed until Syscall_count_replay > Syscall_count. At this time, the re - online variant has recovered to the state of the variant before cleaning and switches to synchronous execution.
[0112] Specifically, the asynchronous recovery mechanism can realize the uninterrupted operation of the multi - variant execution system. When the adjudication is abnormal and the abnormal variant is taken offline, other variants execute normally instead of being in a lock - step state. Different from the asynchronous recovery process described above, the Syscall_count of the variant being recovered continues to increase with the execution of system calls of other variants instead of remaining fixed. Since the variant does not need to perform synchronization and adjudication during recovery, the recovery speed is faster than that of synchronous execution. Therefore, the growth rate of Syscall_count_replay is faster than that of Syscall_count. When Syscall_count_replay > Syscall_count, the re - online variant recovers to the state of other execution entities and joins the synchronous execution of other variants. Figure 5Taking the example shown, Variant 3 is cleaned at the second synchronization point. At this time, Variant 1 and Variant 2 continue to execute synchronously. Variant 3 performs asynchronous recovery alone until it catches up with Variant 1 and Variant 2 and then joins the synchronous execution with other variants. The advantage of doing this is that it can achieve uninterrupted operation of the system. When the system detects an unknown attack and performs a cleaning operation, the variants that have not been attacked can continue to execute synchronously to maintain the service.
[0113] In step S105, the variant sends the system calls that need to be executed during the recovery process to the system call coordinator. The system call is the core link of variant state recovery, and the accurate replay of the system call is realized by using the key information recorded in step S102. The replay process of the system call is based on the category of the system call and the recorded key information, and adopts corresponding execution strategies to ensure that the system calls executed during the recovery phase are consistent with the system calls executed before recovery in terms of function and effect, so as to ensure the integrity and consistency of the variant recovery state.
[0114] Specifically, Syscall_Coordinator will classify the system calls and adopt different recovery rules to avoid repeated reading and writing and inconsistencies before and after recovery. According to the category of the system call, Syscall_Coordinator will adopt different strategies for execution. For other system calls that do not affect the consistency between variants, the variant can be executed normally during replay. The classification is as follows:
[0115] (1) Uncertain system calls, such as getpid, random, and stat, may return different results when executed in different variants or at different times. Specifically, getpid is used to obtain the process ID (PID) of the current process, and each call may return a different value, especially in a multi-process environment; random is used to generate pseudo-random numbers, and usually each call will return a different value, and it depends on the seed and algorithm; stat is used to obtain the status information of a file or directory (such as size, permissions, last modification time, etc.), and its results will change when the file system status changes. Therefore, the execution results of these system calls are uncertain. During the variant recovery phase, executing these system calls may obtain different results from those during the previous execution, resulting in inconsistencies before and after recovery, affecting the correctness and stability of the program. Sensitive system calls, such as mkdir, write, and send, will affect the state outside the variant. Specifically, mkdir is used to create a new directory. If the directory already exists at a certain path, executing it again may cause errors or duplicate operations; write is used to write data to a file or device. If such operations are repeated during the recovery phase, it may lead to duplicate data writing or damage to data integrity; send is used to send data to a network socket. If repeated sending occurs during the recovery phase, it may cause duplicate data transmission or incorrect network responses. Therefore, the system calls during the recovery phase have been correctly executed before cleaning. Repeating such system calls may cause problems such as duplicate writing and error reporting, thereby affecting the stability and consistency of the program. For these two types of system calls, Syscall_Coordinator will choose to skip the execution, modify the system call number of the variant to an irrelevant system call, such as getppid, and then use the result recorded by ret to replace the result returned by the system call.
[0116] (2) External input system calls, such as read and recv, are usually used to read data from file descriptors or network sockets. Among them, read is used to read data from files or devices, while recv is usually used to receive data from network sockets. Repeating the previous input operations during the variant recovery stage is not a reasonable approach. For this reason, Syscall_Coordinator copies the input-related content to input during the variant synchronous execution. When an input system call such as standard input is made during variant recovery, Syscall_Coordinator will intercept it, obtain the memory address of the variant where the system call is about to write through PTRACE_GETREGS, and then write the input content buffer saved in input[] to the variant memory through PTRACE_POKEDATA. Finally, the modified system call number is changed to an irrelevant system call, and the length of the buffer is used as the return result of the system call after the execution ends.
[0117] (3) System calls for creating processes and threads, such as clone and fork, will return the process and thread IDs of the created processes and threads after the parent process finishes execution. Different variants will return different values after executing such system calls. Different from indeterministic system calls, after the parent process executes, it often uses the returned process ID to manage the child processes and threads. For example, by calling wait(pid) to wait for a specific child process or thread to end. If only the returned result (i.e., the process ID pid) is rewritten to ensure consistency between variants, it will cause errors in the program when the variant operates on the child process. To solve this problem, Syscall_Coordinator introduces virtual process identifiers to solve this problem. When synchronously executing such system calls in the variant, Syscall_Coordinator will save the returned child process ID pid, and then return a virtual process identifier vid to the variant in order, and establish a mapping between vid and the actual process ID pid. When the variant synchronously executes and adjudicates the system call parameters, the monitor will look up the virtual identifier vid corresponding to each pid, and then adjudicate the consistency of vid. The introduction of virtual process identifiers not only solves the false alarm problem of fork-like system calls, but also ensures that the variant can still maintain consistency with before replay during replay.
[0118] Step S106, the system call coordinator returns the key information and execution policy to the variant. Specifically, the above steps include
[0119] Step c1: The system call coordinator selects the corresponding execution policy according to the category of the system call record.
[0120] Step c2: For an uncertain system call, the system call coordinator sends the recorded return value to the variant.
[0121] Step c3: For a sensitive system call, the system call coordinator chooses to skip repeated execution and sends the corresponding execution policy to the variant.
[0122] Step c4: For an external input system call, the system call coordinator returns the pre-recorded input data to the variant.
[0123] Step c5: For a process or thread creation system call, the system call coordinator replaces the real process identifier with a virtual process identifier and sends it to the variant.
[0124] Specifically, the system call coordinator writes the system call related information into the variant memory through PTRACE_POKEDATA. PTRACE_POKEDATA is an operation type in the ptrace debugging interface, which is used to write data into the memory of the process being debugged. This command is usually used for memory modification during debugging, allowing the parent process (debugger) to write values to specific memory addresses of the child process. For system calls that the variant needs to skip execution, the system call coordinator will modify the variant's system call to an irrelevant system call, such as getppid, and then use the result recorded by ret to replace the result returned by the system call.
[0125] Step S107, the variant completes the replay of the system call and determines whether the recovery is over.
[0126] Step 7.1: The variant determines whether the recovery has been completed by the number of system calls that have completed replay.
[0127] Step 7.2: If the recovery is not completed, the variant continues the recovery. When encountering the next system call, it passes it to the system call coordinator as described in Step S105.
[0128] Step 7.3: If the recovery is completed, the variant ends the asynchronous recovery and joins the synchronous execution of other variants.
[0129] Specifically, if Syscall_count_replay > Syscall_count, the variant that is back online has recovered to the state where the variant before cleaning was executed, and switches to synchronous execution. Otherwise, the variant continues to remain in the recovery phase.
[0130] Embodiment 2
[0131] Based on the same inventive concept, this embodiment discloses a multi-variant recovery device based on record replay, including:
[0132] Monitoring module, which uses a monitor to monitor the system calls of variants. Among them, a child process is generated by a thread as a variant;
[0133] Interception and adjudication module, which is used to intercept and adjudicate the system calls of variants by the monitor according to the key information of the system calls. Among them, the key information of the system calls includes the system call number and the system call parameters. For the system calls adjudged to be legal, the system call coordinator records the key information of the system calls. The system call coordinator is a component of the monitor;
[0134] Recovery module, which is used to, when the system call of a certain variant is adjudged to be illegal, then this variant is an abnormal variant, take the offline processing of the abnormal variant, and when re - online the abnormal variant, use an asynchronous execution mechanism to perform state recovery. Among them, when the variant in the state recovery executes a system call, the system call coordinator determines an execution strategy according to the category of the system call, and then replays the system call of the variant in the state recovery according to the execution strategy and the recorded key information of the legal system calls. Among them, the categories of system calls include indeterminate system calls, sensitive system calls, external input system calls, and system calls for process or thread creation.
[0135] In one implementation, the monitoring module is specifically used for:
[0136] The monitor uses a preset tool to monitor the system calls of variants. Among them, the variant synchronizes at each system call and waits for the system to adjudicate it. The monitor is the core module in the variant parent process and is responsible for the monitoring and synchronization of system calls.
[0137] In one implementation, the interception and adjudication module is specifically used for:
[0138] After the variant synchronizes at the system call, it writes the key information of the system call into the first shared buffer. The first shared buffer is used for the monitor to adjudicate the variant and allows the variant to access it;
[0139] The monitor adjudicates the system call according to the information in the first shared buffer. If the system call numbers, system call parameters, and return results of each variant are all the same, the adjudication result is legal; otherwise, the adjudication result is illegal. Among them, for the system calls adjudged to be legal, the system call coordinator records the key information in the second shared buffer. The second shared buffer is used for the recording of system calls and allows other variants to access it. Among them, when the second shared buffer is full, the system call coordinator batches the content in the second shared buffer and writes it into an external file for persistent storage, and clears the second shared buffer.
[0140] In one implementation, the recovery module is also used for:
[0141] The system call coordinator is solely responsible for the state recovery of the variant and does not interfere with the monitor's interception and adjudication of normal variants;
[0142] When the system call coordinator performs state recovery on an abnormal variant, it continues to record the key information of the system calls of the normal variant.
[0143] In one embodiment, the recovery module is further configured to:
[0144] The system call coordinator batch-reads the key information of the system calls from the external file into the third shared buffer in the recorded order. The third shared buffer is used for the replay of system calls. After all the information in the third shared buffer is used, all the content will be cleared and the key information of the system calls will be read again from the external file;
[0145] The system call coordinator captures the system calls of the variant whose state is being recovered;
[0146] Determine the corresponding execution policy according to the category of the system call;
[0147] The system call coordinator performs deterministic replay of the system call according to the execution policy and the key information of the read system call.
[0148] In one embodiment, the recovery module is further configured to:
[0149] For non-deterministic and sensitive system calls, skip execution during the recovery phase and send the recorded return value to the variant;
[0150] For external input class system calls, return the pre-recorded input data to the variant;
[0151] For process or thread creation class system calls, use a virtual process identifier to replace the real process identifier and send it to the variant.
[0152] In one embodiment, the method further includes a judgment module for judging whether the recovery has been completed according to the number of system calls that have been replayed.
[0153] Since the device introduced in the second embodiment of the present invention is the device adopted for implementing the multi-variant recovery method based on recording and replay in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and deformation of the device, so it will not be elaborated here. Any device adopted in the method of the first embodiment of the present invention falls within the scope of protection of the present invention.
[0154] Embodiment III
[0155] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in Embodiment 1 is implemented.
[0156] Since the computer-readable storage medium introduced in Embodiment 3 of the present invention is the computer-readable storage medium adopted for implementing the method of multi-variant recovery based on recording and playback in Embodiment 1 of the present invention, based on the method described in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the computer-readable storage medium, so it will not be elaborated herein. Any computer-readable storage medium adopted by the method of Embodiment 1 of the present invention falls within the scope of protection of the present invention.
[0157] Embodiment 4
[0158] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in Embodiment 1 is implemented.
[0159] Since the computer device introduced in Embodiment 4 of the present invention is the computer device adopted for implementing the method of multi-variant recovery based on recording and playback in Embodiment 1 of the present invention, based on the method described in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the computer device, so it will not be elaborated herein. Any computer device adopted by the method of Embodiment 1 of the present invention falls within the scope of protection of the present invention.
[0160] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0161] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 a process or multiple processes and / or blocks Figure 1a device for the functions specified in one or more boxes.
[0162] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and variations.
Claims
1. A method for multi-variant recovery based on record replay, characterized in that, Including: Monitoring the system calls of the variant using a monitor, where a subprocess is generated by a thread as the variant; Intercepting and adjudicating the system calls of the variant by the monitor according to the key information of the system calls. For the system calls adjudicated as legal, the key information of the system calls is recorded by the system call coordinator, and the system call coordinator is a component of the monitor; where the key information of the system calls includes the system call number, system call parameters, and return results. When the system call of a certain variant is adjudicated as illegal, then the variant is an abnormal variant, and the abnormal variant is taken offline. When the abnormal variant is restarted, an asynchronous execution mechanism is used for state recovery. Where when the variant in state recovery executes a system call, the system call coordinator determines the execution policy according to the category of the system call, and then replays the system call of the variant in state recovery according to the execution policy and the recorded key information of the legal system calls. Where the categories of system calls include indeterminate system calls, sensitive system calls, external input system calls, and system calls for process or thread creation.
2. The variant recovery method based on recording and replay according to claim 1, characterized in that Monitoring the system calls of the variant using a monitor, including: The monitor uses a preset tool to monitor the system calls of the variant, where the variant synchronizes at each system call and waits for the system to adjudicate it. The monitor is the core module in the variant parent process and is responsible for monitoring and synchronizing system calls.
3. The method for multi-variant recovery based on recording and replay according to claim 1, wherein Intercepting and adjudicating the system calls of the variant by the monitor according to the key information of the system calls, including: After the variant synchronizes at the system call, the key information of the system call is written into the first shared buffer, and the first shared buffer is used for the monitor to adjudicate the variant and allows the variant to access it; The monitor adjudicates the system call according to the information in the first shared buffer. If the system call numbers, system call parameters, and return results of each variant are all the same, the adjudication result is legal, otherwise, the adjudication result is illegal. Where for the system calls adjudicated as legal, the system call coordinator records the key information in the second shared buffer, and the second shared buffer is used for recording system calls and allows other variants to access it. Where when the second shared buffer is full, the system call coordinator batches the content in the second shared buffer and writes it into an external file for persistent storage, and clears the second shared buffer.
4. The method for multi-variant recovery based on recording and playback according to claim 1, characterized in that When restarting the abnormal variant, an asynchronous execution mechanism is used for state recovery, including: The system call coordinator is solely responsible for the state recovery of the variant and does not interfere with the monitor's interception and adjudication of normal variants; When the system call coordinator performs state recovery on the abnormal variant, it continues to record the key information of the system calls of normal variants.
5. The method for multi-variant recovery based on record replay according to claim 1, characterized in that Replaying the system call according to the execution policy and the key information recorded by the system call coordinator, including: The system call coordinator batch reads the key information of system calls from the external file in the recorded order into the third shared buffer, which is used for the replay of system calls. After all the information in the third shared buffer is used, all the content will be cleared and the key information of system calls will be read again from the external file; The system call coordinator captures the system calls of the variant that are in the process of state recovery; Determine the corresponding execution policy according to the category of the system call; The system call coordinator performs deterministic replay of the system call according to the execution policy and the key information of the read system call.
6. The method for variant recovery based on recording and replay according to claim 5, characterized in that, Determine the corresponding execution policy according to the category of the system call, including: For non-deterministic and sensitive system calls, skip execution during the recovery phase and send the recorded return value to the variant; For external input system calls, return the pre-recorded input data to the variant; For system calls related to process or thread creation, use a virtual process identifier to replace the real process identifier and send it to the variant.
7. The method for multi-variant recovery based on recording and replay according to claim 1, characterized in that The method further includes judging whether the recovery has been completed according to the number of system calls that have been replayed.
8. A multi-variant recovery device based on record and replay, characterized in that, Including: A monitoring module for monitoring the system calls of the variant using a monitor; An interception and adjudication module for intercepting and adjudicating the system calls of the variant by the monitor according to the key information of the system call. For the system calls adjudged to be legal, the system call coordinator records the key information of the system call. The system call coordinator is a component of the monitor. Among them, the key information of the system call includes the system call number, system call parameters, and return result; A recovery module for when the system call of a certain variant is adjudged to be illegal, then the variant is an abnormal variant, perform offline processing on the abnormal variant, and when re-online the abnormal variant, use an asynchronous execution mechanism for state recovery. Among them, when the variant in state recovery executes a system call, the system call coordinator determines the execution policy according to the category of the system call, and then performs replay on the system call of the variant in state recovery according to the execution policy and the key information of the recorded legal system call. Among them, the category of the system call includes non-deterministic system calls, sensitive system calls, external input system calls, and system calls related to process or thread creation.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the record replay-based multi-variant recovery method as described in any one of claims 1 to 7.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the record replay-based multi-variant recovery method as described in any one of claims 1 to 7.