A method and system for obtaining file operation logs in a cloud environment
The virtual machine monitor intercepts and analyzes block instructions and uses the write-pre-log semantic dynamic recovery method to process messages, solving the security and compatibility issues of file operation log records in the cloud environment, achieving safe, convenient and complete log acquisition, and improving the security and compatibility of the monitoring system.
Patent Information
- Application Number
- CN202110472277.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-04-29
AI Technical Summary
The file operation logging method in the existing cloud environment has security and compatibility issues, and the deployment is complex, making it difficult to solve the two major issues of security and ease of use at the same time.
Block instructions are intercepted and analyzed through the virtual machine monitor, and the message is processed consistently by using the write-pre-log semantic dynamic recovery method to obtain the file operation log in the virtual machine.
It realizes the secure, convenient and complete acquisition of file operation logs, eliminates the behavior of modifying and deleting logs in the virtual machine, improves the security and compatibility of the monitoring system, and reduces the overhead of deployment and updates.
Smart Images

Figure CN115269300B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer virtualization, and in particular relates to a method and system for obtaining file operation logs in a cloud environment. Background Art
[0002] With the continuous development of cloud computing over the past decade, logging the access and modification of data in the cloud has become of great significance, helping cloud administrators to manage, audit, and protect data. However, the file operation log records in the cloud environment are still imperfect. There are two main types of file operation log collection, specifically:
[0003] One is a common method for modifying the client operating system, including a Linux file operation monitoring method and device (CN106557694B), which loads a preset hijacking function library and dynamic library; a kernel-based file operation behavior monitoring method and device (CN109388538A), which is based on the modification of the kernel; an operation log recording method and device (CN109117420A), which receives file operation requests sent by users; and a low-overhead file operation log collection method (CN111159117A), which uses kernel probes to collect file operation log information in the kernel. The above methods have security and compatibility issues because the monitoring components run inside the virtual machine. After gaining control of the virtual machine, malicious attackers can force the monitoring components to stop or provide false log records. In addition, there are also problems with the deployment of components in the cloud environment. The cloud platform administrator needs to enter each virtual machine for deployment.
[0004] The other is through VMI technology, including a VMI-based file system fine-grained monitoring method (CN102799484A), which simulates the instructions of the virtual machine and obtains memory page and register information to perform fine-grained restoration of the virtual machine. However, this technology requires adaptation of a large number of kernel versions, resulting in complex compatibility issues. At the same time, due to excessive interrupts, it has a significant impact on the performance of the virtual machine.
[0005] In summary, the above-mentioned existing methods cannot solve the problems of security and ease of use at the same time.
[0006] Glossary:
[0007] Cloud computing: refers to breaking down huge data computing programs into countless small programs through the network "cloud", and then processing and analyzing these small programs through a system composed of multiple servers to obtain results and return them to users;
[0008] Virtual Machine (VM): A complete computer system with complete hardware system functions simulated by software and running in a completely isolated environment;
[0009] Hypervisor (VMM): Also known as a virtual machine monitor, a software layer that runs between the host and the client, responsible for coordinating access to hardware resources and applying protection and isolation between virtual machines. Summary of the invention
[0010] In order to solve the above problems, a method and system for intercepting and analyzing block instructions through a virtual machine monitor so as to obtain file operation logs safely, conveniently and completely is provided. The present invention adopts the following technical solutions:
[0011] The present invention provides a method for obtaining a file operation log in a cloud environment, which is used for processing the file access situation of a virtual machine in a cloud computing environment so as to obtain a corresponding file operation log. The method is characterized in that the method comprises the following steps: step S1, intercepting a request from a virtual block device through a virtual machine monitor, and dividing the request into multiple messages of the same length, each message representing a virtual block; step S2, using a pre-write log semantic dynamic recovery method to perform consistency processing on the message, and obtaining a target correspondence relationship between an inode table and a file name by processing the message, and further obtaining a file operation log in the virtual machine according to the target correspondence relationship and the inode table, wherein step S2 comprises the following sub-steps: step S2-1, continuously receiving blocks until the block is a log descriptor block and discarding all blocks before the log descriptor block; step S2-2, when the block is a log descriptor block, parsing the log descriptor block to generate a to-be-processed list, and calculating the disk write position corresponding to each subsequent log content block through the to-be-processed list, and further judging whether the log content block is an inode table block or a directory entry block through the actual write position; step S2-3, When the block is an inode table block, the inode table block is parsed and a corresponding event list is generated according to a predetermined event determination rule, and the event list at least records the inode number corresponding to the inode table block and the corresponding event name; step S2-4, when the block is a directory entry block, the initial correspondence between the inode table block and the file name is parsed from the directory entry block, and the inode number corresponding to the parent directory of each inode table is recorded; step S2-5, when the block is a log submission block, the current time is recorded, and the items in the event list are generated in the order to form a file operation log output, and the file operation log includes the inode number, path name, file name and event name.
[0012] The method for obtaining file operation logs in a cloud environment provided by the present invention may also have the following technical features: wherein, step S2 is processed by an analysis module, the obtained message in step S1 is sent to a message queue in an immediate return manner, and the messages in the message queue are continuously read by the analysis module.
[0013] The method for obtaining file operation logs in a cloud environment provided by the present invention may also have the following technical features: wherein the message includes a virtual block device name, an offset and a message length, and each message represents a virtual block.
[0014] The method for obtaining file operation logs in a cloud environment provided by the present invention may also have the following technical features: wherein, the method for obtaining the path name, file name, and event name is as follows: for each inode, its file name and event name can be directly output, and its path name recursively outputs its parent directory file name until the root directory is reached, at which time the list of all file names output is the path name of the file.
[0015] The method for obtaining file operation logs in a cloud environment provided by the present invention may also have the following technical features: wherein, the blocks received by the analysis module are of three types, namely, log blocks, data blocks, and metadata blocks. The log blocks include descriptor blocks, log content blocks, and log submission blocks. A file operation log starts with a log descriptor block, includes multiple log content blocks, and ends with a submission block, and the log content block is a backup of the metadata block.
[0016] The method for obtaining file operation logs in a cloud environment provided by the present invention may also have the following technical features: wherein the log content block includes an inode table block, a directory entry block, and other log contents.
[0017] The present invention also provides a system for acquiring file operation logs in a cloud environment, which is used for acquiring file operation logs in a cloud environment, and is characterized in that it includes: a message segmentation module, which intercepts requests from virtual block devices through a virtual machine monitor, and segments the requests to obtain multiple messages of the same length, each message representing a virtual block; a file operation log acquisition module, which uses a pre-write log semantic dynamic recovery method to perform consistency processing on the messages and obtain a target correspondence between an inode table and a file name, and further obtains the file operation log in the virtual machine based on the target correspondence and the inode table. The module for obtaining file operation log obtains the file operation log under the virtual environment through the following steps: Step S2-1, continuously receiving blocks until the block is a log descriptor block and discarding all blocks before the log descriptor block; Step S2-2, when the block is a log descriptor block, parsing the log descriptor block to generate a pending list, and calculating the disk write position corresponding to each subsequent log content block through the pending list, and further judging whether the log content block is an inode table block or a directory entry block through the actual write position; Step S2-3, when the block is an inode table block, parsing the inode table block and generating a corresponding event list according to a predetermined event judgment rule, the event list at least records the inode number corresponding to the inode table block and the corresponding event name; Step S2-4, when the block is a directory entry block, parsing the initial correspondence between the inode table block and the file name from the directory entry block, and recording the inode number corresponding to the parent directory of each inode table; Step S2-5, When the block is a log commit block, the current time is recorded, and the items in the event list are output as a file operation log in the order in which they are generated. The file operation log includes the inode number, path name, file name, and event name.
[0018] Function and Effect of the Invention
[0019] According to the method and system for obtaining file operation logs in a cloud environment provided by the present invention, since the block device instructions of the virtual machine are intercepted, the analysis module analyzes the file reading and writing behavior inside the virtual machine, thereby preventing the behavior of modifying and deleting logs in the virtual machine, and also preventing the monitoring system from being affected by malware in the virtual machine. At the same time, the monitoring system is implemented for the file system rather than the operating system, and does not need to be frequently updated. The characteristic of the method of the present invention is that the monitoring is performed on the host machine rather than inside the virtual machine.
[0020] In this way, agentless monitoring can be achieved, which does not rely on the function library and dynamic library inside the virtual machine, nor does it need to modify the kernel. It also does not rely on kernel probes and VMI technology. It can directly intercept block device operations in the driver and reconstruct the file system and its semantics as needed. Compared with existing methods, on the one hand, it has the characteristics of easy deployment, security and transparency, and good compatibility. In the case of a large number of virtual machines in the cloud environment, it can reduce the deployment and update overhead, and on the other hand, it also improves the security of the monitoring software itself. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a system architecture diagram for obtaining file operation logs in a cloud environment in this embodiment.
[0022] Figure 2 This is a flowchart of obtaining file operation logs in a cloud environment in this embodiment
[0023] Figure 3 This is a diagram showing the relationship between various timestamp changes and file operations in this embodiment.
[0024] Figure 4 This is a flowchart of the process of performing consistency processing on messages in this embodiment.
[0025] Figure 5 This is a schematic diagram of the results of the performance loss test in this embodiment. DETAILED DESCRIPTION
[0026] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following is a detailed description of the method and system for obtaining file operation logs in a cloud environment of the present invention in combination with embodiments and drawings.
[0027] <Example>
[0028] The method for obtaining file operation logs in a cloud environment of this embodiment is specifically implemented through a system, which is a host machine based on a Linux platform (equivalent to a cloud environment) and uses kvm-qemu as a virtual machine monitor.
[0029] Figure 1 It is an architecture diagram of a system for acquiring file operation logs in a cloud environment of the present invention.
[0030] like Figure 1As shown, the system 100 includes a virtual machine 101, a virtual machine block device driver 102, a virtual block device 103, a sending end 104, an analysis module 105, a data storage system 106, and a host machine 107. When the virtual machine 101 accesses a file, it sends instructions to the virtual block device 103 through the virtual block device driver 102, and the sending end 104 in the virtual block device driver 102 intercepts these instructions and sends them to the analysis module 105, which parses these instructions into file operations of the virtual machine 101 through a restoration algorithm.
[0031] Figure 2 This is a flowchart of obtaining file operation logs in a cloud environment in an embodiment of the present invention.
[0032] like Figure 2 As shown, obtaining file operation logs in a cloud environment specifically includes the following steps S1 to S2:
[0033] Step S1, modify the virtual block device related code in the virtual block device 103 through the virtual machine monitor, and embed it into the sending end 104. The sending end 104 intercepts and splits each block device instruction, and splits to obtain multiple messages of the same length, each message represents a virtual block, and uses an immediate return method to send the message generated by the sending end 104 to the corresponding message queue, and continuously reads the messages in the message queue through the analysis module 105, and reads the messages sent by the sending end 104 to the message queue to the log system.
[0034] In this embodiment, the sending end 104 is a block device interception plug-in set in the virtual fast device driver 102, which can intercept the virtual block device 103 in qcow2 format and parse the io vector of qemu. Each io vector of qemu is composed of ordinary vectors of different lengths. The sending end 104 of this embodiment cuts it into messages of 4k in length. The message contains the virtual block device name, offset and message length. The sending end 104 uses an immediate return method to send it to the pre-applied linux message queue.
[0035] In addition, since the sending end 104 returns immediately, in order to avoid message queue overflow, the analysis module 105 needs to read messages more frequently than the sending end 104. When the number of messages in the message queue exceeds half of the capacity, the analysis module 105 can adjust the system operation status by increasing the message queue capacity or adding receiving threads.
[0036] Step S2, using the pre-write log semantic dynamic recovery method to process the message for consistency, and obtain the target correspondence between the inode table and the file name by processing the message. Next, according to the target correspondence, the received message is sorted and analyzed, the content of the inode table is taken out from the message, the access time, modification time and deletion time entry of each inode are obtained, and these three are compared with the last write of the inode to obtain the operation content in the virtual machine (i.e., the file operation log).
[0037] Since there is a sequence in which blocks are written, the correspondence between inode and file name may not match, and consistency processing is required. Since most mainstream file systems use block-level logs to ensure the consistency of the file system, this embodiment adopts a solution of parsing the pre-write log, and uses a pre-write log semantic dynamic recovery method to process the message to obtain the target correspondence between the inode table and the file name. This method can be used in a variety of pre-log systems.
[0038] Figure 4 It is a flow chart of the process of performing consistency processing on messages in this embodiment.
[0039] like Figure 4 As shown, the processing process specifically includes the following steps S2-1 to S2-5:
[0040] Step S2-1, the analysis module continuously receives blocks and waits for a log descriptor block to start. Other blocks received at this time are not processed and are directly discarded. This is because these blocks contain metadata blocks and data blocks, and the metadata blocks that are important to the target have been written in the previous log.
[0041] There are three types of blocks that the analysis module may receive, namely log blocks, data blocks, and metadata blocks. Among them, log blocks include descriptor blocks, log content blocks, and log commit blocks. A file operation log starts with a descriptor block, contains multiple log content blocks, and ends with a log commit block. The log content block is a backup of the metadata block.
[0042] In addition, the log content blocks include inode table blocks, directory entry blocks, and other log contents.
[0043] exist Figure 4 In the , the super block is the block of the management block, which is used to store management data. When the super block is received, no processing is done and other blocks continue to be received.
[0044] Step S2-2, when the received block is a log descriptor block, the log descriptor block is parsed to generate a list to be processed. For each subsequent log content block, the actual write position of the log content block on the disk can be calculated, and the actual write position can be used to determine whether the log content block is an inode table block or a directory entry block.
[0045] Step S2-3, when the received block is an inode table block, the inode table block is parsed to obtain the access time, modification time and deletion time entries of each inode, and these three are compared with the last write of the inode, and the inode table block is updated according to the time. Figure 3 The rules in generate a corresponding event list, which at least records the inode number corresponding to the inode table block and the corresponding event name, and stores these events in the buffer.
[0046] Step S2-4, when the log content block is a directory entry block, parse the directory entry, obtain the corresponding relationship between the inode table block and the file name, and save the file name in the inode. At the same time, record the inode number of each inode parent directory, so as to output the file path when the log is submitted.
[0047] Step S2-5, when the log content block is a log submission block, records the current time, and the items in the event list are formed into a file operation log output in the order generated, and the file operation log includes inode number, path name, file name and event name. The corresponding relationship of inode at this moment and file name is consistent with that in the actual file system, and for each inode, its file name and event name can be directly output, and its path name outputs its parent directory file name in a recursive manner, until reaching the root directory, and the list of all file names output at this moment is the path name of the file. Output all access records to the storage system, and clear the buffer, return to S2-1 and continue to run.
[0048] In the above process, if the received block is not any of the log descriptor block, inode table block, directory entry block or log commit block, but other log content, the analysis module will not process it.
[0049] Through the above steps, you can output all file operation logs.
[0050] The method of this embodiment is tested for performance loss. The test results are as follows: Figure 5 As shown, from Figure 5 It can be seen that the performance of the virtual machine file reading and writing system introduced by the present invention decreases by about 7%, and can be applied to cloud environments.
[0051] Example Function and Effect
[0052] According to the method and system for obtaining file operation logs in a cloud environment provided by this embodiment, since the block device instructions of the virtual machine are intercepted, the analysis module analyzes the file reading and writing behavior inside the virtual machine, thereby preventing the behavior of modifying and deleting logs in the virtual machine, and also preventing the monitoring system from being affected by malware in the virtual machine. At the same time, the monitoring system is implemented for the file system rather than the operating system, and does not need to be updated frequently. The characteristic of the method of this embodiment is that the monitoring is performed on the host machine rather than inside the virtual machine.
[0053] In this way, agentless monitoring can be achieved, which does not rely on the function library and dynamic library inside the virtual machine, nor does it need to modify the kernel. It also does not rely on kernel probes and VMI technology. It can directly intercept block device operations in the driver and reconstruct the file system and its semantics as needed. Compared with existing methods, on the one hand, it has the characteristics of easy deployment, security and transparency, and good compatibility. In the case of a large number of virtual machines in the cloud environment, it can reduce the deployment and update overhead, and on the other hand, it also improves the security of the monitoring software itself.
[0054] In this embodiment, due to the immediate return method, the impact on the read and write performance of the virtual machine is reduced. Sending messages of the same length each time helps standardize the sending end and the analysis module, and can achieve the effect of reducing subsequent processing overhead.
[0055] The above embodiments are only used to illustrate specific implementation modes of the present invention, and the present invention is not limited to the description scope of the above embodiments.
[0056] For example, in the above embodiment, the block device interception plug-in intercepts the virtual block device in qcow2 format, thereby cutting the message. For other virtual block device formats such as raw, vdi, vmdk, vhd, etc., the present invention can be directly applied with simple modifications.
[0057] For example, in the above embodiment, the io vector of qemu is cut to generate message contents with a length of 4k. In actual applications, the message length can be adjusted according to actual needs.
[0058] For example, in the above embodiment, the analysis module adopts an extensible design and is currently implemented for the ext4 (Fourth extended filesystem) file system and the jbd2 (journal block device 2) write-ahead log system. In actual applications, the analysis module can be simply modified to extend to other file systems such as XFS, JFS, ReiserFS and Btrfs.
Claims
1. A method for obtaining file operation logs in a cloud environment, for processing file access conditions of a virtual machine in a cloud computing environment to obtain corresponding file operation logs, characterized in that: The following steps are involved: Step S1, intercepting a request from a virtual block device through a virtual machine monitor, and dividing the request into multiple messages of the same length, each message representing a virtual block; Step S2, using a write-ahead log semantic dynamic recovery method to perform consistency processing on the message, and obtaining a target correspondence relationship between an inode table and a file name by processing the message, and further obtaining a file operation log in the virtual machine according to the target correspondence relationship and the inode table, Wherein, the step S2 includes the following sub-steps: Step S2-1, continuously receiving the blocks through the analysis module until the block is a log descriptor block and discarding all the blocks before the log descriptor block; Step S2-2, when the block is a log descriptor block, the log descriptor block is parsed to generate a pending list, and the disk write position corresponding to each subsequent log content block is calculated through the pending list, and further the disk write position is used to determine whether the log content block is an inode table block or a directory entry block; Step S2-3, when the block is the inode table block, the inode table block is parsed and a corresponding event list is generated according to a predetermined event determination rule, the event list at least recording the inode number corresponding to the inode table block and the corresponding event name; Step S2-4, when the block is a directory entry block, the initial correspondence between the inode table block and the file name is parsed from the directory entry block, and the inode number corresponding to the parent directory of each inode table is recorded; Step S2-5, when the block is a log commit block, record the current time, and output the items in the event list in the order of generation to form the file operation log, where the file operation log includes the inode number, path name, file name, and event name. The blocks received by the analysis module are of three types, namely log blocks, data blocks and metadata blocks. The log block includes the log descriptor block, the log content block and the log commit block. A section of the file operation log starts with the log descriptor block, includes multiple log content blocks, and ends with the log commit block. The log content block is a backup of the metadata block.
2. The method for obtaining file operation logs in a cloud environment according to claim 1, characterized in that: in, The step S2 is processed by the analysis module. In the step S1, the obtained message is sent to the message queue in an immediate return manner, and the message in the message queue is continuously read by the analysis module.
3. The method for obtaining file operation logs in a cloud environment according to claim 1, characterized in that: in, The message includes a virtual block device name, an offset and a message length, and each of the messages represents a virtual block.
4. The method for obtaining file operation logs in a cloud environment according to claim 1, characterized in that: in, The method for obtaining the path name, the file name, and the event name is: For each inode, its file name and event name can be directly output, and its path name recursively outputs the parent directory file name until the root directory is reached, at which time the list of all the file names output is the path name of the file.
5. The method for obtaining file operation logs in a cloud environment according to claim 1, characterized in that: in, The log content blocks include inode table blocks, directory entry blocks and other log contents.
6. A system for acquiring file operation logs in a cloud environment, used for acquiring file operation logs in a cloud environment, characterized in that: include: Message segmentation module: intercepts the request from the virtual block device through the virtual machine monitor, and segments the request into multiple messages of the same length, each of which represents a virtual block; File operation log acquisition module: uses a write-ahead log semantic dynamic recovery method to perform consistency processing on the message and obtains a target correspondence between an inode table and a file name, and further obtains a file operation log in the virtual machine based on the target correspondence and the inode table; The file operation log acquisition module acquires the file operation log in the virtual environment through the following steps: Step S2-1, continuously receiving the blocks through the analysis module until the block is a log descriptor block and discarding all the blocks before the log descriptor block; Step S2-2, when the block is a log descriptor block, the log descriptor block is parsed to generate a pending list, and the disk write position corresponding to each subsequent log content block is calculated through the pending list, and further the disk write position is used to determine whether the log content block is an inode table block or a directory entry block; Step S2-3, when the block is the inode table block, the inode table block is parsed and a corresponding event list is generated according to a predetermined event determination rule, the event list at least recording the inode number corresponding to the inode table block and the corresponding event name; Step S2-4, when the block is a directory entry block, the initial correspondence between the inode table block and the file name is parsed from the directory entry block, and the inode number corresponding to the parent directory of each inode table is recorded; Step S2-5, when the block is a log commit block, record the current time, and output the items in the event list in the order of generation to form the file operation log, where the file operation log includes the inode number, path name, file name, and event name. The blocks received by the analysis module are of three types, namely log blocks, data blocks and metadata blocks. The log block includes the log descriptor block, the log content block and the log commit block. A section of the file operation log starts with the log descriptor block, includes multiple log content blocks, and ends with the log commit block. The log content block is a backup of the metadata block.
Citation Information
Patent Citations
Method and device for running multiple operating systems by mobile terminal
CN102799484A
Linux File Operation Monitoring Methods and Devices
CN106557694B
Operation log recording method and device
CN109117420A
A method and apparatus for monitoring file operation behavior based on kernel
CN109388538A
Low-overhead file operation log collection method
CN111159117A