Data insertion method and device, electronic equipment, storage medium and program product
By introducing the standard B+tree insertion algorithm and using the file granularity switch in the OCFS2 file system, the write amplification problem of the Extent record insertion algorithm in severely fragmented scenarios is resolved, improving the file system's write performance and storage device lifespan while ensuring system compatibility and flexibility.
Patent Information
- Application Number
- CN202510803465.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-23
AI Technical Summary
In scenarios where file fragmentation is severe, the OCFS2 file system's Extent record insertion algorithm has complex insertion performance, significantly increasing the file system's write amplification, reducing write performance, and accelerating storage device wear.
The standard B+tree insertion algorithm is used to replace the Extent record insertion algorithm in the OCFS2 file system. By introducing a file granularity switch and corresponding tool support in the file system, data format conversion and metadata updates are implemented to optimize insertion performance.
It effectively alleviates the write amplification phenomenon of the file system, improves write performance and extends the service life of storage devices, while maintaining system compatibility and flexibility, making it suitable for dynamically changing workload environments.
Smart Images

Figure CN120687420A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of task scheduling technology, and in particular to a data insertion method, device, electronic device, storage medium, and program product. Background Art
[0002] OCFS2 (Oracle Cluster File System 2) is a cluster file system widely used in large-scale data storage and sharing scenarios. The OCFS2 file system uses Extent B+tree to manage file data storage. Extent is a contiguous data block allocation method that effectively reduces file fragmentation and improves disk sequential read and write performance. Extent B+tree is a B+tree-based data structure used to organize and manage file Extent information. Each node contains information such as the file offset, cluster number, and disk block number. The B+tree structure can quickly locate the location of a file's data blocks. In scenarios with severe file fragmentation, the Extent B+tree insertion operation has some limitations. The traditional Extent B+tree insertion algorithm uses a subtree rotate operation (rotating the entire subtree) when there are no free slots in the node, resulting in a large right shift from the insertion position to the rightmost leaf node, with a time complexity of O(N). Frequent occurrence of this large-scale right-shift operation will significantly increase write amplification (write amplification mainly refers to the phenomenon in solid-state storage where a larger area of internal data needs to be rewritten in order to write a small amount of new data), reduce the write performance of the file system, and increase wear on the storage device.
[0003] Regarding the related technologies, the insertion performance of the OCFS2 file system's Extent record insertion algorithm is relatively complex. In scenarios with severe file fragmentation, this problem will significantly increase the file system's write amplification phenomenon, which has not yet been effectively resolved. Summary of the Invention
[0004] The present application provides a data insertion method, apparatus, electronic device, storage medium, and program product to at least address the problem in the related art that the insertion performance of the OCFS2 file system Extent record insertion algorithm is relatively complex, which significantly increases the file system's write amplification phenomenon in scenarios with severe file fragmentation.
[0005] The present application provides a data insertion method, comprising: determining the activation state of a first insertion algorithm allowed to be used in a target file system; when the activation state of the first insertion algorithm is enabled, converting a first data format of first data to be stored in the target file system into a second data format corresponding to the first insertion algorithm, wherein the target file system is pre-adapted with the first insertion algorithm, and the time complexity of the first insertion algorithm is less than the time complexity of the second insertion algorithm originally in the target file system; and inserting the first data in the second data format into a first data structure of the target file system using the first insertion algorithm.
[0006] The present application also provides a data insertion device, comprising: a determination module, configured to determine the activation state of a first insertion algorithm allowed to be used in a target file system; a conversion module, configured to, when the activation state of the first insertion algorithm is enabled, convert a first data format of first data to be stored in the target file system into a second data format corresponding to the first insertion algorithm, wherein the target file system is pre-adapted with the first insertion algorithm, and a time complexity of the first insertion algorithm is less than a time complexity of the second insertion algorithm originally in the target file system; and an insertion module, configured to insert the first data in the second data format into a first data structure of the target file system through the first insertion algorithm.
[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned data insertion methods when executing the computer program.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned data insertion methods are implemented.
[0009] The present application also provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned data insertion methods when executed by a processor.
[0010] Through this application, the activation state of the first insertion algorithm allowed to be used in the target file system is first determined; when the activation state of the first insertion algorithm is enabled, the first data format of the first data to be stored in the target file system is converted into the second data format corresponding to the first insertion algorithm, wherein the target file system is pre-adapted with the first insertion algorithm, and the time complexity of the first insertion algorithm is less than the time complexity of the original second insertion algorithm in the target file system; and then the first data in the second data format is inserted into the first data structure of the target file system through the first insertion algorithm. Therefore, the problem in the related art that the insertion performance of the Extent record insertion algorithm of the OCFS2 file system is relatively complex and will significantly increase the write amplification phenomenon of the file system in a scenario with severe file fragmentation can be solved, and the write amplification phenomenon of the file system can be effectively alleviated. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 This is a hardware structure block diagram of a computer terminal for a data insertion method according to an embodiment of the present application;
[0013] Figure 2 is a flow chart of a data insertion method according to an embodiment of the present application;
[0014] Figure 3 is another flow chart of a data insertion method according to an embodiment of the present application;
[0015] Figure 4 4 is a structural block diagram of a data insertion device according to an embodiment of the present application. DETAILED DESCRIPTION
[0016] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0017] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0018] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0019] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the data insertion method depends, the specific application environment architecture or specific hardware architecture is described here.
[0020] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of a computer terminal of a data insertion method according to an embodiment of the present application. Figure 1 As shown, the computer terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. The computer terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. For example, the computer terminal may also include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0021] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the data insertion method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0022] The transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a computer terminal's communications provider. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0023] Figure 2 This is a flow chart of a data insertion method according to an embodiment of the present application, which is applied to a metadata server cluster of a distributed storage system. Figure 2 As shown, the process includes the following steps:
[0024] Step S202, determining the activation status of the first insertion algorithm allowed to be used in the target file system;
[0025] Step S204: If the activation state of the first insertion algorithm is enabled, converting the first data format of the first data to be stored in the target file system into a second data format corresponding to the first insertion algorithm, wherein the target file system is pre-adapted with the first insertion algorithm, and the time complexity of the first insertion algorithm is less than the time complexity of the second insertion algorithm originally used in the target file system;
[0026] Optionally, the target file system is an OCFS2 cluster file system. The first insertion algorithm is a standard B+tree insertion algorithm; the second insertion algorithm is an Extent record insertion algorithm.
[0027] The first data format is the data format corresponding to the second insertion algorithm, and the second data format is the data format corresponding to the first insertion algorithm. In other words, the data is converted from the first data corresponding to the Extent record insertion algorithm to the second data format corresponding to the standard B+tree insertion algorithm.
[0028] Step S206: insert the first data in the second data format into the first data structure of the target file system using the first insertion algorithm.
[0029] Through the above steps, first determine the enabled state of the first insertion algorithm allowed in the target file system; when the enabled state of the first insertion algorithm is enabled, convert the first data format of the first data to be stored in the target file system into the second data format corresponding to the first insertion algorithm, wherein the target file system is pre-adapted with the first insertion algorithm, and the time complexity of the first insertion algorithm is less than the time complexity of the original second insertion algorithm in the target file system; and then insert the first data in the second data format into the first data structure of the target file system through the first insertion algorithm. Therefore, it can solve the problem in the related art that the insertion performance of the OCFS2 file system Extent record insertion algorithm is relatively complex, and in the scenario of severe file fragmentation, it will significantly increase the write amplification phenomenon of the file system, and effectively alleviate the write amplification phenomenon of the file system.
[0030] An embodiment of the present application provides a data insertion method, and the method is described in detail in conjunction with the execution process of the data insertion method.
[0031] In an exemplary embodiment, determining the enabled state of a first insertion algorithm allowed to be used in a target file system includes: determining a value written to a first flag bit in a target structure of the target file system, wherein the first flag bit is added to the target structure when the first insertion algorithm is adapted to the target file system, wherein the target structure includes a pointer to the first data structure; when the value is a first value, determining that the enabled state of the first insertion algorithm is enabled; when the value is a second value, determining that the enabled state of the first insertion algorithm is not enabled.
[0032] A first flag, bt_opt_flag, is added to the inode structure (equivalent to the target structure) of the OCFS2 file system to indicate whether the first insertion algorithm (i.e., the standard B+tree insertion algorithm) optimization is enabled for this file. bt_opt_flag = 0 indicates that the existing second insertion algorithm implementation logic is used. bt_opt_flag = 1 indicates that the standard B+tree insertion algorithm optimization is enabled.
[0033] By pre-adding the first flag bit in the target structure, the activation state of the first insertion algorithm can be determined by the specific value written in the first flag bit, thereby making it possible to clearly identify the insertion algorithm currently used by the target file system.
[0034] In an exemplary embodiment, before converting the first data format of the first data to be stored in the target file system into the second data format corresponding to the first insertion algorithm, the method further includes: determining the degree of data fragmentation of the target file system through a first command according to a preset period; enabling the first insertion algorithm when the degree of data fragmentation is greater than a preset threshold; receiving a second command of the target object, and enabling the first insertion algorithm based on the second command.
[0035] It should be noted that the first command is, for example: o2info--filestat <disk file path>; the second command is, for example: ocfs2-btree-ctl-e <file path>#enable the optimization feature.
[0036] That is, the activation schemes of the first insertion algorithm include: Scheme 1: periodically querying the fragmentation level of the target file system through a first command, and automatically activating the first insertion algorithm when the fragmentation level exceeds a preset threshold (e.g., 80%) to insert the first data into the target file system through the first insertion algorithm. Scheme 2: Alternatively, upon receiving a second command from a target party (e.g., an administrator of the target file system), the first insertion algorithm may be activated to insert the first data into the target file system through the first insertion algorithm.
[0037] Through this application, the target file system or the device on which the target file system is deployed can intelligently switch to a more optimal insertion algorithm when data fragmentation reaches a critical point, allowing users to actively intervene, thereby maximizing file system performance and efficiency while ensuring data security. This approach is particularly suitable for dynamically changing workload environments, and can adjust the file system's internal mechanisms in real time to adapt to changing performance requirements without interrupting service.
[0038] In an exemplary embodiment, before determining the enabled state of the first insertion algorithm allowed to be used in the target file system, the method further includes: unmounting the target file system through a third command to stop all read and write operations on the target file system; an extraction step: extracting second data from the target node of the first data structure, wherein the second data includes: stored data in the target node and metadata information corresponding to the stored data, and the target node includes: a root node; a conversion step: converting the second data from the first data format to the second data format, and storing the format-converted second data back to the target file system; an updating step: determining the next node to be processed in the first data structure according to a preset traversal order, and updating the next node to the target node; looping through the extraction step, the conversion step, and the update step until all nodes in the first data structure have been traversed according to the preset traversal order, and the second data corresponding to all nodes have been stored back to the target file system after the format conversion is completed; and remounting the target file system after the data is re-stored, wherein the target file system after the data is re-stored is adapted to the first insertion algorithm.
[0039] Optionally, the third command is, for example, a umount command, wherein a preset traversal order is, for example, depth-first, breadth-first, and the like.
[0040] In this application, the solution for adapting the first insertion algorithm to the target file system includes:
[0041] First, the target file system is unmounted through a third command to stop all read and write operations on the target file system to ensure data integrity and consistency and avoid conflicts or damage caused by other processes accessing the file system during the conversion process.
[0042] Secondly, data extraction and conversion operations are performed. That is, data and metadata are extracted from the data structure currently used by the file system (extent B+tree), converted into a format that can be understood by the new insertion algorithm (standard B+tree), and then the converted data is stored back in the file system. Specifically, it includes: 1) Data extraction: Starting from the root node of the extent B+tree, loop through all nodes to extract the data stored in the node (such as the file offset, block number, extent length, etc.) and metadata information (file size, inode information, etc.). 2) Format conversion: Convert the extracted extent B+tree data and metadata into the data format of the standard B+tree. This may involve reorganizing and storing the data to adapt to the node structure and storage rules of the standard B+tree. 3) Data writeback: Store the converted data back into the file system in the standard B+tree format. This requires modifying the file system metadata and updating the superblock, inode table, b+tree index, etc. to reflect the changes in the data structure.
[0043] Finally, before remounting, perform a data integrity check to ensure that all data and metadata have been correctly converted and are compatible with the new insertion algorithm. Remount the file system using the mount command. This will now use the standard B+tree insertion algorithm instead of the original extent B+tree insertion algorithm. Ensure that the file system operates stably after adopting the new insertion algorithm, and that users can continue to use it without noticing any functional differences.
[0044] Through the above implementation, the target file system can safely convert its internal data structure from extent B+tree to standard B+tree, reducing write amplification and improving file system write performance and storage device lifespan. This approach effectively upgrades and optimizes the file system's internal mechanisms through a process of unmounting, data extraction and conversion, traversal update, and remounting, without interrupting service.
[0045] Alternatively, a Livepatch module can be developed to replace the kernel's OCFS2 file system's Extent B+tree operation logic to implement the first insertion algorithm in the target file system. The Livepatch module can also be used to implement online updates of the target file system, the first insertion algorithm, the second insertion algorithm, and so on.
[0046] In an exemplary embodiment, after inserting the first data in the second data format into the first data structure of the target file system through the first insertion algorithm, the method further includes: extracting third data from the first data structure according to a received fourth command, wherein the fourth command is used to request a query or update of the third data inserted into the first data structure through the first insertion algorithm; converting the third data from the second data format to the first data format; and feeding back the third data in the first data format to the sender of the fourth command.
[0047] After implementing the first insertion algorithm (standard B+tree insertion algorithm) and successfully inserting data (first data) into the first data structure (extent B+tree) of the target file system (OCFS2) in the second data format (data format suitable for standard B+tree), the system must be able to process query or update requests using the first insertion algorithm. These requests are sent using the fourth command, which can be any command that can trigger a file system query or update operation, for example: ocfs2-btree-ctl -q <file path>#Query optimization status.
[0048] Optional, specific steps to process query or update requests:
[0049] Step 1: Receive the fourth command.
[0050] The system monitors and receives a fourth command sent by a user or application. This command carries identification information of the third data to be queried or updated, such as a file name and offset. The command processing module of the OCFS2 file system can be pre-developed or modified to enable the command processing module to recognize and process the fourth command. For example, the user command parser in the kernel can be modified to enable it to parse the new command format and call the corresponding query or update function.
[0051] Step 2: Extract the third data from the first data structure.
[0052] Based on the request in the fourth command, third data is extracted from the first data structure. The third data is the data that the user or application wants to query or update. Specifically, a standard B+tree search algorithm is used to quickly locate the data location based on the key value (e.g., file offset) provided in the fourth command, and then the data and metadata information at that location are extracted.
[0053] Step 3: Data format conversion.
[0054] The third data is converted from the second data format back to the first data format (i.e., the original data format of the OCFS2 file system) so that other system components can understand and process the data. The converted third data (in the first data format) is returned to the sender of the fourth command (e.g., the user or application).
[0055] Through steps 1 through 3 above, we ensure that file system performance is optimized while maintaining compatibility with existing systems. Users do not need to worry about converting the underlying data format and can use the file system as usual. This approach not only improves the internal efficiency of the file system (by reducing write amplification) but also maintains consistency in the user interface, providing users with a seamless upgrade experience. Furthermore, since query and update operations can also be performed online, this further enhances system resiliency and availability.
[0056] In an exemplary embodiment, after inserting the first data in the second data format into the first data structure of the target file system through the first insertion algorithm, the method further includes: when a system upgrade of the target file system is required, checking the enabled state of the first insertion algorithm; if the enabled state is enabled, turning off the first insertion algorithm and changing the enabled state of the first insertion algorithm to disabled.
[0057] That is, the default value of the first flag bit should be a value indicating that the first insertion algorithm is not enabled, for example, bt_opt_flag = 0. Optionally, the first flag bit can be used to enable and disable the first insertion algorithm. By controlling the first flag bit to maintain the default value, the first insertion algorithm can be disabled, ensuring that the original behavior is maintained before and after the system upgrade, thereby reducing the upgrade risk.
[0058] In an exemplary embodiment, after inserting the first data in the second data format into the first data structure of the target file system through the first insertion algorithm, the method further includes: receiving a fifth command initiated by the target object based on a preset tool, wherein the fifth command includes at least one of the following: a command for disabling the first insertion algorithm, a command for querying the optimization status of the fourth data stored in the first data structure; executing the fifth command, and feeding back the execution result of the fifth command to the target object.
[0059] The fifth command is, for example: ocfs2-btree-ctl-d<file path># disable the optimization feature. It is understandable that the system can disable the first insertion algorithm based on the fifth command and switch the target file system to use the original second insertion algorithm. The fifth command is also, for example: ocfs2-btree-ctl-q<file path># query the optimization status. The system can then return the current optimization status of the first data structure of the file system based on the fifth command, including whether the first insertion algorithm is enabled, optimized performance indicators, etc. When the fifth command is executed, the result is returned to the target object through preset tools (such as command line output, logging, API response, etc.).
[0060] By receiving and executing the fifth command, the system not only provides the ability to dynamically manage optimization algorithms but also allows users to monitor optimization results. This is crucial for performance-sensitive production environments or scenarios requiring flexible adjustment of optimization strategies. This design ensures the flexibility and manageability of the file system while improving system transparency and responsiveness through real-time feedback.
[0061] In order to better understand the process of the above-mentioned data insertion method, the implementation process of the above-mentioned data insertion method is described below in combination with an optional embodiment, but it is not used to limit the technical solution of the embodiment of this application.
[0062] The traditional OCFS2 file system's extent record insertion performance is complex and closely related to factors such as the file allocation pattern, extent size, and distribution. When files are written continuously, extent record insertion performs well because it significantly reduces disk fragmentation and improves the disk's sequential read and write performance. However, when files are written randomly or extents are more dispersed, insertion performance may be affected, requiring more disk seeks and metadata updates. Furthermore, the insertion of extent records also needs to consider merging with adjacent records, which may increase computational overhead. When a node has no free slots, the insertion algorithm uses a subtree rotate operation. Because the node is always full, the insertion operation results in a large right shift from the insertion position to the rightmost leaf node, with a time complexity of O(N). In the case of severe file fragmentation, this large right shift operation occurs frequently, greatly exacerbating the write amplification problem.
[0063] To overcome the performance bottleneck of the OCFS2 file system in scenarios with severe file fragmentation, the extent B+tree insertion algorithm needs to be optimized. Replacing the existing extentrecord insertion implementation in the OCFS2 file system with a standard B+tree insertion algorithm effectively resolves write amplification issues, improving file system write performance and storage device lifespan.
[0064] Specifically, the embodiment of the present application aims to provide an OCFS2 file system write amplification optimization method based on a file granularity switch (equivalent to the data insertion algorithm in the above embodiment), which replaces the current extent record insertion implementation of OCFS2 with a standard B+tree insertion algorithm, and introduces a file granularity feature switch and corresponding tool support to effectively solve the write amplification problem while ensuring the compatibility and stability of the system. The standard B+tree is a classic data structure that is widely used in databases and file systems. The standard B+tree dynamically adjusts the tree structure by splitting and merging nodes, and the time complexity of the insertion operation is O(log m N), where m is the order of the B+tree. Since the maximum depth of a B+tree in practical applications typically does not exceed 6, the actual complexity is approximately constant time O(1). Non-leaf nodes in a standard B+tree store only key-value information, while all data is stored in leaf nodes, which are connected via a linked list. This structure allows for more centralized disk I / O, allowing data to be read and written with fewer disk seeks.
[0065] Among them, such as Figure 3 As shown in the figure, the standard B+tree insertion algorithm is used to replace the original extentrecord insertion implementation in the OCFS2 file system, including:
[0066] Step S31: Standard B+tree insertion algorithm replacement (replacement here should be understood as addition or adaptation).
[0067] Algorithm Adaptation and Modification: Adapt the standard B+tree insertion algorithm to the OCFS2 file system architecture and data management. This involves modifying the extent record insertion code logic in the file system kernel module to use the standard B+tree insertion algorithm instead of the existing one. Furthermore, based on the OCFS2 file system's requirements for data consistency and concurrency control, the standard B+tree insertion algorithm was modified and optimized, including the addition of necessary locking mechanisms and transaction support to ensure data integrity in the event of concurrent insertions and system failures.
[0068] Data structure conversion and mapping: Due to differences between the standard B+tree and the original extent record data structures, a data structure conversion and mapping mechanism was designed. During insert operations, extent record data is converted to the standard B+tree node format for insertion. During query and update operations, standard B+tree node data is converted back to the extent record format for use by other file system components. This conversion must be accurate and efficient to avoid performance losses caused by data conversion.
[0069] Metadata Update and Synchronization: After an insert operation, file system metadata, such as file size and data block allocation, is promptly updated. This ensures metadata synchronization between memory and disk to maintain file system consistency. Given the cluster nature of the OCFS2 file system, metadata synchronization and coordination between nodes within the cluster are also crucial to ensure data consistency and availability within the cluster environment.
[0070] Alternatively, an algorithm replacement solution includes:
[0071] 1) Unmount the file system: Before replacing the b+tree, unmount the OCFS2 file system to ensure that no other processes access or modify the file system during the replacement process to avoid data loss or corruption. You can use the umount command to unmount the file system.
[0072] 2) Developing Algorithm Replacement Code: Based on the structural differences between OCFS2 and standard b+trees, we developed code to reorganize OCFS2 b+tree data and store it in a standard b+tree format. This code allows us to read the OCFS2 b+tree structure, extract the data and metadata, and convert it to the standard b+tree format.
[0073] 3) Replace the OCFS2 b+tree. Specific steps: Starting from the root node, traverse the b+tree in OCFS2 layer by layer, and reorganize and store the data and metadata in the node according to the standard b+tree structure.
[0074] 4) Synchronize metadata: When replacing the b+tree structure, first modify the metadata in memory, and then synchronize the metadata to disk regularly to ensure data persistence.
[0075] It's important to note that all metadata updates in OCFS2 are performed on leaf nodes, and the key order cannot be disrupted. A key is a logical offset within a file (cpos). If data is severely fragmented, updating the data will require a large number of right-shift operations.
[0076] The standard b+tree metadata update is as follows: Start from the root node to find the insertion position, traverse down the b+tree layer by layer according to the size of the key value to be inserted, compare the key values of each node until a suitable leaf node is found. After finding the insertion position in the leaf node, if the leaf node is not full (the number of key-value pairs has not reached the preset upper limit), the new key-value pair is directly inserted into this position. If the leaf node is full, the node needs to be split. The key-value pair in the leaf node is divided into two parts, the middle key-value pair is promoted to the parent node, and a new leaf node is created at the same time.
[0077] Step S32: Specify the feature switch of the file granularity.
[0078] 1) Enhanced feature switching at file granularity.
[0079] Flag bit extension: A flag bit, bt_opt_flag, is added to the OCFS2 file system's inode structure (equivalent to the first flag bit in the above embodiment) to indicate whether the standard B+tree insertion algorithm optimization is enabled for the file. A version number field, bt_opt_ver, is also added to identify the version of the feature switch, facilitating subsequent expansion and compatibility management.
[0080] Among them, the addition of flag bits and version number fields is as follows:
[0081]
[0082] bt_opt_flag=0: Use the original extent B+tree implementation logic.
[0083] bt_opt_flag=1: Enable standard B+tree insertion algorithm optimization.
[0084] bt_opt_ver: The initial version number is set to 1 and incremented for subsequent expansions.
[0085] 2) Feature switch management mechanism: New command-line options and corresponding implementation logic have been added to support online management of the B+tree optimization feature. For example, the following operations can be performed:
[0086] It allows batch operation functions, making it convenient for users to manage a large number of files in a unified manner.
[0087] Allow default settings: The feature switch is turned off by default (bt_opt_flag=0), ensuring that the system maintains its original behavior after an upgrade, reducing upgrade risks.
[0088] Online modification support: The enhanced ocfs2-tools tool allows users to open, close, and query the B+tree optimization features of specified files online. The tool commands are as follows:
[0089] ocfs2-btree-ctl -e <file path> # Enable optimization features;
[0090] ocfs2-btree-ctl -d <file path> #Disable optimization features;
[0091] ocfs2-btree-ctl -q <file path>#Query the optimization status.
[0092] Allow batch operation support: support batch operations on directories or file systems, making it easier for users to manage a large number of files in a unified manner and optimize features.
[0093] ocfs2-btree-ctl-E<directory path>#Batch enable the optimization features of all files in the directory;
[0094] ocfs2-btree-ctl -D <directory path> #Batch disable the optimization features of all files in the directory.
[0095] Step S33: Livepatch module development.
[0096] The Livepatch module was developed to replace the extent B+tree operation logic of the OCFS2 file system in the kernel. Version checking and compatibility processing logic were added to the module to ensure its correctness and stability. This module supports online updates without restarting the system or file system services, reducing system maintenance risks and ensuring business continuity. This optimization is suitable for customers with temporary needs.
[0097] Step S34: Enable the feature.
[0098] The present embodiment also provides a compatibility assurance mechanism, including:
[0099] 1) Fallback mechanism: After enabling the optimization feature, if you encounter compatibility issues or exceptions, you can quickly disable the optimization feature through the tool to restore the file system to its original behavior.
[0100] 2) Data consistency check: When the feature switch status changes, a data consistency check is automatically triggered to ensure the integrity of the file system. debugfs.ocfs2 is used to check the file system status, view the internal structure and metadata information of the file system, and help analyze whether the file system has data inconsistencies.
[0101] 3) Version compatibility: The bt_opt_ver field ensures compatibility between optimization features of different versions. When loading Livepatch modules, the version number is checked to ensure that incompatible modules are not loaded.
[0102] The embodiment of the present application adopts a standard B+tree insertion algorithm to avoid large-scale node right shift operations, greatly reduce the write amplification phenomenon, and improve the service life of the storage device and the write performance of the file system. The embodiment of the present application introduces a feature switch at the file granularity, and files that do not have the optimization feature enabled still use the original extent B+tree logic, ensuring compatibility with existing systems and reducing upgrade risks. Through the embodiment of the present application, users can use the ocfs2-tools tool to open, close or query the B+tree optimization feature of a specified file online according to actual needs, providing a flexible management method. In addition, the embodiment of the present application also uses Livepatch technology to achieve online enabling of optimization features, reducing system downtime and improving the availability of the cluster system.
[0103] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0104] This embodiment also provides a data insertion device for implementing the above-mentioned embodiments and preferred implementations. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented using software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0105] Figure 4 is a structural block diagram of a data insertion device according to an embodiment of the present application, such as Figure 4 As shown, the device includes:
[0106] A determination module 42, configured to determine an enabled state of a first insertion algorithm permitted to be employed in a target file system;
[0107] a conversion module 44 configured to, when the activation state of the first insertion algorithm is enabled, convert a first data format of the first data to be stored in the target file system into a second data format corresponding to the first insertion algorithm, wherein the target file system is pre-adapted with the first insertion algorithm, and a time complexity of the first insertion algorithm is less than a time complexity of the second insertion algorithm originally used in the target file system;
[0108] The inserting module 46 is configured to insert the first data in the second data format into the first data structure of the target file system by using the first inserting algorithm.
[0109] Through the above-mentioned device, the activation state of the first insertion algorithm allowed to be used in the target file system is first determined; when the activation state of the first insertion algorithm is enabled, the first data format of the first data to be stored in the target file system is converted into the second data format corresponding to the first insertion algorithm, wherein the target file system is pre-adapted with the first insertion algorithm, and the time complexity of the first insertion algorithm is less than the time complexity of the original second insertion algorithm in the target file system; and then the first data in the second data format is inserted into the first data structure of the target file system through the first insertion algorithm. Therefore, the problem in the related art that the insertion performance of the OCFS2 file system Extent record insertion algorithm is relatively complex and will significantly increase the write amplification phenomenon of the file system in a scenario with severe file fragmentation can be solved, and the write amplification phenomenon of the file system can be effectively alleviated.
[0110] In an exemplary embodiment, a determination module is used to: determine a value written to a first flag bit in a target structure of the target file system, wherein the first flag bit is added to the target structure when the first insertion algorithm is adapted to the target file system, wherein the target structure includes a pointer to the first data structure; when the value is a first value, determine that the enablement state of the first insertion algorithm is enabled; when the value is a second value, determine that the enablement state of the first insertion algorithm is not enabled.
[0111] In an exemplary embodiment, the device also includes an enabling module for: determining the degree of data fragmentation of the target file system through a first command according to a preset period; enabling the first insertion algorithm when the degree of data fragmentation is greater than a preset threshold; receiving a second command of the target object, and enabling the first insertion algorithm based on the second command.
[0112] In an exemplary embodiment, the device further includes an adaptation module, which is used to: unmount the target file system through a third command to stop all read and write operations on the target file system; an extraction step: extracting second data from the target node of the first data structure, wherein the second data includes: stored data in the target node and metadata information corresponding to the stored data, and the target node includes: a root node; a conversion step: converting the second data from the first data format to the second data format, and storing the second data after format conversion back to the target file system; an update step: determining the next node to be processed in the first data structure according to a preset traversal order, and updating the next node to the target node; looping through the extraction step, the conversion step, and the update step until all nodes in the first data structure have been traversed according to the preset traversal order, and the second data corresponding to all nodes have been stored back to the target file system after the format conversion is completed; and remounting the target file system after the data is re-stored, wherein the target file system after the data is re-stored has been adapted to the first insertion algorithm.
[0113] In an exemplary embodiment, the device also includes a first feedback module, which is used to: extract third data from the first data structure according to a received fourth command, wherein the fourth command is used to request a query or update of the third data inserted into the first data structure by the first insertion algorithm; convert the third data from the second data format to the first data format; and feed back the third data in the first data format to the sender of the fourth command.
[0114] In an exemplary embodiment, the device also includes a change module for: when a system upgrade of the target file system is required, checking the enabled state of the first insertion algorithm; when the enabled state is enabled, turning off the first insertion algorithm and changing the enabled state of the first insertion algorithm to disabled.
[0115] In an exemplary embodiment, the device also includes a second feedback module for: receiving a fifth command initiated by the target object based on a preset tool, wherein the fifth command includes at least one of the following: a command for disabling the first insertion algorithm, a command for querying the optimization status of the fourth data stored in the first data structure; executing the fifth command and feeding back the execution result of the fifth command to the target object.
[0116] For the description of the features in the embodiment corresponding to the data insertion device, reference can be made to the relevant description of the embodiment corresponding to the data insertion method, which will not be repeated here.
[0117] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above data insertion method embodiments.
[0118] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned data insertion method embodiments when run.
[0119] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0120] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned data insertion method embodiments are implemented.
[0121] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned data insertion method embodiments are implemented.
[0122] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0123] The above is a detailed introduction to a data insertion method provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A data insertion method, characterized in that: include: Determining an enabled state of a first insertion algorithm allowed to be employed in the target file system; When the activation state of the first insertion algorithm is enabled, converting a first data format of first data to be stored in the target file system into a second data format corresponding to the first insertion algorithm, wherein the target file system is pre-adapted with the first insertion algorithm, and a time complexity of the first insertion algorithm is less than a time complexity of the second insertion algorithm originally in the target file system; The first data in the second data format is inserted into the first data structure of the target file system by using the first insertion algorithm.
2. The data insertion method according to claim 1, characterized in that Determine the enabling status of the first insertion algorithm allowed in the target file system, including: determining a value written to a first flag bit in a target structure of the target file system, wherein the first flag bit is added to the target structure when the first insertion algorithm is adapted to the target file system, wherein the target structure includes a pointer to the first data structure; When the numerical value is the first value, determining that the enabling state of the first insertion algorithm is enabled; When the numerical value is the second value, it is determined that the enabling state of the first insertion algorithm is disabled.
3. The data insertion method according to claim 1, wherein: Before converting the first data format of the first data to be stored in the target file system into a second data format corresponding to the first insertion algorithm, the method further includes: determining the data fragmentation degree of the target file system through a first command according to a preset period; and activating the first insertion algorithm when the data fragmentation degree is greater than a preset threshold; A second command of the target object is received, and the first insertion algorithm is enabled based on the second command.
4. The data insertion method according to claim 1, wherein: Before determining the activation state of the first insertion algorithm allowed to be adopted in the target file system, the method further includes: Unmounting the target file system through a third command to stop all read and write operations on the target file system; Extraction step: extracting second data from a target node of the first data structure, wherein the second data includes: stored data in the target node and metadata information corresponding to the stored data, and the target node includes: a root node; Converting step: converting the second data from the first data format to the second data format, and storing the converted second data back to the target file system; Updating step: determining the next node to be processed in the first data structure according to a preset traversal order, and updating the next node to the target node; The extracting step, the converting step, and the updating step are executed in a loop until all nodes in the first data structure have been traversed according to the preset traversal order, and the second data corresponding to all nodes have been stored back into the target file system after format conversion has been completed; The target file system after the data is restored is remounted, wherein the target file system after the data is restored is adapted to the first insertion algorithm.
5. The data insertion method according to claim 1, wherein: After inserting the first data in the second data format into the first data structure of the target file system using the first insertion algorithm, the method further includes: extracting third data from the first data structure according to a received fourth command, wherein the fourth command is used to request querying or updating the third data inserted into the first data structure by the first insertion algorithm; converting the third data from the second data format to the first data format; The third data in the first data format is fed back to the sender of the fourth command.
6. The data insertion method according to claim 1, characterized in that: After inserting the first data in the second data format into the first data structure of the target file system using the first insertion algorithm, the method further includes: In a case where a system upgrade is required for the target file system, checking the enabled state of the first insertion algorithm; When the enabling state is enabled, the first insertion algorithm is disabled, and the enabling state of the first insertion algorithm is changed to not enabled.
7. The data insertion method according to claim 1, characterized in that: After inserting the first data in the second data format into the first data structure of the target file system using the first insertion algorithm, the method further includes: receiving a fifth command initiated by the target object based on a preset tool, wherein the fifth command includes at least one of the following: a command for disabling the first insertion algorithm, and a command for querying an optimization state of fourth data stored in the first data structure; The fifth command is executed, and the execution result of the fifth command is fed back to the target object.
8. A data insertion device, characterized in that: include: A determination module, configured to determine an enabled state of a first insertion algorithm allowed to be adopted in a target file system; a conversion module, configured to, when the activation state of the first insertion algorithm is enabled, convert a first data format of the first data to be stored in the target file system into a second data format corresponding to the first insertion algorithm, wherein the target file system is pre-adapted with the first insertion algorithm, and a time complexity of the first insertion algorithm is less than a time complexity of the second insertion algorithm originally used in the target file system; An insertion module is configured to insert the first data in the second data format into the first data structure of the target file system by using the first insertion algorithm.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the data insertion method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the data insertion method according to any one of claims 1 to 7 are implemented.