High-performance file processing method and system

Through file path mapping with metadata nodes, file size shard merging and processing, and dynamic load monitoring, the network congestion and resource bottleneck problems of existing file processing systems are solved, efficient and reliable file processing and management are achieved, and the flexibility and maintainability of the system are improved.

CN120295979APending Publication Date: 2025-07-11深圳市领星网络科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510193441.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the reading and writing of high concurrent data, existing file processing systems are prone to significantly increase in latency due to network congestion or node resource bottlenecks, low file processing efficiency, waste of storage space and access efficiency, making it difficult to meet the dual needs of efficiency and reliability.

Method used

By mapping the file path with the metadata node, sharding or merging processing based on the file size, dynamically monitor the storage node load, selecting the appropriate target metadata node for storage, and monitoring the shard health status in real time, and data recovery is carried out using replica redundancy strategies and erasure coding technology.

Benefits of technology

It improves file access speed and storage efficiency, reduces network congestion and node resource bottlenecks, enhances the scalability, reliability and stability of the system, optimizes the utilization of storage resources, and simplifies the system configuration and adjustment process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295979A_ABST
    Figure CN120295979A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of file processing, in particular to a high-performance file processing method and system. The method comprises the following steps: when a target file is received, mapping a path of the target file with a metadata node; performing fragmentation / combination processing on the file based on the size of the file to obtain processing information; storing the processing information on the target metadata node based on the mapping allocation; when the target file is read, identification information and position size information of a target metadata node are returned, and the position size information comprises a network address of the target metadata node and an initial address and a data size of processing information stored in the node; reading processing information based on the identification information and the position size information; the storage types of the processing information are merged / decompressed, the target file is obtained, and the storage types comprise fragmented storage and merged storage. According to the invention, the files in the system can be efficiently processed and managed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of file processing, and in particular, to a high-performance file processing method and system. Background Art

[0002] Currently, the amount of data is growing exponentially, and traditional file processing systems can no longer meet the dual requirements of high efficiency and reliability. Although existing technologies such as Hadoop Distributed File System (High-Availability HDFS), Amazon S3 (Amazon Simple Storage Service), and Ceph (Distributed File System) provide certain solutions for large-scale file processing, they still have the following defects: 1. Performance bottleneck: In high-concurrency data reading and writing, the system is prone to significant delays due to network congestion or node resource bottlenecks.

[0003] 2. Low file processing efficiency: Many systems lack support for file reading, writing, and management, resulting in wasted storage space and decreased access efficiency.

[0004] The above problems severely limit the practical applications of existing file processing systems in the rapid development, and there is an urgent need for an efficient and flexible file processing solution. Summary of the Invention

[0005] The purpose of this application is to overcome the above technical problems and provide a high-performance file processing method and system that can efficiently process and manage files in the system.

[0006] In a first aspect, an embodiment of this application discloses a high-performance file processing method, which adopts the following technical solutions: A high-performance file processing method includes: When receiving a target file, map the path of the target file to a metadata node; Perform sharding / merging processing on the file based on the size of the file to obtain processing information; Allocate and store the processing information on a target metadata node based on the mapping; When reading the target file, return the identification information and location size information of the target metadata node, where the location size information includes the network address of the target metadata node and the start address and data size of the internal storage of the processing information in the node; Read the processing information based on the identification information and the location size information; Perform merging / decompression processing on the storage type of the processing information to obtain the target file, where the storage type includes sharded storage and merged storage.

[0007] By adopting the above technical solutions, files in the system can be efficiently processed and managed. Specifically: By mapping the file path to the metadata node, rapid positioning from the file path to the metadata node can be achieved, improving the file access speed. Through the sharding / merging process based on the file size, large files can be sharded for storage to reduce the burden on a single node; small files can be merged for storage to improve storage efficiency and utilization. By distributing and storing the processing information on the target metadata nodes, the storage pressure can be dispersed, enhancing the scalability and reliability of the system. By returning the identification information and location size information of the target metadata node, the accuracy and efficiency during the file reading process are ensured. Through the merging or decompression process corresponding to the sharding storage and merging storage types, efficient reading and restoration of different types of files can be realized.

[0008] Optionally, the sharding / merging process of the file based on the file size to obtain processing information specifically includes: judging whether the size of the target file exceeds a preset threshold; if so, performing sharding processing on it as a large file to obtain multiple sharding information, where the corresponding target metadata nodes are multiple, and the sharding information is stored in the target metadata nodes in parallel one by one; if not, performing grouped merging processing on it as a small file to obtain data blocks, where the small files are recorded as metadata entries in the data blocks, and the corresponding target metadata node is one; the small files are grouped according to the directory to which the files belong, the file type, and the file size range.

[0009] By adopting the above technical solutions, the file processing efficiency and storage management performance can be effectively improved. Specifically: For large files, when the size of the target file exceeds the preset threshold, it is sharded to generate multiple sharding information, and these sharding information are stored in multiple target metadata nodes in parallel. This method can significantly improve the read / write speed of large files and disperse the storage pressure, avoiding performance degradation caused by overloading a single node. For small files, when the size of the target file does not exceed the preset threshold, they are grouped and merged to form data blocks, and are stored on one target metadata node. At the same time, classification is carried out according to the directory to which the files belong, the file type, and the file size range, making the management and retrieval of small files more efficient and orderly, and reducing the complexity of metadata management. In summary, this solution not only improves the speed and efficiency of file processing, but also optimizes the utilization of storage resources, enhancing the stability and reliability of the system.

[0010] Optionally, the process of merging / unzipping the storage type of the processing information to obtain the target file specifically includes: when the target file is stored in fragments, the multiple fragment information stored in fragments is merged in sequence to obtain the target file; when the target file is stored in a merged manner, the data block is unzipped, the files in each group are extracted, and merged in the original order to obtain the target file.

[0011] By adopting the above technical solution, it is possible to efficiently restore the target files stored in fragments and in a merged manner. For the files stored in fragments, by merging the multiple fragment information in sequence, the integrity and consistency of the files are ensured, and the reading efficiency is improved. For the small files stored in a merged manner, by unzipping the data block and extracting the files in each group, and then merging them in the original order, the rapid and accurate recombination of the small files is achieved, optimizing the storage space utilization rate and access performance.

[0012] Optionally, the process of storing the processing information on the target metadata node based on mapping specifically includes: monitoring the resource usage information of the storage nodes based on mapping, and obtaining the load of each storage node; determining the target metadata node based on the load of each storage node and the processing information, and performing storage.

[0013] By adopting the above technical solution, it is possible to dynamically monitor the resource usage of the storage nodes and select a suitable storage node for storing the processing information according to the current load. This not only improves the flexibility and scalability of the system, but also effectively avoids the problem of overall performance degradation caused by overloading of some nodes. At the same time, determining the target metadata node based on the actual load status of each storage node ensures that the data distribution is more balanced and reasonable, further improving the efficiency and stability of the entire file processing system.

[0014] Optionally, the process of monitoring the resource usage information of the storage nodes based on mapping and obtaining the load of each storage node specifically includes: monitoring the CPU usage rate, memory usage rate, disk I / O, and network bandwidth occupancy information of the storage nodes, and calculating by corresponding matching of preset weight values to obtain the load of each storage node.

[0015] By adopting the above technical solutions, it is possible to monitor the usage of various resources of the storage nodes in real time and accurately, ensuring load balancing. Specifically: Monitoring the CPU usage rate, memory usage rate, disk I / O, and network bandwidth occupancy information of the storage nodes can comprehensively understand the actual working status of each storage node. By calculating corresponding to the preset weight values, the overall load conditions of each storage node can be comprehensively evaluated, thereby providing a scientific basis for selecting appropriate target metadata nodes in the follow-up. This dynamic monitoring and load evaluation mechanism helps to avoid the situation where some nodes are overloaded while others are idle, improving the operating efficiency and stability of the entire system.

[0016] Optionally, determining the target metadata node based on the load of each storage node and the processing information and performing storage includes: When the processing information is the multiple shard information, then based on the number of the multiple shard information, multiple nodes with a load less than a preset load are screened from the storage nodes as the target metadata nodes and stored separately; When the processing information is the data block, then a node with a load less than the preset load is screened from the storage nodes as the target metadata node and stored.

[0017] By adopting the above technical solutions, it is possible to effectively select appropriate storage nodes according to different types of processing information (multiple shard information or data block), ensuring that high-load nodes will not be further burdened, thereby improving the overall performance and stability of the system. For multiple shard information, parallel storage can be achieved to speed up the storage speed; while for data blocks, a single low-load node can be selected for centralized storage to reduce resource waste and improve storage efficiency.

[0018] Optionally, it further includes: Receiving a storage node addition instruction, re-screening multiple nodes with a load less than a preset load based on the load of the newly added node as new storage metadata nodes; When the new storage metadata nodes are different from the target metadata nodes, then the nodes where the target metadata nodes are different from the new storage metadata nodes are screened out as migration nodes; Asynchronously migrating the shard information stored in the migration nodes to the new storage metadata nodes.

[0019] By adopting the above technical solutions, when receiving a storage node addition instruction, it is possible to dynamically adjust the storage resource allocation to ensure the effective utilization of the newly added node. Re-screening new storage metadata nodes that meet the preset conditions based on the load conditions of the newly added node ensures the load balancing and efficient operation of the system. When there are differences between the old and new nodes, asynchronously migrating the shard information on the original target metadata node to the new storage metadata node avoids the performance degradation problem caused by synchronous operations and improves the stability and reliability of the system.

[0020] Optionally, it further includes: receiving a target storage node removal instruction, screening the node with the lowest load from the storage nodes as the new storage metadata node; migrating and storing the shard information / the data block stored on the target storage node to the new storage metadata node.

[0021] By adopting the above technical solution, it is realized that when receiving a target storage node removal instruction, the allocation of storage resources can be dynamically adjusted. First, the new storage metadata node with the lowest load is screened out from the existing storage nodes to ensure that the new node has sufficient resources to bear the migrated data. Then, the shard information or data block on the original target storage node is migrated to the new storage metadata node, ensuring the stability and high availability of the system. This solution can complete the dynamic increase and decrease of storage nodes without affecting normal services, improving the flexibility and maintainability of the system.

[0022] Optionally, when reading the file, it further includes: returning the health status information of the shard information, where the health status information corresponds to whether the shard information is healthy, whether it needs to be rebuilt, and whether there is data corruption; judging whether to adopt the replica redundancy strategy and erasure code technology to recover data based on the health status information, and if so, executing it.

[0023] By adopting the above technical solution, the health status of file shards can be monitored in real time, so that when a fault or damage is detected in the shard information, the replica redundancy strategy and erasure code technology can be used in time to recover data, thereby improving the stability and availability of the system.

[0024] In a second aspect, an embodiment of the present application discloses a high-performance file processing system, adopting the following solution: A high-performance file processing system for executing the method described in any one of the above, includes: A mapping module, when receiving a target file, is used to map the path of the target file to a metadata node; a first processing module, used to perform sharding / merging processing on the file based on the size of the file to obtain processing information; An allocation module, used to allocate and store the processing information on a target metadata node based on the mapping; A return module, when reading the file, is used to return the identification information and location size information of the target metadata node, where the location size information includes the network address of the target metadata node and the start address and data size of the processing information stored inside the node; A reading module, used to read the processing information based on the identification information and the location size information; A second processing module, configured to perform merging / decompression processing on the storage type of the processing information to obtain the target file, where the storage type includes sharded storage and merged storage.

[0025] By adopting the above technical solution, the high-performance file processing system can efficiently manage the storage and access of large-scale files. Specifically: The mapping module realizes the fast mapping between the file path and the metadata node, improving the file indexing speed. The first processing module optimizes the storage efficiency by sharding or merging the file, enabling large files to be stored dispersedly and small files to be centrally managed. The allocation module dynamically selects a suitable storage node based on the monitored load conditions of the storage nodes, avoiding single-point overload and ensuring the stability and reliability of the system. The return module provides detailed identification information and location size information when reading the file, simplifying the file positioning process and accelerating the file reading speed. The reading module accurately reads the required processing information based on the provided detailed information, reducing unnecessary data transmission and enhancing the performance. The second processing module is responsible for restoring the processing information into a complete file, supporting the conversion of multiple storage types, and ensuring the consistency and integrity of the file. Thus, the system not only improves the speed and efficiency of file processing but also enhances the reliability and scalability of the system.

[0026] In summary, the present application includes at least one of the following beneficial technical effects: 1. By mapping the path of the target file to the metadata node and performing sharding or merging processing based on the file size, network congestion and node resource bottlenecks can be effectively reduced, thereby enhancing the response speed and throughput capacity of the system in high-concurrency scenarios.

[0027] 2. For small files, the method of grouped merging processing is adopted. They are grouped according to directories, types, and size ranges and centrally stored in a single metadata node, reducing the management and access overhead of small files, improving the storage space utilization rate and access efficiency.

[0028] 3. By real-time monitoring the load conditions of the storage nodes and dynamically selecting suitable storage nodes according to the load conditions, the system can quickly adapt to the needs of node addition and deletion, simplifying the system configuration and adjustment process, and enhancing the flexibility and maintainability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a schematic flowchart of a high-performance file processing method disclosed in an embodiment of the present application; Figure 2 It is another schematic flowchart of a high-performance file processing method disclosed in an embodiment of the present application; Figure 3Another flowchart of a high-performance file processing method disclosed in an embodiment of the present application; Figure 4 A structural diagram of a high-performance file processing system disclosed in another embodiment of the present application. Detailed implementation manners

[0030] The present application will be further described in detail below with reference to the accompanying drawings.

[0031] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0032] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a" and "the" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0033] It should be understood that although the terms "first", "second", etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.

[0034] The technical solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0035]

First Embodiment

[0036] S20. Shard / merge the file based on the size of the file to obtain processing information; Among them, in this embodiment, large files are sharded and stored on multiple nodes to reduce the burden on a single node; small files are merged and stored on one node to improve storage efficiency and utilization. Specifically, it includes the following steps: S21. Determine whether the size of the target file exceeds a preset threshold; Among them, the preset threshold is, for example, 4M to determine whether the size of the target file exceeds 4M, thereby distinguishing large files and small files.

[0037] S22. If so, perform sharding processing on it as a large file to obtain multiple shard information; Among them, large files are divided into several shard information according to a predetermined size such as 128MB, and the corresponding target metadata nodes are multiple. The shard information is stored in the target metadata nodes in parallel one by one. In this way, the read and write speed of large files can be significantly improved and the storage pressure can be dispersed, avoiding performance degradation caused by overloading of a single node.

[0038] S23. If not, perform grouping and merging processing on it as a small file to obtain data blocks; Among them, in this embodiment, multiple small files can be initially grouped according to features such as the directory to which the file belongs, the file type, and the file size range (such as 0 - 1MB, 1 - 2MB), etc. The target size of each group is controlled between 64MB and 128MB, and then the grouped data is recorded as a metadata entry. In this way, through the metadata entry, the metadata can be managed and retrieved more efficiently and orderly, reducing the complexity of metadata management. It should be noted here that one data block corresponds to a target metadata node for compressed storage to reduce disk occupancy.

[0039] S30. Store the processing information on the target metadata node based on the mapping assignment; Among them, storing the processing information on the target metadata node can disperse the storage pressure and improve the scalability and reliability of the system. In this embodiment, step S30 specifically includes the following steps: S31. Monitor the resource usage information of the storage node based on the mapping and obtain the load of each storage node.

[0040] Among them, the system will monitor the resource usage information of each storage node in real time, such as CPU usage rate, memory usage rate, disk I / O, and network bandwidth occupancy, and then match these resource usage information with preset weight values (such as CPU usage rate accounting for 40%, disk I / O (cache I / O) accounting for 30%, and network bandwidth accounting for 30%) to calculate the load level of each node. Regarding the monitoring, it can be understood that each storage node runs a lightweight resource monitoring agent to report the node status regularly (such as every 10 seconds). The central scheduler aggregates the node status and adjusts the shard distribution strategy in real time to achieve dynamic load balancing. For details, see the following.

[0041] S32. Determine the target metadata node based on the load and processing information of each storage node, and perform storage; Among them, when the processing information is multiple shard information, select multiple nodes with a lower load less than the preset load (for example, 0.8) as the target metadata nodes, and store the shard information on these nodes respectively. When the processing information is a data block, select a node with a lower or the lowest load less than the preset load as the target metadata node, and store the data block on this node.

[0042] In this way, the dynamic allocation of file locations is realized, and appropriate storage nodes are effectively selected according to different types of processing information (multiple shard information or data blocks), ensuring that high-load nodes will not be further burdened, thereby improving the overall performance and stability of the system. For multiple shard information, parallel storage can be realized to speed up the storage speed; while for data blocks, a single low-load node can be selected for centralized storage to reduce resource waste and improve storage efficiency.

[0043] Furthermore, when the client obtains file information, that is, when reading the target file, it includes the following steps: S40. Return the identification information and location size information of the target metadata node; Among them, the identification information is the unique identifier of the storage node, which is the node ID of the returned storage node. For example: node_id: "node_12345", ensuring that the client can accurately find the node storing the processing information. The location size information includes the network address of the target metadata node and the start address and data size of the processing information stored inside the node, as follows: The network address is the IP address and port number of the returned storage node. For example: ip_address: "192.168.0.1", port: 9000.

[0044] The storage location of the processing information is the position offset of the returned file in the node, that is, the start address and size of the processing information stored inside the node. For example: file_offset: 0, data_length: 128MB.

[0045] In this way, the client can communicate with the target metadata node through this information.

[0046] In addition, when there is replica information corresponding to the processing information, the replica information will also be returned in the following steps. If there are multiple replica nodes, the client can select the optimal replica for reading. For example: replica_nodes: [{"node_id": "node_6789", "ip": "192.168.0.2", "port": 9001}].

[0047] In addition, information about the health status of the shard information / data block will also be returned, such as status: "healthy" or status: "damaged".

[0048] Among them, the health status information can specifically correspond to whether the shard information is healthy, whether reconstruction is required, and whether there is data damage, and based on the health status information, it is determined whether to adopt the replica redundancy strategy and erasure code technology to recover the data. If so, the execution is carried out.

[0049] In this way, in this embodiment, the health status of the file shards can be monitored in real time to ensure the integrity and reliability of the data. When it is detected that there is a fault or damage in the shard information, the data recovery process is automatically started.

[0050] Among them, the replica redundancy strategy: use the pre-created spare replicas to complete the missing or damaged data segments. Erasure code technology: Through mathematical operations, reconstruct the lost data from the remaining intact data segments. Here, the detailed execution details of the replica redundancy strategy and erasure code technology are not limited here.

[0051] S50. Read the processing information based on the identification information and location size information; Among them, based on the identification information and location size information of the returned target metadata node, the accuracy and efficiency in the file reading process are ensured.

[0052] S60. Perform merging / decompression processing on the storage type of the processing information to obtain the target file; Among them, the storage types include shard storage and merged storage. When the target file is shard-stored, the multiple shard information of the shard storage is merged in the shard order to obtain the target file; when the target file is merged-stored, the data block is decompressed, the files in each group are extracted, and merged in the original order to obtain the target file.

[0053] For example, the file file_abc is divided into 3 shards (file_abc_part1, file_abc_part2, file_abc_part3). When the client reads, it obtains the data of each shard from different nodes and combines them in sequence into a complete file corresponding to the target file.

[0054] For the merged and stored small files file_1, file_2, and file_3, after the client reads the data blocks and decompresses them, it needs to restore them in sequence to these 3 files, corresponding to 3 target files.

[0055] See Figure 2 , when the client performs the operation of adding a storage node, that is, receiving the storage node addition instruction, it includes the following steps: S70. Based on the load of the newly added node, re-screen multiple nodes with a load lower than the preset load as new storage metadata nodes; Among them, the calculation of the load of the newly added node can refer to the calculation method in step S31 above. This step S70 re-screens new storage metadata nodes that meet the preset conditions based on the load situation of the newly added node, which can dynamically adjust the storage resource allocation, ensure the effective utilization of the newly added node, and ensure the load balance and efficient operation of the system.

[0056] S80. When the new storage metadata node is different from the target metadata node, screen out the nodes where the target metadata node is different from the new storage metadata node as migration nodes; S90. Asynchronously migrate the shard information stored in the migration nodes to the new storage metadata nodes.

[0057] Among them, based on steps S80 and S90, when there are differences between the old and new nodes, the shard information on the original target metadata node is asynchronously migrated to the new storage metadata node, ensuring that the service is not interrupted and avoiding the performance degradation problem caused by synchronous operations, improving the stability of the system and the expansion efficiency of new nodes.

[0058] See Figure 3 , when the client performs the operation of removing a storage node, that is, receiving the target storage node removal instruction, it includes the following steps: S100. Receive the target storage node removal instruction, and screen out the node with the lowest load from the storage nodes as the new storage metadata node; Among them, screening out the new storage metadata node with the lowest load from the existing storage nodes ensures that the new node has sufficient resources to bear the migrated data.

[0059] S110. Migrate and store the shard information / data blocks stored on the target storage node to the new storage metadata node.

[0060] Among them, migrating the shard information or data blocks on the original target storage node to the new storage metadata node ensures the stability and high availability of the system. This solution can complete the dynamic addition and deletion of storage nodes without affecting normal operations, improving the flexibility and maintainability of the system.

[0061] In summary, a high-performance file processing method disclosed in the first embodiment of the present invention can effectively reduce network congestion and node resource bottlenecks by mapping the path of the target file to the metadata node and performing sharding or merging based on the file size, thereby improving the response speed and throughput capacity of the system in high-concurrency scenarios and enhancing the file processing efficiency. By adopting the method of grouped merging processing, grouping them according to directory, type, and size range and centrally storing them in a metadata node, it reduces the management and access overhead of small files, improves the storage space utilization rate and access efficiency. By monitoring the load conditions of the storage nodes in real time and dynamically selecting appropriate storage nodes according to the load conditions, the system can quickly adapt to the needs of node addition and deletion, simplifies the system configuration and adjustment process, and enhances the flexibility and maintainability of the system.

[0062]

Second Embodiment

[0063] Among them, when receiving the target file, the mapping module 210 is used to map the path of the target file to the metadata node; the first processing module 220 is used to perform sharding / merging processing on the file based on the size of the file to obtain processing information; the allocation module 230 is used to allocate and store the processing information on the target metadata node based on the mapping; when reading the file, the return module 240 is used to return the identification information and location size information of the target metadata node, and the location size information includes the network address of the target metadata node and the start address and data size of the processing information stored inside the node; the reading module 250 is used to read the processing information based on the identification information and the location size information; the second processing module 260 is used to perform merging / decompression processing on the storage type of the processing information to obtain the target file, and the storage type includes sharded storage and merged storage.

[0064] It should be noted that a high-performance file processing method implemented by the high-performance file processing system disclosed in the second embodiment of the present application is the same as that in the first embodiment, so it will not be described in detail here.

[0065] Optionally, each module in this embodiment and the above other operations or functions are respectively for implementing the methods in the foregoing embodiments.

[0066]

Third Embodiment

[0067] The technical effect of an electronic device provided in this embodiment in actual application is the same as that of a high-performance file processing method in the first embodiment.

[0068]

Fourth Embodiment

[0069] In addition, it can be understood that the foregoing embodiments are only exemplary descriptions of the present invention. On the premise that the technical features do not conflict, the structure is not contradictory, and the invention purpose of the present invention is not violated, the technical solutions of the various embodiments can be arbitrarily combined and used in combination.

[0070] In several embodiments provided by the present invention, it should be understood that the disclosed methods, systems, and devices can be implemented in other ways. For example, the modules included in the system described above are only illustrative. The division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0071] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0072] In addition, each functional unit / module in various embodiments of the present invention can be integrated in one processing unit / module, or each unit / module can exist physically alone, or two or more units / module can be integrated in one unit / module. The above integrated unit / module can be implemented in the form of hardware, or in the form of hardware plus software functional unit / module.

[0073] The above integrated unit / module implemented in the form of software functional unit / module can be stored in a computer-readable storage medium. The above software functional unit is stored in a storage medium and includes several instructions for causing one or more processors of a computer device (which can be a personal computer, a server, or a network device, etc.) to execute some steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.

Claims

1. A high-performance file processing method, characterized in that, Including: When receiving a target file, map the path of the target file to a metadata node; Perform sharding / merging processing on the file based on the size of the file to obtain processing information; Based on the mapping, allocate and store the processing information on a target metadata node; When reading the target file, return the identification information and location size information of the target metadata node, where the location size information includes the network address of the target metadata node and the start address and data size of the processing information stored inside the node; Read the processing information based on the identification information and the location size information; Perform merging / decompression processing on the storage type of the processing information to obtain the target file, where the storage type includes sharded storage and merged storage.

2. The method according to claim 1, wherein The performing sharding / merging processing on the file based on the size of the file to obtain processing information specifically includes: Judge whether the size of the target file exceeds a preset threshold; If so, perform sharding processing on it as a large file to obtain multiple sharding information, where the corresponding target metadata nodes are multiple, and the sharding information is stored in the target metadata nodes in parallel one by one; If not, perform grouping and merging processing on it as a small file to obtain data blocks, where the small files are recorded as metadata entries in the data blocks, and the corresponding target metadata node is one; the small files are grouped according to the directory to which the file belongs, the file type, and the file size range.

3. The method according to claim 2, characterized in that, The performing merging / decompression processing on the storage type of the processing information to obtain the target file specifically includes: When the target file is stored in shards, perform data merging on the multiple sharding information stored in shards in the sharding order to obtain the target file; When the target file is stored in a merged manner, decompress the data block, extract the files in each group, and merge them in the original order to obtain the target file.

4. The method according to claim 2, wherein The allocating and storing the processing information on a target metadata node based on the mapping specifically includes: Monitor the resource usage information of the storage nodes based on the mapping to obtain the load of each storage node; Based on the load of each storage node and the processing information, determine the target metadata node and perform storage.

5. The method according to claim 4, wherein The monitoring the resource usage information of the storage nodes based on the mapping to obtain the load of each storage node specifically includes: Monitor the CPU usage rate, memory usage rate, disk I / O, and network bandwidth occupancy information of the storage nodes, and perform calculations by correspondingly matching preset weight values to obtain the load of each storage node.

6. The method according to claim 5, characterized in that, The determining the target metadata node and performing storage based on the load of each storage node and the processing information includes: When the processing information is the multiple sharding information, then based on the number of the multiple sharding information, screen multiple nodes with a load less than the preset load from the storage nodes as the target metadata nodes, and perform separate storage; When the processing information is the data block, then screen the nodes with a load less than the preset load from the storage nodes as the target metadata nodes and perform storage.

7. The method according to claim 6, wherein Also including: Receive a storage node addition instruction, and based on the load of the newly added node, re-screen multiple nodes with a load less than a preset load as new storage metadata nodes; When the new storage metadata node is different from the target metadata node, screen out the nodes where the target metadata node is different from the new storage metadata node as migration nodes; Asynchronously migrate the shard information stored in the migration nodes to the new storage metadata nodes.

8. The method according to claim 6, characterized in that, It also includes: Receive a target storage node removal instruction, and screen out the node with the lowest load from the storage nodes as the new storage metadata node; Migrate and store the shard information / data block stored on the target storage node to the new storage metadata node.

9. The method according to claim 2, wherein When reading the file, it also includes: Return the health status information of the shard information, where the health status information corresponds to whether the shard information is healthy, whether it needs to be rebuilt, and whether there is data corruption; Based on the health status information, determine whether to use the replica redundancy strategy and the erasure code technology to recover data. If so, execute.

10. A high-performance file processing system, characterized in that, For executing the method described in any one of claims 1 to 9 above, it includes: A mapping module, which is used to map the path of the target file to the metadata node when receiving the target file; A first processing module, which is used to perform sharding / merging processing on the file based on the size of the file to obtain processing information; An allocation module, which is used to allocate and store the processing information on the target metadata node based on the mapping; A return module, which is used to return the identification information and the location size information of the target metadata node when reading the file. The location size information includes the network address of the target metadata node and the start address and data size of the processing information stored inside the node; A reading module, which is used to read the processing information based on the identification information and the location size information; A second processing module, which is used to perform merging / decompression processing on the storage type of the processing information to obtain the target file, and the storage type includes sharded storage and merged storage.

Citation Information

Cited By

  • Fault detection method and device based on kernel mode and user mode, medium and product

    CN120768748A