Method and device for improving IO reading performance based on object affinity

By optimizing data layout and access strategies in a distributed storage system and ensuring that GC objects and storage objects are on the same node, the problems of read amplification and low cross-network transmission efficiency are solved, and IO read performance is improved.

CN120687043AActive Publication Date: 2025-09-23CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD

Patent Information

Application Number
CN202511185969.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-09-23
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

The existing ROW+append write architecture has problems with read amplification and low cross-network transmission efficiency in distributed storage systems, resulting in degraded read performance.

Method used

By pre-configuring the affinity between garbage collection (GC) objects and cluster nodes, we optimize data layout and access strategies, ensuring that indexes, storage objects, and corresponding GC objects are located on the same node, and reducing cross-network transmission and disk read operations.

Benefits of technology

It significantly improves the IO read performance of distributed storage systems, reduces the performance loss of data transmission across the network and reading slow media, and improves the read cache hit rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687043A_ABST
    Figure CN120687043A_ABST
Patent Text Reader

Abstract

The invention relates to a method and device for improving IO reading performance based on object affinity. The method comprises the steps that a certain number of GC objects are configured in advance according to the cluster scale, and the affinity relation between the GC objects and all nodes is calculated; when a storage object is applied through aggregation writing or GC transfer writing, determining a GC object associated with the storage object according to the identification information of the storage object, and judging whether the GC object is compatible with a target node or not; if not, discarding the storage object and reselecting, and if yes, writing the data into the storage object; and performing a read service according to the local cache of the storage object in the target node. According to the method, by combining a read-write cache mechanism, the affinity relationship between the cluster node and the GC object is calculated in advance, and the data is written into the storage object compatible with the GC object, so that even if GC moving operation is triggered due to continuous coverage write, the foreground read IO still can obtain effective data from the cache of the original node, and the data storage efficiency is improved. The performance loss of data cross-network transmission and slow medium reading is greatly reduced, and the reading performance is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of hard disk reading and writing technology, and in particular to a method, device, and electronic device for improving IO reading performance based on object affinity. Background Art

[0002] In distributed storage systems, to optimize hard disk write performance, the industry typically uses Redirect-On-Write (ROW) technology: data written to hard disk pool objects is redirected and aggregated into large IOs, and then sequentially written to objects allocated by the hard disk pool. Specifically, when data needs to be written to different storage units in the hard disk pool, the system hashes it to different storage disks (i.e., targets) based on data layout algorithms (such as consistent hashing) to achieve load balancing. In addition, the distributed storage field often uses the Append-Write mechanism to further improve write performance, and ROW technology also uses this feature. However, appending the same data will generate redundant old versions of data (i.e., garbage data), so this invalid space needs to be reclaimed through the Garbage Collection (GC) algorithm.

[0003] While the above technologies effectively improve write performance, they introduce read performance issues that require attention. First, the aggregation nature of foreground data can lead to read amplification. For example, if the data aggregation granularity is 4KB and the user only reads 1KB of data, the system still needs to read the smallest 4KB aggregation unit from the disk, resulting in a read amplification factor of 4. Frequent disk reads of this type can significantly degrade read performance. Second, overwrite writes invalidate old data in old objects, triggering the garbage collection process. The data must be migrated to new objects. Due to the hash consistency mechanism of the redundancy algorithm, the migrated data is distributed across various nodes. Directly reading this data requires accessing it from different nodes across the network, resulting in inefficient transmission.

[0004] In summary, the existing ROW+append write architecture performs well in terms of write performance and space utilization, but has dual bottlenecks in the read path: "read amplification" and "cross-node access", which urgently needs to be optimized. Summary of the Invention

[0005] In order to solve the problems of read amplification and low cross-network transmission efficiency in the existing ROW+append write architecture, this paper proposes a new method to improve IO read performance based on object affinity, aiming to effectively improve the IO read performance of distributed storage systems by optimizing data layout and access strategies.

[0006] The main technical strategies of the present invention include: pre-configuring garbage collection (GC) objects based on the cluster size, and pre-calculating the affinity relationship between each GC object and each node in the cluster; when applying for storage objects and storing data through aggregate write operations or GC garbage collection migration write operations, the system will check the affinity between the GC object associated with the storage object to be applied for and the target node of the storage object; if an inaffinity is detected, the inaffinity storage object will be actively discarded; through the above mechanism, it is ultimately ensured that the index object, storage object and corresponding GC object always maintain affinity, and the three are located on the same node.

[0007] In order to achieve the above objectives, this application provides the following technical solutions: A first aspect of the present application provides a method for improving IO read performance based on object affinity, the method comprising: S1. Pre-configure a certain number of garbage collection (GC) objects based on the size of the distributed storage cluster and pre-calculate the affinity between each GC object and each node in the cluster. S2. When applying for a storage object by aggregate write or GC migration write, determine its associated GC object based on the storage object's identification information (such as the object ID), and determine whether the GC object is affinity with the target node of the storage object; S3. If the result is incompatible, the storage object is discarded and a new one is selected. S4. If the result of the judgment is affinity, the data is written to the storage object, and the storage object and its associated GC object and the corresponding index object are kept on the same node; S5. Perform a read service based on the storage object in the local cache of the node.

[0008] Furthermore, in the method of the present application, the pre-configured number N of GC objects in step S1 satisfies: N is an integer power of 2, and N≥number of nodes×2000.

[0009] Furthermore, in the method of the present application, the pre-calculation of the affinity relationship between each GC object and each node in the cluster in step S1 includes: GC objects are evenly mapped to each node based on consistent hashing, and the set of GC object identifiers that each node is responsible for is saved.

[0010] Furthermore, in the method of the present application, the step S2 of determining whether the GC object is compatible with the target node of the storage object includes: Calculate the GC object identifier obtained by taking the storage object identifier modulo N, and check whether the identifier falls into the GC object identifier set saved by the current node.

[0011] Furthermore, the present application method also includes, during the GC migration and writing process: (1) When the garbage ratio of a storage object reaches a threshold, the GC object associated with the storage object triggers a GC migration operation at the current node; (2) Move valid data (read first and then write) to the newly generated affinity storage object in the same node and update the cache; (3) After the GC migration is completed, the metadata is updated to point to the newly generated affinity storage object, and the original storage object is removed from the cache.

[0012] Furthermore, in the method of the present application, the aggregation granularity of the aggregate write in step S2 is 1 MB, and the small IOs are uniformly written to the disk after aggregation is completed in the node local cache.

[0013] Furthermore, in the method of this application, step S4 also includes: Cache data to the target node: Cache the aggregated written data in units of storage objects into the memory or cache medium of the node where the associated GC object is located; Record storage object metadata: persistently store the metadata of storage objects (including associated GC object information, cache location, etc.); Furthermore, in the method of the present application, step S5 further includes: When a user reads data, the target storage object is located through the index object and the storage object is checked to see if it exists in the cache of the current node. If so, the data is read directly from the local cache.

[0014] A second aspect of the present application provides a device for improving IO read performance based on object affinity, the device comprising: The pre-configuration module is used to pre-configure a certain number of garbage collection (GC) objects based on the scale of the distributed storage cluster and pre-calculate the affinity between each GC object and each node in the cluster; An affinity check module is used to determine the associated GC object based on the identification information (such as the object ID) of the storage object when applying for a storage object through aggregate write or GC migration write, and to determine whether the GC object is compatible with the target node of the storage object; The reselection module is used to actively discard the storage object and reselect when the judgment result is incompatible; The data writing module is used to write data to the storage object when the affinity is determined, and keep the storage object, its associated GC object, and the corresponding index object on the same node; The cache module is used to cache the data of affinity storage objects in the node local cache; The data reading module is used to provide read services based on the local cache of the storage object in the target node.

[0015] The device implements the steps of the aforementioned method for improving IO read performance based on object affinity during operation.

[0016] A third aspect of the present application provides an electronic device, comprising: a memory and a processor; Memory: used to store computer programs; Processor: used for executing the computer program to implement the steps of the aforementioned method for improving IO read performance based on object affinity.

[0017] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the aforementioned method for improving IO read performance based on object affinity are implemented.

[0018] In summary, this application solution pre-calculates the affinity between cluster nodes and garbage collection (GC) objects by combining a read-write cache mechanism, and then writes data to storage objects with affinity for GC objects. With this mechanism, even if GC migration operations are triggered by continuous overwrites, foreground read IO can still obtain valid data from the original node's cache, effectively reducing the performance loss of data transmission across the network and reading slow media, significantly improving read performance.

[0019] Other features and advantages of this application will be described in detail in the following description, or will be understood through the implementation of the relevant technical solutions of this application. The objectives and other advantages of this application can be achieved through the technical features and technical means clearly indicated in the description, claims, and drawings, and obtained through the implementation of these technical contents. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings involved in the description of the embodiments. It should be noted that the drawings only illustrate some embodiments of the present application. Those skilled in the art can deduce other relevant drawings based on these drawings without engaging in creative work.

[0021] Figure 1 This is a flowchart of the overall implementation of the method for improving IO read performance based on object affinity in this application.

[0022] Figure 2 This is a topological diagram of the distributed storage cluster involved in the embodiments of this application.

[0023] Figure 3 This is a schematic diagram of node affinity GC object screening in the embodiment of this application.

[0024] Figure 4 This is a schematic diagram of the data reading and writing process in an embodiment of the present application.

[0025] Figure 5 This is a schematic diagram of the GC moving process in the embodiment of this application.

[0026] Figure 6 This is a structural diagram of the device for improving IO read performance based on object affinity in this application.

[0027] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0029] In this document, the term "including" and any variations thereof (such as "including," "comprising," etc.) are open-ended expressions and should be understood as meaning "including but not limited to," meaning that the listed contents are not exhaustive and may include other contents not explicitly mentioned. The term "based on" should be understood as meaning "based at least in part on," meaning that the basis or condition referred to may not be the only factor and may also involve other relevant factors. The term "one embodiment" should be understood as meaning "at least one embodiment," meaning that the described embodiment is not the only possible implementation method and that other similar embodiments may exist.

[0030] In this application, the terms "a" and "a plurality" are used to modify related elements or features in an illustrative, non-restrictive manner. Unless the context clearly indicates otherwise, "a" should be understood as meaning "at least one," and "a plurality" should be understood as meaning "at least two." Those skilled in the art should interpret these terms appropriately based on the semantics and logical relationships of the context to ensure that they encompass the possibility of "one or more."

[0031] Figure 1 The overall implementation process of the method for improving IO read performance based on object affinity provided by this application is shown, including the following steps: S1. Pre-configure a certain number of garbage collection (GC) objects based on the size of the distributed storage cluster and pre-calculate the affinity between each GC object and each node in the cluster. S2. When applying for a storage object by aggregate write or GC migration write, determine its associated GC object based on the storage object's identification information (such as the object ID), and determine whether the GC object is affinity with the target node of the storage object; S3. If the result is incompatible, the storage object is discarded and a new one is selected. S4. If the result of the judgment is affinity, the data is written to the storage object, and the storage object and its associated GC object and the corresponding index object are kept on the same node; S5. Perform a read service based on the storage object in the local cache of the node.

[0032] In order to more clearly illustrate the technical solution of the present application, the following will further illustrate it through embodiments of specific scenarios.

[0033] like Figure 2 As shown, in the distributed storage cluster scenario involved in this embodiment, it is assumed that the cluster contains three nodes, each node deploying four disks. The user data write process is as follows: first, it is mapped to the index object. The master node of the index object is assigned according to the rules (for example, the master node of index_obj0 is node0, the master node of index_obj1 is node1, and the master node of index_obj2 is node2). Then, through aggregate write, the respective storage objects are requested (assuming three-copy redundancy is used). For example, the storage objects obj0 (associated with tgt0, tgt4, and tgt8), obj1 (associated with tgt1, tgt5, and tgt9), and obj2 (associated with tgt2, tgt5, and tgt9) are generated. Obj0, obj1, and obj2 are respectively associated with the garbage collection (GC) objects gc_obj0, gc_obj1, and gc_obj2 through a hash algorithm.

[0034] In the existing technology, differences in layout algorithms may cause the index object master node and the GC object master node to be incompatible (that is, the index object master node and the GC object master node are not on the same node, for example, index_obj0 is mastered on node0, while gc_obj0 is mastered on node1). When the aggregate writes the storage object obj0, the data will be cached in the memory or cache disk of node0. At this time, reading the data directly from node0 can avoid disk reading and network transmission, significantly improving read performance. However, as the additional overwrite write continues, the storage object obj0 data is gradually overwritten and becomes garbage, triggering the gc_obj0 recycling process. Since the gc_obj0 master node is on node1, the GC move operation will be performed on node1: the valid data of obj0 must first be read and then written to the new storage object obj00 (cached in the memory of node1). At the same time, the metadata will be updated from the old storage object obj0 to point to obj00. At this point, when the foreground IO reads again, it will still try to read the new object obj00 from node0. However, because obj00 was not generated on node0, its cache cannot hit it, resulting in a decrease in read performance. To address this type of read cache hit problem caused by object master node affinities, the present invention proposes a method for improving read performance based on object affinity.

[0035] The specific steps of this method are as follows: 1. Pre-configure GC objects: Pre-configure an appropriate number of GC objects based on the scale of the deployed cluster. For example, for a cluster with fewer than 32 nodes, you can set the number of cluster GC objects to 64K. Based on the hash consistency principle, each node can hash to 64K / 32 = 2K objects. 2. Calculate the affinity between GC objects and nodes: For example, based on the current cluster map information (the map records the cluster nodes and the disk distribution under the nodes, such as node normal / faulty, disk normal / faulty status, etc.), calculate the GC objects that are friendly to the node, that is, select the GC objects that are friendly from 64K GC objects and save them, such as Figure 3 As shown; 3. Data aggregation write: For example, taking a certain aggregation granularity of 1M as an example, if the customer issues an IO size of 4K, 256 4K small IOs need to be aggregated into a large IO of 1M before being sent to the disk. The purpose is to improve the performance of disk writing. When applying for storage object obj during aggregate write, the GC object with obj affinity is hashed out, that is, the ID of the storage object obj is calculated to be remainder 64K, such as Figure 4 As shown; 4. GC object affinity check: Check whether the GC object obtained in step 3 is within the 2K range of this node. If not, it means that the storage object is not compatible with this node. Repeat steps 3-4 until the GC object calculated for the selected storage object is within the 2K range of this node. 5. Write cache: When the storage object obj is written, the data is cached in the corresponding node memory or high-speed media such as SSD in units of obj; 6. Record metadata of storage objects: save metadata of storage objects to disk; 7. Reading data: When a user reads data, the data is first hashed to the index master and read. The object obj is then found through metadata and checked to see if it is in the cache. If obj is in the cache, the read cache hits, and the data can be read directly from the node memory or high-speed media such as SSDs, eliminating the need for cross-network transmission or reading from slow HDDs, thereby improving read performance. 8. GC migration: When the amount of garbage in the storage object obj1 reaches the recycling threshold, the master of the GC object triggers the GC garbage collection mechanism, and moves the valid data of obj1 to the storage object obj2 with affinity to the node by reading first and then writing (the selection method of the storage object obj2 is the same as steps 3-4); at the same time, the obj2 data is written to the cache, such as Figure 5 As shown; 9. Metadata modification and old cache removal: After the GC migration is completed, the metadata of the object obj1 is changed to obj2, and the storage object obj1 is removed from the cache to save cache space; 10. Read data: Same as step 7, except that the metadata is adjusted from obj1 to obj2.

[0036] As data is continuously overwritten, this method can solve the problem of degraded read performance caused by object incompatibility due to continuous GC migration, greatly improving the cache hit rate and read performance.

[0037] Figure 6 FIG. 1 shows a device for improving IO read performance based on object affinity proposed in this application, the device comprising: The pre-configuration module is used to pre-configure a certain number of garbage collection (GC) objects based on the scale of the distributed storage cluster and pre-calculate the affinity between each GC object and each node in the cluster; An affinity check module is used to determine the associated GC object based on the identification information (such as the object ID) of the storage object when applying for a storage object through aggregate write or GC migration write, and to determine whether the GC object is compatible with the target node of the storage object; The reselection module is used to actively discard the storage object and reselect when the judgment result is incompatible; The data writing module is used to write data to the storage object when the affinity is determined, and keep the storage object, its associated GC object, and the corresponding index object on the same node; The cache module is used to cache the data of affinity storage objects in the node local cache; The data reading module is used to provide read services based on the local cache of the storage object in the target node.

[0038] When the above device is running, the steps of the method for improving IO read performance based on object affinity disclosed in this application are implemented.

[0039] The flowcharts and block diagrams in the accompanying drawings illustrate possible implementations of the apparatus, methods, and computer program products according to various embodiments of the present application, including architecture, functions, and operations. In these figures, each box may represent a module, a program segment, or a portion of a code, which contains one or more executable instructions for implementing a specified logical function. It should be noted that each box in the block diagram and / or flowchart, and the combination of these boxes, can be implemented using a dedicated hardware-based system to implement the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0040] like Figure 7 As shown, an embodiment of the present application further discloses an electronic device, comprising: a processor 310, a communication interface 320, a memory 330 for storing a computer program executable by the processor, and a communication bus 340. The processor 310, the communication interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 executes the executable computer program to implement the steps of the above-mentioned method for improving I / O read performance based on object affinity.

[0041] It is understood that, in addition to the memory and processor, the electronic device may also include an input device (e.g., a keyboard), an output device (e.g., a display), and other communication modules. These input devices, output devices, and other communication modules all communicate with the processor via an I / O interface (i.e., an input / output interface).

[0042] The operation of the present application can be implemented by writing computer program code using one or more programming languages ​​or a combination thereof. The programming languages ​​include but are not limited to the following types: Object-oriented programming languages, such as Java, Smalltalk, C++, etc.; A conventional procedural programming language, such as "C" or a similar programming language.

[0043] The execution methods of the program code include but are not limited to: Executes entirely on the user's computer; Partially executed on the user's computer and partially on a remote computer; Executed as a standalone software package; Executes entirely on the remote computer or server.

[0044] In scenarios involving a remote computer, the remote computer can be connected to the user's computer via any type of network, including but not limited to a local area network (LAN) or a wide area network (WAN). Additionally, the remote computer can be connected to an external computer via an Internet service provider, such as the Internet.

[0045] Furthermore, the present application also discloses a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the various steps of the method for improving IO read performance based on object affinity disclosed in the present application.

[0046] In the context of this application, computer-readable storage media refers to tangible media that can store computer program code and related data. Specific examples include, but are not limited to, the following: (1) Portable computer disk: A removable magnetic storage medium such as a floppy disk.

[0047] (2) Hard disk: includes fixed storage devices such as mechanical hard disks and solid-state hard disks.

[0048] (3) Random Access Memory (RAM): Volatile storage medium used for temporary storage of data and program code.

[0049] (4) Read-only memory (ROM): A non-volatile storage medium used to store fixed programs and data.

[0050] (5) Erasable Programmable Read-Only Memory (EPROM) or Flash Memory: A non-volatile storage medium that supports multiple erasing and programming.

[0051] (6) Fiber optic storage device: storage medium based on fiber optic technology.

[0052] (7) Compact Disc Read-Only Memory (CD-ROM): A read-only medium that stores data in the form of an optical disc.

[0053] (8) Optical storage devices: storage media based on optical principles, such as DVDs and Blu-ray discs.

[0054] (9) Magnetic storage devices: storage media based on magnetic principles, such as magnetic tapes and disks.

[0055] (10) Any suitable combination of the above: for example, combining multiple storage media to meet different storage requirements.

[0056] These computer-readable storage media can be used to store the program code and related data described in this application to support the operation of the program and the persistent storage of data.

[0057] In particular, according to an embodiment of the present application, the process described in the flowchart can be implemented as a computer software program. For example, an embodiment of the present application relates to a computer program product, which includes a computer program carried on a non-transitory computer-readable medium. The computer program includes program code for executing the method for improving IO read performance based on object affinity disclosed in the present application. When the computer program is executed by a processing device, it can implement the above-mentioned functions defined in the embodiments of the present application.

[0058] Although the above discussion contains several specific implementation details, these details should not be interpreted as limiting the scope of this application. The above description is only a preferred embodiment of the present application and an illustration of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features. At the same time, this application should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concepts.

[0059] Those skilled in the art should also understand that they may modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein with equivalents, without departing from the spirit and scope of the technical solutions of the embodiments of the present application. Such modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the core spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for improving IO read performance based on object affinity, characterized in that: The method comprises: S1. Pre-configure a certain number of garbage collection (GC) objects based on the size of the distributed storage cluster and pre-calculate the affinity between each GC object and each node in the cluster. S2. When applying for a storage object by aggregate write or GC migration write, the GC object associated with the storage object is determined based on its identification information, and determines whether the GC object and the target node of the storage object are affinity; S3. If the result is incompatible, the storage object is discarded and a new one is selected. S4. If the result of the judgment is affinity, the data is written to the storage object, and the storage object and its associated GC object and the corresponding index object are kept on the same node; S5. Perform a read service based on the storage object in the local cache of the node.

2. The method according to claim 1, characterized in that The preconfigured number N of GC objects in step S1 satisfies: N is an integer power of 2, and N≥number of nodes×2000.

3. The method according to claim 1, characterized in that The pre-calculation of the affinity relationship between each GC object and each node in the cluster in step S1 includes: GC objects are evenly mapped to each node based on consistent hashing, and the set of GC object identifiers that each node is responsible for is saved.

4. The method according to claim 3, characterized in that The step S2 of determining whether the GC object is compatible with the target node of the storage object includes: Calculate the GC object identifier obtained by taking the storage object identifier modulo N, and check whether the identifier falls into the GC object identifier set saved by the current node.

5. The method according to claim 1, wherein The method also includes, during the GC migration and writing process: (1) When the garbage ratio of a storage object reaches a threshold, the GC object associated with the storage object triggers a GC migration operation at the current node; (2) Move valid data to the newly generated affinity storage object in the same node and update the cache; (3) After the GC migration is completed, the metadata is updated to point to the newly generated affinity storage object, and the original storage object is removed from the cache.

6. The method according to claim 1, wherein The aggregation granularity of the aggregate write in step S2 is 1 MB, and small IOs are aggregated in the node local cache and then written to the disk.

7. The method according to claim 1, characterized in that Step S4 also includes: Cache data to the target node: Cache the aggregated written data in units of storage objects into the memory or cache medium of the node where the associated GC object is located; Record storage object metadata: Persistently store the metadata of storage objects.

8. The method according to claim 1, characterized in that Step S5 also includes: When a user reads data, the target storage object is located through the index object and the storage object is checked to see if it exists in the cache of the current node. If so, the data is read directly from the local cache.

9. A device for improving IO read performance based on object affinity, characterized in that: The device comprises: The pre-configuration module is used to pre-configure a certain number of garbage collection (GC) objects based on the scale of the distributed storage cluster and pre-calculate the affinity between each GC object and each node in the cluster; An affinity check module is used to determine the associated GC object based on the identification information of the storage object when applying for a storage object through aggregate write or GC migration write, and to determine whether the GC object is compatible with the target node of the storage object; The reselection module is used to actively discard the storage object and reselect when the judgment result is incompatible; The data writing module is used to write data to the storage object when the affinity is determined, and keep the storage object, its associated GC object, and the corresponding index object on the same node; The cache module is used to cache the data of affinity storage objects in the node local cache; The data reading module is used to provide read services based on the local cache of the storage object in the target node.

10. An electronic device, characterized in that: include: memory and processor; Memory: used to store computer programs; Processor: configured to execute the computer program to implement the steps of the method for improving IO read performance based on object affinity as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data storage method and device

    CN104298681A

  • CRDT junk data recovery method and device, equipment and storage medium

    CN112559383A

  • Data processing method, system and equipment and storage medium

    CN116069261A

  • Garbage collection method and equipment for partition storage equipment and storage medium

    CN118733482A

  • Hard disk data writing method and device, electronic equipment and storage medium

    CN120315633A

Cited By

  • Hard disk garbage collection method and device

    CN121166563A