Internal object processing method and device, storage medium, electronic equipment and program product
By pre-storing metadata in memory and determining the recycling strategy in response to business object changes, the problem of low internal object recycling efficiency is solved, and more efficient storage resource management and system performance improvement is achieved.
Patent Information
- Application Number
- CN202412000543.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-31
Smart Images

Figure CN120045129A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computers, and more particularly, to a method, apparatus, storage medium, electronic device, and program product for processing internal objects. Background Art
[0002] Currently, distributed storage systems play a core role in the fields of big data and cloud computing, and their high performance, high reliability, and easy scalability are widely favored.
[0003] In the related art, metadata management has become the key to optimizing system performance and resource utilization, especially when dealing with small input / output (I / O) scenarios. However, for the update of reverse metadata, usually the metadata of the internal object is first read and then modified. Therefore, the above method has the technical problem of low efficiency in recycling internal objects. Summary of the Invention
[0004] The embodiments of the present application provide a method, apparatus, storage medium, electronic device, and program product for processing internal objects, so as to at least solve the problem of low efficiency in recycling internal objects in the related art.
[0005] According to an embodiment of the present application, a method for processing internal objects is provided. The method may include: in response to an initial service object changing to a target service object, determining, from at least one first metadata stored in memory, an initial metadata corresponding to the initial service object, where the first metadata is used to at least characterize the storage state of data in historical internal objects associated with the initial service object, and the historical internal objects are used to store the service data of the initial service object; changing the initial metadata to obtain target metadata; and based on the target metadata, determining a recycling policy for the historical internal objects, where the recycling policy is a rule for indicating whether to recycle the historical internal objects.
[0006] In an exemplary embodiment, changing the initial metadata to obtain target metadata includes: determining the data volume corresponding to the service data; and changing the valid data length information in the initial metadata according to the data volume, where the valid data length information is used to characterize the data storage volume in the historical internal objects.
[0007] In an exemplary embodiment, based on the target metadata, determining a recycling policy for the historical internal objects includes: in response to the valid data length information being less than a first target value, determining that the recycling policy is to recycle the historical internal objects.
[0008] In an exemplary embodiment, based on the target metadata, a recycling policy for historical internal objects is determined, including: in response to the valid data length information being greater than or equal to the first target value, determining that the recycling policy is to not perform recycling processing on the historical internal objects.
[0009] In an exemplary embodiment, in response to the initial business object changing to a target business object, the initial metadata corresponding to the initial business object is determined from at least one first metadata stored in the memory, including: in response to the initial business object changing to a target business object, determining the reverse relationship corresponding to the initial business object, where the reverse relationship is used to represent the mapping relationship between the historical internal object and the initial business object; based on the reverse relationship, determining the historical internal object mapped by the initial business object; determining the identity information of the historical internal object; and determining the first metadata including the identity information among the at least one first metadata as the initial metadata.
[0010] In an exemplary embodiment, determining the first metadata including the identity information among the at least one first metadata as the initial metadata includes: searching the database according to the identity information to obtain the initial metadata, where the database is used to store the first metadata.
[0011] In an exemplary embodiment, the method may further include: when the memory restarts, in response to the need to recycle historical internal objects or the data storage situation of the historical internal objects changing, loading the initial metadata into the memory.
[0012] In an exemplary embodiment, the method may further include: in response to a data storage instruction issued by the client, obtaining a plurality of data to be stored corresponding to the data storage instruction; performing an aggregation process on the plurality of data to be stored to obtain a first internal object; constructing second metadata of the first internal object based on the data volume of the data to be stored, where the second metadata is at least used to represent the data storage situation of the first internal object; constructing a mapping relationship between the data to be stored and the first internal object to obtain a first forward relationship; and storing the first forward relationship and the second metadata in the memory.
[0013] In an exemplary embodiment, the method may further include: obtaining a data acquisition request issued by the client, where the data acquisition request is used to acquire data to be acquired stored in the data pool; determining a first business object corresponding to the data acquisition request; obtaining the first forward relationship corresponding to the first business object from the memory; based on the first forward relationship, determining the identity information of the first internal object mapped by the target business object; and acquiring the data to be acquired from the first internal objects deployed in the data pool according to the identity information.
[0014] In an exemplary embodiment, the method may further include: transmitting data to be acquired to a client, and in response to acquiring the data to be acquired from a first internal object, modifying the valid data length information of the third metadata corresponding to the first internal object, wherein the third metadata is stored in memory.
[0015] In an exemplary embodiment, the method may further include: in response to a change in service data, performing an aggregation process on the changed service data to obtain a second internal object, where the second internal object is different from a historical internal object; based on the data volume of the changed service data, constructing fourth metadata of the second internal object, where the fourth metadata is at least used to characterize the data storage condition of the second internal object; constructing a mapping relationship between the changed service data and the second internal object to obtain a second forward relationship; and storing the second forward relationship and the fourth metadata in memory.
[0016] According to another embodiment of the present application, there is also provided a processing device for internal objects, the device may include: a first determination unit, configured to, in response to an initial service object changing to a target service object, determine initial metadata corresponding to the initial service object from at least one first metadata stored in memory, where the first metadata is at least used to characterize the data storage condition of a historical internal object associated with the initial service object, and the historical internal object is used to store service data in the initial service object; a change unit, configured to change the initial metadata to obtain target metadata; and a second determination unit, configured to determine a recycling policy for the historical internal object based on the target metadata, where the recycling policy is a rule for indicating whether to recycle the historical internal object.
[0017] According to still another embodiment of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0018] According to still another embodiment of the present application, there is also provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0019] According to still another embodiment of the present application, there is also provided a computer program product, including a computer program, where the computer program implements the steps in any one of the above method embodiments when executed by a processor.
[0020] Through this application, in response to an initial service object being changed into a target service object, an initial metadata corresponding to the initial service object is determined from at least one piece of first metadata stored in a memory, where the first metadata is used to at least characterize a storage state of data in historical internal objects associated with the initial service object, and the historical internal objects are used to store service data of the initial service object; the initial metadata is changed to obtain target metadata; based on the target metadata, a recycling policy for the historical internal objects is determined, where the recycling policy is a rule for indicating whether to perform a recycling process on the historical internal objects. That is to say, in the embodiments of this application, the first metadata is pre-stored in the memory, and the first metadata can be used to determine the data storage situation of the historical internal objects. Therefore, when the service data of the initial service object changes and the initial service object is changed into a target service object, the initial metadata corresponding to the initial service object can be further determined, and the initial metadata is adjusted to obtain target cloud data. Further, based on the target metadata, it can be determined whether to perform a recycling process on the historical internal objects, thereby solving the technical problem of low efficiency in recycling internal objects and achieving the technical effect of improving the efficiency of recycling internal objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a hardware structure block diagram of a server device for a method of processing internal objects according to an embodiment of this application;
[0022] Figure 2 is a flowchart of a method of processing internal objects according to an embodiment of this application;
[0023] Figure 3 is a schematic diagram of the correspondence between service objects and internal objects according to an embodiment of this application;
[0024] Figure 4 is a flowchart of a method of recycling internal objects according to an embodiment of this application;
[0025] Figure 5 is a flowchart of a method of flushing internal objects according to an embodiment of this application;
[0026] Figure 6 is an example diagram of a write operation in a distributed storage system according to an embodiment of this application;
[0027] Figure 7 is a structure block diagram of a device for processing internal objects according to an embodiment of this application;
[0028] Figure 8 is a computer system structure block diagram of an electronic device according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0031] In this embodiment, a method for processing an internal object is also provided. The system is used to implement the embodiment and the preferred implementation mode, and the description has been made no further. As used below, the terms "module" and "unit" are a combination of software and / or hardware that can implement a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0032] As an optional implementation, the method embodiment provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 1 is a hardware structure block diagram of a server device of an internal object processing method according to an embodiment of the present application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the server device may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0033] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as computer programs corresponding to the processing methods of internal objects in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, to implement the above-mentioned methods. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the server device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0034] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of a server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0035] In this embodiment, a method for processing internal objects is provided. Figure 2 It is a flowchart of the method for processing internal objects according to an embodiment of the present application, as Figure 2 shown, and the process includes the following steps:
[0036] Step S202, in response to the initial service object changing to a target service object, determine the initial metadata corresponding to the initial service object from at least one piece of first metadata stored in the memory, where the first metadata is used to at least characterize the storage state of data in historical internal objects associated with the initial service object, and the historical internal objects are used to store the service data of the initial service object;
[0037] Step S204, modify the initial metadata to obtain target metadata;
[0038] Step S206, based on the target metadata, determine the recycling policy of the historical internal objects, where the recycling policy is used to represent the rule for indicating whether to perform recycling processing on the historical internal objects.
[0039] In this embodiment, the above initial service object can be data written by a client, can be a data entity created, read, written, and managed by a user or an application program in a distributed storage system, can be a data unit, and can contain any type of service data, such as files, images, or database records, etc. The above service data can be service object data, can be documents, pictures, or video files uploaded by users. It should be noted that only examples are given here, and the type of the above service data is not specifically limited.
[0040] Optionally, the above first metadata can be newly added metadata added in this embodiment, a newly added memory data structure, can be metadata information of an internal object, and the first metadata can be used to characterize the state of the corresponding internal object. The above internal object can be a data unit stored in a data pool, can be used to store service data, and the above internal object can be a data structure or storage unit introduced in a distributed storage system to optimize data storage and operations.
[0041] Optionally, when processing write operations on business objects, the system aggregates the written data and converts multiple small, random data write operations into sequential write operations on internal objects. Internal objects usually have a fixed size, such as 4 megabytes (MB), which can be used to store aggregated data segments.
[0042] Optionally, the internal object can be a physical storage carrier of the business object data, and there is a forward and reverse metadata relationship between it and the business object. The forward relationship is from the business object to the internal object for data retrieval; the reverse relationship is from the internal object to the business object for garbage collection and storage space management. When the business object data changes, the system will create a new internal object to store the updated data, and update the forward and reverse metadata at the same time to maintain the integrity of the data and the efficient operation of the system.
[0043] Optionally, during the garbage collection process, the valid data length information of the internal objects can be used to determine which internal objects are no longer referenced by any business object. These internal objects are regarded as garbage data, and the storage space they occupy can be reclaimed for reuse.
[0044] In this embodiment, in order to manage internal objects, a distributed storage reverse metadata management method is designed according to the characteristics of distributed storage. This method avoids the need to read old internal object metadata every time a business object is modified by adding a set of metadata. The metadata can be first metadata, and the constructed first metadata can be stored in the memory to manage the internal object using the first metadata.
[0045] Optionally, each storage unit (PG) stores its own internal object information, that is, the first metadata corresponding to the internal object. Each internal object can use a 64-bit unsigned integer and a 32-bit unsigned integer to record information, wherein the above 64-bit unsigned integer and a 32-bit unsigned integer record information can constitute the first metadata. Among them, the 64-bit unsigned integer records the internal object identity (Identification, referred to as ID), and the 32-bit unsigned integer records the valid data length of the internal object. This information resides in the memory and is stored on the disk. Each time a business object is written, this information is written to the disk together with the metadata of the business object. In this way, there is no need to read the metadata of the original internal object. When the internal object is aged, these metadata information needs to be updated synchronously.
[0046] The newly added metadata (ie, the first metadata) may be organized as follows:
[0047] bool is_loaded = false;
[0048] unordered_map<uint32_t, unordered_map<uint64_t, uint32_t>> agg_shards;
[0049] Among them, is_loaded is used to represent whether this set of metadata has been loaded into memory. The default value is false, indicating that it has not been loaded into memory. This set of metadata can be loaded into memory during garbage collection or batch metadata submission, and is_loaded is set to true.
[0050] The first 32-bit key in Agg_item can be used to represent the sequence number of a shard (for example, a shard shard), the second 64-bit key represents the ID of an internal object, and the third 32-bit value represents the valid data length of the internal object.
[0051] Optionally, taking the effective capacity of the data pool as 27.075 billion bytes (Terabyte, abbreviated as TB) as an example, each internal object is 4MB in size. Then there are 70,975,488 internal objects in the cluster. There are 4096 PGs in the cache pool of the cluster. Then there are 17,328 internal objects distributed in each PG. The metadata information (i.e., the first metadata) of 16 internal objects is stored in each shard. The space occupied by a single shard is 192. There are 1084 shards in each PG. In an 8k random write scenario, in the extreme case, when flushing 4MB of business object data at a time, it corresponds to 512 internal objects and 512 shards. Then the amount of metadata that needs to be additionally submitted to the key-value database (kv) for a single flush is 96KB.
[0052] Optionally, the above first metadata can be stored in the key-value database in the form of key-value pairs.
[0053] Optionally, in response to the change of the initial business object, that is, the initial business object changes to the target business object, at least one first metadata can be obtained from the storage unit in memory. From at least one first metadata, the initial metadata corresponding to the initial business object can be determined, and the initial metadata can be used to represent the storage state of the data in the historical internal object.
[0054] Optionally, in this embodiment, the business object is associated with the specific physical location (i.e., the internal object) in the storage system through the first metadata, so as to achieve efficient access and management of data.
[0055] Optionally, in response to the modification of a business object, i.e., when the initial business object changes to a target business object, the system determines the effective data length of the internal object (historical internal object) corresponding to the initial business object from the first metadata (such as agg_shards) stored in memory. This step is crucial for updating the forward and reverse metadata relationships, ensuring the correct creation of new internal objects and the update of the effective data length of old internal objects.
[0056] For example, in response to the change of the initial business object to a target business object, the internal object corresponding to the target business object is reconstructed, and it is determined whether to recycle the historical internal object storing the initial business object. In response to the change of the initial business object to a target business object, the system can first read the agg_shards data structure (i.e., the initial metadata) in memory, which stores the effective data length information of all internal objects. According to the forward metadata information of the initial business object, the associated historical internal object can be located. Further, in response to the change of the initial business object, the effective data length information in the initial metadata can be changed to obtain the target metadata. Based on the target metadata, the effective data length stored in the historical internal object can be determined. If part of the data is overwritten by the new target business object, the effective length of the overwritten data part becomes 0, indicating that this part of the data can be garbage collected.
[0057] For another example, assume that the initial business object A contains data segments D1 and D2, which are stored in internal objects agg1 and agg2 respectively. Now the business object A is modified, D1 is replaced by new data D1′, and D2 remains unchanged, forming the target business object A′. From at least one first metadata, the initial metadata corresponding to the changed initial business data is determined. In agg_shards, there are the following records:
[0058]
[0059] Among them, agg1 and agg2 can represent the IDs of internal objects respectively, and 4096 and 2048 can represent their respective effective data lengths respectively.
[0060] When writing the target business object A′, the system will check the effective data length of agg1 and update it to 0 (assuming D1′ completely overwrites D1). At the same time, a new internal object agg1′ is created for D1′, and the information in agg_shards is updated to obtain the target metadata:
[0061]
[0062] Further, by judging the target metadata, it is determined whether it is necessary to recycle the historical internal object. It should be noted that the magnitudes of the above numbers are only for illustration, and the English of the above parameters is only for illustration. There are no specific restrictions on the above content here.
[0063] In this embodiment, the update process of the initial metadata is completely completed in the memory, without reading the metadata of the internal object from the hard disk, saving the input / output (I / O) operation.
[0064] For example, if the effective data length of agg1 becomes 0, it indicates that agg1 is now garbage data and can be recycled. The new internal object agg1′ is used to store the updated data D1′, and its effective data length is recorded in agg_shards.
[0065] Optionally, when the initial business object changes to the target business object, the system can first read the internal object metadata information (initial metadata) related to the initial business object from the memory. The initial metadata can be stored in agg_shards and can reflect the storage state of the data in the historical internal object, that is, the effective data length.
[0066] Optionally, the system analyzes the newly written data and updates the effective data length of the relevant internal object in agg_shards. If the new data completely overwrites the old data, the effective data length of the old internal object will be set to 0, indicating that all the data of the internal object is no longer valid, and thus the changed target cloud data can be obtained.
[0067] Optionally, after obtaining the target metadata, the system can determine the garbage collection policy based on the internal object with an effective data length of 0 in the target metadata. This policy generally means that if the effective data length of an internal object is 0, the internal object can be marked as recyclable, thereby releasing the storage space it occupies. Among them, the above recycling policy can be a garbage collection policy, which can be used as a rule for determining whether to recycle the historical internal object.
[0068] For example, assume that the initial business object B contains two pieces of data, Data1 and Data2, which are stored in internal objects agg1 and agg2 respectively. The effective data lengths of agg1 and agg2 in agg_shards are 4096 and 2048 bytes respectively. In response to the initial business object changing to the target business object, the effective data length information of the internal objects agg1 and agg2 associated with B in agg_shards can be read. B changes, Data1 is completely overwritten by the new Data1′, while Data2 remains unchanged. The effective data length of agg1 in agg_shards can be updated to 0, and a new internal object agg3 is created for Data1′, and its effective data length is updated to 4096. At this time, the target metadata includes the effective data length of 0 for agg1, 2048 for agg2, and 4096 for the new object agg3. The agg1 in the target metadata can be checked and it is found that its effective data length is 0, so it is confirmed that agg1 is a recyclable internal object. The effective data length of agg2 is not 0, indicating that there are still business objects referencing the data in it, so it will not be recycled. - For agg3, since it is newly created, it has nothing to do with the recycling policy.
[0069] The present invention is a method for reverse metadata management in distributed storage. This method reduces the I / O paths for metadata updates in the write-back process and the garbage collection process, reduces the access to hard disk data, reduces the traversal of target data, and directly accesses the target data in memory and writes it to disk after update, which can save central processing unit resources and hard disk I / O resources and achieve the purpose of improving I / O performance.
[0070] Optionally, through the above steps, the system not only efficiently processes the update of business objects, but also determines the garbage collection policy based on the updated metadata (target metadata), so as to quickly identify and recycle the internal objects that are no longer referenced, and optimize the use of storage space. This method significantly reduces the read and write operations of the hard disk and improves the response speed and performance of the system.
[0071] Through the above method, when the business data of the initial business object changes, the initial business object changes to the target business object, then further the initial metadata corresponding to the initial business object can be determined, and the initial metadata is adjusted to obtain the target cloud data. Further, based on the target metadata, it can be determined whether to perform a recycling process on the historical internal objects, thereby solving the technical problem of low efficiency in recycling internal objects and achieving the technical effect of improving the efficiency of recycling internal objects.
[0072] In an exemplary embodiment, changes are made to the initial metadata to obtain target metadata, including: determining the data volume corresponding to the service data; changing the valid data length information in the initial metadata according to the data volume, where the valid data length information is used to characterize the data storage volume in the historical internal object.
[0073] In this embodiment, the data volume of the changed service data can be determined, and according to this data volume, the valid length information in the initial metadata can be changed to obtain the target metadata. At this time, the modified target metadata can be used to characterize the size of the data stored in the historical internal object. Based on the size, it can be determined whether to perform a recycling process on the historical internal object.
[0074] Optionally, when the initial service object changes to the target service object, the system needs to change the metadata (initial metadata) stored in the memory to reflect the updated situation of the service data. The valid data length information in the initial metadata is crucial and is used to indicate the actual occupied space of the data in the internal object. The process of changing the metadata includes determining the data volume of the new data and accordingly adjusting the valid data length information of the internal object. This process ensures that the system's garbage collection mechanism can accurately identify which data can be recycled and optimizes the management of storage resources.
[0075] Optionally, first calculate the actual size of the newly written service data Data_new. This step is to ensure that the change in the data volume can be accurately reflected when the metadata is updated. Based on the above data volume, the internal object associated with Data_new can be found, and then the valid data length information of the internal object and its information in agg_shards (i.e., the initial metadata) can be updated. If Data_new completely overwrites the old data, the valid data length of the old internal object will be set to 0; if Data_new is an incremental update (i.e., only a part of the old data is replaced), the valid data length of the corresponding internal object will be adjusted according to the data volume of Data_new.
[0076] For example, assume that the initial business object B contains two pieces of data, Data1 (4096 bytes) and Data2 (2048 bytes), which are stored in internal objects agg1 and agg2 respectively. The initial valid data lengths of agg1 and agg2 in agg_shards are 4096 and 2048 bytes respectively. It can be determined that the data volume B corresponding to the business data has changed. Data1 is partially overwritten by new Data1′ (2048 bytes), while Data2 remains unchanged. The system changes the valid data length information in the initial metadata to determine that the data volume of Data1′ is 2048 bytes. Since Data1′ only replaces part of the data in Data1, the valid data length of agg1 needs to be reduced from 4096 to 2048 to reflect the actual data occupancy of Data1′. The valid data length of agg2 remains unchanged at 2048 bytes.
[0077] Furthermore, after Data1′ is written, the system can also consider creating a new internal object agg3 to store Data1′, and setting the valid length of the overwritten data segment in agg1 to 0, indicating that this part of the data has become garbage data. At the same time, update the valid data length information about agg3 in agg_shards to 2048 bytes.
[0078] Figure 3 is a schematic diagram of the correspondence between business objects and internal objects according to an embodiment of the present application. As Figure 3 shown, the data fragments contained in business objects (such as object 301, object 302, object 303) can be aggregated into internal objects (such as 4M object 304) in the system. These internal objects can be stored in the data pool, and the relationship between business objects and internal objects is recorded through metadata, forming forward and reverse relationship chains.
[0079] Optionally, as Figure 3 shown, the data fragments of each business object exist both in the cache pool 305 and in the data pool 306, reflecting the dual-pool storage characteristic of the data. Among them, business data is written from the client to the cache pool 305, aggregated to form a 4M object 304), and then flushed down to the data 306 to finally achieve persistent storage of the data.
[0080] In summary, through the above implementation methods, the system can efficiently manage the valid data length information of internal objects, so that when business objects are updated, the storage state can be quickly adjusted. The process of changing metadata is carried out in memory, avoiding frequent read and write operations on the hard disk and improving the performance of the storage system. At the same time, the timely updated valid data length information provides accurate data basis for subsequent garbage collection, which helps to optimize the use of storage space.
[0081] In an exemplary embodiment, based on the target metadata, a recycling strategy for historical internal objects is determined, including: in response to the valid data length information being less than a first target value, determining that the recycling strategy is to recycle the historical internal objects.
[0082] In an exemplary embodiment, based on the target metadata, a recycling strategy for historical internal objects is determined, including: in response to the valid data length information being greater than or equal to the first target value, determining that the recycling strategy is not to recycle the historical internal objects.
[0083] In this embodiment, the above first target value can be a preset value, for example, it can be 0. The above valid data length information can be used to characterize the data volume of the historical internal object data.
[0084] Optionally, in a distributed storage system, when business data (such as initial business data or target data) is updated, the system modifies the initial metadata corresponding to the business data to generate target metadata, where the target metadata may include the latest status of the valid data length information about the internal object. Therefore, based on the target metadata, the system can determine whether to recycle the historical internal objects, and whether to recycle the historical internal objects depends on the comparison result between the valid data length information and the preset first target value.
[0085] Optionally, in response to the valid data length information being less than the first target value, determining that the recycling strategy is to recycle the historical internal objects. Among them, the above first target value can be 0, indicating that there is no valid data in the internal object. When the valid data length information in the target metadata is equal to or less than this first target value, the system will determine that the internal object is in a discarded or no longer used state, and thus determine that the recycling strategy is to recycle this historical internal object. This ensures the effective management of storage space and avoids useless data occupying storage resources.
[0086] Optionally, in response to the valid data length information being greater than or equal to the first target value, it can be determined that the recycling strategy is not to recycle the historical internal objects. That is, when the valid data length information of the internal object in the target metadata is greater than or equal to the first target value, this indicates that there is still valid data stored in the internal object, and it may still be referenced by one or more business objects. In this case, the system determines that the recycling strategy is not to recycle this historical internal object to maintain data integrity and business operation continuity.
[0087] For example, assume that the internal object agg1 initially stores 4MB of data of the business object X, and the valid data length information is 4MB. Later, the business object X is modified, and the new data completely overwrites the original data in agg1, with a data volume of 2MB. After the modified write operation, the valid data length information of agg1 is updated to 2MB in the target metadata. When checking the target metadata of agg1 and finding that its valid data length information is 2MB, the system determines that there is no need to recycle agg1 in response to this information being greater than or equal to the first target value of 0. However, if the modified write operation of X overwrites all the data in agg1 again and its valid data length becomes 0MB, then the system determines the recycling strategy as recycling agg1 in response to this information being less than the first target value of 0.
[0088] In this embodiment, for internal objects shared or partially covered by multiple business objects, the system will more carefully track and update the valid data length information. For example, if the original data of the internal object agg2 is 4MB, 2MB of which is overwritten by the business object X, and the remaining 2MB is still used by other business objects, then the valid data length in the updated target metadata of agg2 may be 2MB. In this case, the system determines the recycling strategy as not recycling agg2, but retaining its remaining valid data part until all business objects referencing this internal object have completed data updates and the valid data length information in agg2 drops to 0.
[0089] In summary, the recycling strategy decision based on the valid data length information in the target metadata not only improves the management efficiency of storage space and system performance, but also simplifies the logical processing of garbage collection, and is a key strategy for efficiently handling data changes and storage space management in a distributed storage system.
[0090] In an exemplary embodiment, in response to an initial business object changing to a target business object, determining, from at least one first metadata stored in memory, initial metadata corresponding to the initial business object includes: in response to the initial business object changing to the target business object, determining a reverse relationship corresponding to the initial business object, where the reverse relationship is used to represent the mapping relationship between a historical internal object and the initial business object; based on the reverse relationship, determining the historical internal object mapped by the initial business object; determining the identity information of the historical internal object; and determining, as the initial metadata, the first metadata including the identity information from the at least one first metadata.
[0091] In this embodiment, the above reverse relationship can be used to recycle internal objects to avoid waste of space. During the operation of the business, the business objects are modified and written. Each time the modified and written data is aggregated into a new internal object, while the data of the old internal object is discarded as old data to generate garbage data, and the storage space of the data pool occupied by these garbage data needs to be recycled.
[0092] For the modification and writing of business objects, in addition to writing data, it is also necessary to update their forward and reverse relationships simultaneously. Updating the reverse relationship includes two parts. One is to update the valid data length in the old internal object, and the other is to insert the reverse relationship into the new internal object. To update the valid data length in the old internal object, it is necessary to first read the metadata of the internal object and then write it after modification. In this way, the method of updating the reverse relationship of the internal object requires reading first and then writing, which has a greater impact on performance.
[0093] In this embodiment, if the initial business object is transformed into the target business object, the reverse relationship corresponding to the initial business object can be determined. Based on the reverse relationship, the historical internal object mapped by the initial business object can be determined. Based on the identity information corresponding to the historical internal object, the initial metadata in at least one first metadata can be determined.
[0094] Optionally, in a distributed storage system, when a business object is updated, the system needs to determine which internal objects are associated with the business object in order to update its initial metadata. This process involves the search and analysis of the reverse relationship, as well as the determination of the identity information of the internal object, and finally filters out the metadata related to the change of the business object from the metadata set in the memory.
[0095] Optionally, when a business object (initial business object) changes, the system first searches for the reverse metadata relationship of the business object, which is usually stored in agg_shards. The reverse relationship indicates which internal objects store the data of the business object and the specific location of the data in the internal object. By analyzing the reverse relationship, the system can determine all the historical internal objects related to the initial business object. This step ensures that the system can find all the internal object metadata that may need to be updated. Each internal object has a set of identity information, such as the internal object ID and storage location, etc. The system needs to determine this information so that it can accurately locate the correct internal object when updating the metadata later. Finally, the system can filter out the metadata containing the identity information of the historical internal object from all the first metadata stored in the memory. These metadata are the initial metadata directly related to the change of the business object. After determining the initial metadata, the system can further update its valid data length information to reflect the actual state of the business object data.
[0096] For example, assume that the initial business object A is stored in the internal objects agg_A1, agg_A2, and agg_A3. After an update operation, part of the data of A is overwritten by the new data Data_new (2MB), and Data_new is aggregated into the new internal object agg_A4. At this time, it is necessary to determine which historical internal objects have been modified and update their metadata. The initial cloud data corresponding to the initial business object can be modified through the following steps: The reverse relationship of A can be searched to determine that agg_A1, agg_A2, and agg_A3 are associated with A. Based on the reverse relationship, the mapping relationship between A and agg_A1, agg_A2, and agg_A3 can be confirmed, and these historical internal objects can be identified. Determine the identity information of the historical internal objects. For example, the ID of agg_A1 is 0x1A2B3C4D, and the storage location is / data / Pool1 / agg_A1. Further, from the first metadata stored in the memory, the metadata containing the identity information of agg_A1, agg_A2, and agg_A3 can be filtered out. These metadata are the initial metadata related to the change of A. For example, the initial metadata may record that the effective data length of agg_A1 is 4MB, the effective data length of agg_A2 is 2MB, and the effective data length of agg_A3 is 1MB.
[0097] In this embodiment, in a distributed storage environment, a business object may be associated with multiple internal objects, and these internal objects may be distributed on different storage nodes. Therefore, when determining the internal objects and metadata related to the change of a business object, it is necessary to consider the data distribution and consistency issues in a distributed system.
[0098] In summary, through the above steps, the distributed storage system can efficiently process the update of business objects, ensure that the metadata of the associated internal objects can be updated accurately and in a timely manner, while optimizing the use of storage space and the garbage collection mechanism, and improving the overall performance and data management ability of the system.
[0099] As an alternative embodiment, Figure 4 is a flowchart of a method for recycling internal objects according to an embodiment of the present application. As Figure 4 shown, the method may include the following steps:
[0100] Step S402, find the internal objects that need garbage collection through agg_shards.
[0101] In this embodiment, in the system, the agg_shards data structure is used to store and manage the metadata information of internal objects. When the system needs to perform garbage collection, it directly accesses the agg_shards in the memory to find those internal objects with a zero effective data length. This is because when a modification write operation of a business object occurs, the old internal object data is overwritten by the new internal object data, and the old internal object data thus becomes garbage data. By checking the effective data length of the internal objects in agg_shards, the system can quickly locate which internal objects no longer contain valid data, thereby determining the target internal objects that need to be garbage collected.
[0102] Step S404, read the found internal objects for garbage collection.
[0103] In this embodiment, after determining the internal objects that need to be garbage collected, the system no longer needs to read the metadata information of these internal objects from the hard disk. This is because the metadata information is already resident in the memory in agg_shards, including key information such as the effective data length of the internal objects. By directly accessing this information in the memory, the system can avoid reading operations on the hard disk, reduce the I / O burden, and improve the efficiency of garbage collection. Once it is confirmed that the internal objects no longer contain valid data, the system will perform a garbage collection operation, release the corresponding storage space, update the effective data length information of the internal objects in agg_shards, and synchronously write these updates to the kv database to ensure the persistent storage of the metadata.
[0104] In an exemplary embodiment, determining, in at least one first metadata, the first metadata including identity information as initial metadata includes: searching the database according to the identity information to obtain the initial metadata, where the database is used to store the first metadata.
[0105] In this embodiment, each shard can be written into the kv database with the index as the key.
[0106] In an exemplary embodiment, the method may further include: when the memory restarts, in response to the need to recycle historical internal objects or the data storage situation of historical internal objects changes, loading the initial metadata into the memory.
[0107] In this embodiment, after the service restarts, the effective data length information of the internal objects in the memory is loaded into the memory during down brushing or garbage collection, avoiding loading during the service startup process and affecting the recovery efficiency of the fault scenario. The effective data length of the internal objects in the memory is written to the kv database for disk storage in the form of shard sharding, reducing the amount of data written to the disk each time.
[0108] In an exemplary embodiment, the method may further include: in response to a data storage instruction issued by a client, obtaining a plurality of data to be stored corresponding to the data storage instruction; performing an aggregation process on the plurality of data to be stored to obtain a first internal object; based on the data volume of the data to be stored, constructing second metadata of the first internal object, where the second metadata is at least used to characterize the data storage situation of the first internal object; constructing a mapping relationship between the data to be stored and the first internal object to obtain a first forward relationship; and storing the first forward relationship and the second metadata in a memory.
[0109] In this embodiment, when the distributed storage system receives a data storage instruction from a client, it can perform a series of operations according to the instruction content to efficiently store data and maintain metadata relationships.
[0110] Optionally, after receiving the instruction from the client, the instruction content can be parsed to extract a plurality of data blocks to be stored. These data blocks may come from different business objects, but they will be aggregated into the same internal object for storage. The above data blocks may contain a plurality of data to be stored, and the data to be stored may be business data. The obtained plurality of data blocks can be aggregated to create an internal object (the first internal object). The purpose of aggregation can be to improve storage efficiency and data processing speed. Especially for small I / O scenarios, through aggregation, multiple random small I / Os can be converted into sequential large I / Os, thereby increasing the write bandwidth and reducing latency. Further, metadata information (second metadata) of the first internal object can be constructed according to the aggregated data volume. The second metadata can at least include the data storage situation of the internal object. For example, internal object ID, data size, valid data length, etc. It is the basis for subsequent data retrieval, garbage collection and other operations. It should be noted that this is only an example, and no specific restrictions are imposed on the content included in the second metadata.
[0111] Further, a mapping relationship between the data to be stored and the first internal object can be constructed to obtain a first forward relationship. In this embodiment, to ensure the retrievability of data, a mapping relationship from the business object to the internal object (the first forward relationship) can be constructed. In this way, when the business object data needs to be read, the system can quickly locate the corresponding internal object through the forward relationship. To improve data processing speed, the constructed first forward relationship and the second metadata can be stored in the memory. In this way, in subsequent read, update, garbage collection and other operations, the metadata information can be directly read from the memory, avoiding frequent hard disk read and write operations, thereby improving the system response speed and performance.
[0112] For example, assume that the client issues a storage instruction containing multiple small data blocks, which are from business objects Obj1, Obj2, and Obj3 respectively, with data volumes of 1KB, 2KB, and 3KB. The 1KB data of Obj1, the 2KB data of Obj2, and the 3KB data of Obj3 can be parsed from the storage instruction. These data blocks are aggregated to create a first internal object Agg_Obj with a size of 4KB. The second metadata of Agg_Obj can be constructed, including information such as the ID of Agg_Obj, the data size of 4KB, and the valid data length of 4KB. Further, the mapping relationships between Obj1, Obj2, and Obj3 and Agg_Obj can be established to obtain the first forward relationship, such as Obj1->Agg_Obj(1KB), Obj2->Agg_Obj(2KB), Obj3->Agg_Obj(3KB). The first forward relationship and the second metadata of Agg_Obj can be stored in memory for subsequent metadata management and data retrieval operations.
[0113] Optionally, in the actual operation of the distributed storage system, the processes of data aggregation and metadata construction may be more complex. For example, the system may need to determine data aggregation and the creation of internal objects based on factors such as the data access pattern and the distribution of storage space. At the same time, to ensure data consistency and reliability, the system may need to perform some data verification and consistency check operations before storing the metadata in memory.
[0114] In summary, the processes of data aggregation and metadata construction help improve storage efficiency. Especially for small I / O operations, they can significantly reduce the number of hard disk read and write operations and increase the write speed. Storing the metadata in memory can speed up data retrieval and update, reduce the dependence on hard disk resources, and improve the overall performance of the system. By constructing the forward relationship, the system can quickly locate the physical storage location of business object data, simplify the data retrieval process, and improve the data access speed. The metadata management mechanism in memory enables the system to respond more quickly to the read and write requests of the client, improving the system's response ability and user experience. The efficient management of data aggregation and metadata helps avoid waste of storage space, optimizes the utilization of storage resources, and reduces storage costs.
[0115] Optionally, through the above steps, the distributed storage system can efficiently process the data storage instructions issued by the client, not only improving the efficiency of storage operations, but also optimizing the metadata management and data retrieval processes, enhancing the system's response ability and user experience, while reducing storage costs and improving the utilization rate of storage resources.
[0116] In an exemplary embodiment, the above method may further include: obtaining a data acquisition request sent by a client, where the data acquisition request is used to acquire data to be acquired stored in a data pool; determining a first service object corresponding to the data acquisition request; obtaining a first forward relationship corresponding to the first service object from a memory; based on the first forward relationship, determining identity information of a first internal object mapped by a target service object; and acquiring the data to be acquired from the first internal object deployed in the data pool according to the identity information.
[0117] In this embodiment, the memory may contain forward relationships corresponding to multiple service objects. The forward relationship may include the first forward relationship, and the forward relationship may be used to read data of a service object.
[0118] For example, after receiving a request from a client to read a service object oid in the range [offset1, length2], retrieve the forward relationship of the object oid, obtain the corresponding internal object name agg_oid and the range [offset2, length2] of this range, then read data from the data pool according to the obtained forward relationship, and then return it to the client.
[0119] Optionally, in a distributed storage system, a data acquisition request initiated by a client is a trigger point for the system to perform a data reading operation. When the system receives a data acquisition request, it will execute a series of steps to efficiently retrieve and return the data required by the client. These steps make full use of the metadata information stored in the memory, avoid unnecessary hard disk reads, and improve the data reading efficiency.
[0120] Optionally, obtain a data acquisition request sent by a client: The client sends a data acquisition request to the system, requesting to acquire specific data stored in the data pool. The above request can be parsed to extract key information in the request, such as the identifier of a service object, the offset and length of the data, etc. Based on the parsed request information, the system can determine the service object for which the client requests to acquire data. For example, the client may request to acquire data starting from offset1 and with a length of length1 in the Obj1 service object. Using the service object identifier, the system quickly retrieves the forward relationship information related to the service object from the memory. The forward relationship information records the storage location of the service object data in the internal object, which is crucial for quickly retrieving data. By analyzing the forward relationship, the identity information of the internal object storing the requested data can be determined, such as the internal object ID. This means that the system has located the actual physical storage location of the data. Further, according to the lD of the internal object, the corresponding data can be directly read from the data pool and returned to the client. This process becomes very efficient because it uses the metadata information in the memory.
[0121] For example, assume that the client sends a data acquisition request to obtain data starting from offset1 with a length of length1 in the business object ObjX. The system receives the client's request and parses information such as the ObjX identifier, offset1, and length1 in the request. If the system determines that the client requests to obtain data of ObjX, it can obtain the forward relationship information of ObjX from the memory and find that part of the data of ObjX is stored in the internal object agg_X1. Through the forward relationship information, the system can determine the actual storage location of the requested data of ObjX as agg_X1 and obtain identity information such as the ID of agg_X1. Further, based on the ID of agg_X1, the system reads the corresponding data from the data pool and returns this part of the data to the client, satisfying the client's data acquisition request. It should be noted that the above method for data processing is only for illustrative purposes and is not specifically limited here.
[0122] Optionally, in actual applications, business objects may contain a large amount of data, which are scattered and stored in multiple internal objects. Therefore, the system needs to support querying and integrating data in multiple internal objects to satisfy the client's complete data acquisition request. In addition, to improve efficiency, the system may also pre-cache some popular data so that for frequently accessed business object data, it can be directly obtained from the cache, further reducing the read operations on the data pool.
[0123] In summary, by utilizing the forward relationship information stored in the memory, the system can quickly locate the storage location of the data, avoid unnecessary hard disk reads, and significantly improve the data access speed. The use of the forward relationship reduces the I / O operations of the hard disk, simplifies the logical processing of data retrieval, and makes the process of data retrieval and integration simpler and more efficient.
[0124] Optionally, by utilizing the forward relationship information stored in the memory, the distributed storage system can efficiently and accurately respond to the client's data acquisition request, not only improving the data access speed, reducing the I / O burden, but also optimizing the data retrieval logic, enhancing data consistency, and improving the user experience. It is one of the key mechanisms for efficient data management in the distributed storage system.
[0125] As an optional embodiment, Figure 5 is a flowchart of the method for flushing the internal object according to the embodiment of the present application, as Figure 5 shown, the method may include the following steps:
[0126] Step S502, obtain the forward relationship of the business object.
[0127] In this embodiment, the business object [offset, length] is obtained. According to the identifier of the business object (such as ObjX) and the offset (offset) and length (length) of its data, the forward relationship information associated with the business object can be obtained from memory or persistent storage. The forward relationship information records the storage location of the business data in the internal object, including the ID of the internal object and the specific range of the data in the internal object.
[0128] Step S504, update agg_shards according to the forward relationship.
[0129] In this embodiment, after obtaining the forward relationship, these information can be read and analyzed to determine which internal object's valid data length needs to be updated. The system applies these update operations to the agg_shards data structure in memory. The information in agg_shards includes the ID of each internal object and its valid data length. By updating agg_shards, the system can ensure that the metadata information of the internal object is consistent with the actual situation of the business object data.
[0130] Step S506, write the updated agg_shards to disk.
[0131] In this embodiment, in order to persistently store the updates of these metadata information, the system needs to write the content of the updated agg_shards data structure to the kv database for disk writing operation. In a distributed storage system, the kv database is usually used to store metadata information to ensure that the metadata information can still be correctly read and used even after the system restarts or fails.
[0132] In an exemplary embodiment, the method may further include: transmitting the data to be obtained to the client, and in response to obtaining the data to be obtained from the first internal object, modifying the valid data length information of the third metadata corresponding to the first internal object, where the third metadata is stored in memory.
[0133] Optionally, in a distributed storage system, when the data in the data pool is requested by the client and successfully obtained, the system needs to update the metadata information stored in memory to reflect the change in the valid state of the data. This update process is crucial for maintaining the accuracy and integrity of the data.
[0134] In this embodiment, after successfully obtaining the data required by the client from the first internal object deployed from the data pool, the system transmits this data back to the client to satisfy the data acquisition request. Since the client has successfully obtained the data, the state of this data in the internal object may have changed. For example, if the obtained data is part of a business object, then this part of the data may no longer be considered "pending" or "valid". Therefore, the system needs to update the metadata information (the third metadata) stored in the memory to reflect the change in the effective data length of the internal object. This helps subsequent operations such as garbage collection to ensure the effective management of storage space.
[0135] For example, assume that the client requests to obtain data with a length of length1 starting from offset1 in the business object ObjA, and this part of the data is stored in the internal object agg_A1. The above process may include the following steps: The system reads the data from agg_A1 and then transmits this part of the data back to the client through the network. Since the data has been obtained by the client, the system updates the corresponding third metadata in the memory for agg_A1 to reduce its effective data length information. For example, if the original effective data length of agg_A1 is 10MB and the client obtains data with length1 of 2MB, then the updated effective data length information should be 8MB.
[0136] In summary, by timely updating the effective data length information of the third metadata of the first internal object, the system can ensure the accuracy of the metadata, thereby making correct decisions in subsequent data retrieval, storage space management, garbage collection, and other operations. The update of the effective data length information helps the system accurately identify which data is invalid, thereby releasing storage space in the garbage collection operation and optimizing the utilization of storage resources. In a concurrent reading scenario, ensuring that the effective data length information of the metadata is updated in a timely manner after each reading operation helps improve data consistency and accuracy, and avoids errors caused by inconsistent data states.
[0137] Optionally, the update of the effective data length information of the metadata is performed in the memory, avoiding frequent hard disk read and write operations, and improving the system's response speed and I / O efficiency.
[0138] In summary, by updating the effective data length information of the third metadata corresponding to the first internal object in the memory after the data is successfully obtained, the distributed storage system can ensure the accuracy of the metadata, optimize the storage space management, and improve data consistency.
[0139] In an exemplary embodiment, the method may further include: in response to a change in service data, performing an aggregation process on the changed service data to obtain a second internal object, where the second internal object is different from the historical internal object; constructing fourth metadata of the second internal object based on the data volume of the changed service data, where the fourth metadata is at least used to characterize the data storage condition of the second internal object; constructing a mapping relationship between the changed service data and the second internal object to obtain a second forward relationship; and storing the second forward relationship and the fourth metadata in a memory.
[0140] In this embodiment, when the service data changes, the distributed storage system needs to execute a series of steps to process the new data and update the metadata to maintain the data correctness and system consistency.
[0141] Optionally, perform an aggregation process on the modified part of the service data to form a new internal object (the second internal object), which is different in content from the original internal object (the historical internal object). The purpose of the aggregation process is to improve the storage efficiency and data processing speed. Especially in the small I / O scenario, multiple random small I / Os can be converted into sequential large I / Os through aggregation, thereby improving the write bandwidth and reducing the latency. Further, metadata information (the fourth metadata) of the second internal object can be constructed according to the size of the changed service data. The fourth metadata may at least include information such as the internal object ID, data size, and valid data length, and is used for subsequent data retrieval, storage space management, and other operations. A mapping relationship (the second forward relationship) can be established between the changed service data and the newly created internal object. In this way, when the client requests to read the changed service data, the system can quickly locate the correct internal object to ensure the data retrievability and consistency.
[0142] Optionally, to accelerate subsequent data retrieval and metadata update operations, the constructed second forward relationship and fourth metadata information can be stored in the memory, which avoids frequent hard disk read and write operations and improves the system response speed and performance.
[0143] For example, assume that a part of the data of business object ObjX has been stored in internal object agg1, and this part of the data of ObjX is modified, with the new data size being 2MB. The modified new data can be aggregated to create a new internal object agg2 for storing the updated data. Metadata information (the fourth metadata) of agg2 can be constructed, including information such as the ID of agg2, the data size of 2MB, and the valid data length of 2MB. Further, a mapping relationship between ObjX and agg2 can be established to obtain the second forward relationship, such as ObjX->agg2(2MB). The second forward relationship and the fourth metadata information of agg2 can be stored in memory for subsequent metadata management and data retrieval operations.
[0144] In summary, by aggregating the changed business data, the system can optimize storage operations, improve write efficiency, and reduce latency. Especially in small I / O scenarios, the effect is remarkable. Storing the mapping relationship between the changed business data and the newly created internal object, as well as the metadata information of the new internal object in memory simplifies the metadata management and update process and improves the data retrieval efficiency.
[0145] Optionally, through the above steps, the distributed storage system can effectively handle changes in business data, not only improving the data storage efficiency and processing speed, simplifying metadata management, but also ensuring data consistency and accuracy, significantly enhancing the user experience and the utilization efficiency of storage space. It is an important mechanism for data change processing and metadata management in the distributed storage system.
[0146] As an optional implementation manner, Figure 6 is an example diagram of a write operation in a distributed storage system according to an embodiment of the present application. As Figure 6 shown, the write operation in the distributed system may include the following: The client initiates a write request and sends the business data to be stored to the cache pool of the distributed storage system to complete the client write operation (Clientwrite). The business object can be encapsulated as Write-Ahead Logging (WAL) data and submitted to disk.
[0147] Optionally, the cache pool can merge the business object data into memory for caching.
[0148] Optionally, after receiving the client's write request, the cache pool temporarily stores the data in memory for caching to accelerate the write process and reduce direct operations on the hard disk. Memory caching is the starting point for data aggregation and processing.
[0149] Optionally, when the business data in the memory reaches a certain amount or meets specific conditions, these business data can be aggregated to form an internal object. The internal object is the aggregation unit of data, and it has its own metadata for recording the status and location information of the data.
[0150] Optionally, after the aggregation is completed, the system writes (flushes) the internal object from the memory to the data pool for persistent storage. Flushing is the process of storing data from the unstable memory to the stable data pool.
[0151] Optionally, when the data is flushed to the data pool, the data in the memory becomes aged data, and these data will no longer be used by the cache pool, but wait to be cleaned or recycled.
[0152] Optionally, after the flushing operation is completed, the system will trigger a callback mechanism to ensure that the data is correctly written to the data pool and synchronously update the metadata and forward relationships. The flushing callback mechanism is a key step in the data writing process to ensure the integrity and consistency of the data.
[0153] Optionally, the data pool is a component in the distributed storage system for persistent storage of data. The internal object data will be stored in the data pool after flushing to ensure the security and persistence of the data.
[0154] In this embodiment, when the client sends a write request, the system first stores the data in the memory cache to reduce the frequency of direct writing to the hard disk and improve the writing efficiency. When the data in the memory cache reaches a certain threshold, the system merges these data into an internal object and performs a flushing operation to write the internal object to the data pool for persistent storage. After the data is flushed, it will trigger the processing of aged memory data and aged WAL objects, which usually involves cleaning the data in the memory and updating the metadata to ensure the storage efficiency and data consistency of the system. After the flushing operation is completed, the system uses the flushing callback mechanism to confirm whether the data is successfully written to the data pool and synchronously update the metadata and forward relationships in the memory to ensure that subsequent data reading and retrieval operations can be carried out correctly. The data in the data pool is persistent and is used to store internal objects and ensure the reliability of the data. The cache pool is a temporary storage area for caching data and metadata to improve the efficiency of read and write operations.
[0155] Among them, the execution subject of the above steps can be a server, a terminal, etc., but is not limited thereto.
[0156] To facilitate the understanding of the implementation manner of this application, relevant scenarios are now explained, but it does not limit this application.
[0157] With the continuous development of information technology, data, as a valuable resource, has gradually attracted people's attention. How to quickly process data resources and obtain the expected results has become one of the key issues in the transformation from resources to assets. People's various activities in work and life will generate data. Collecting this data and then through analysis and processing can obtain very useful information, realizing the transformation from resources to assets, thus catalyzing the rapid development of big data and high-performance computing. Enterprises' operations, traffic management, and public security management will also generate a large amount of data. This type of data often has a specific storage cycle, and it needs to be retrievable and viewable at any time within the cycle. Based on the generation of massive data, data storage, as one of the core elements of data resources, has also entered a period of rapid development.
[0158] Traditional network storage systems use centralized storage servers to store all data. The storage server becomes the bottleneck of system performance and the focus of reliability and security, and cannot meet the needs of large-scale storage applications. Distributed network storage systems adopt an expandable system structure, which not only improves the reliability, availability, and access efficiency of the system, but also is easy to expand, so it is increasingly accepted and recognized by enterprises and institutions. Distributed storage systems generally exist in the form of a storage server cluster. A typical storage server cluster contains 10 storage server nodes. Currently, the largest-scale storage server cluster consists of 1024 storage server nodes, which are used to provide high-performance and massive data storage. Distributed storage servers can be divided into mechanical disk storage, hybrid flash storage, and all-flash storage according to different storage media, and generally make a trade-off plan according to the customer's business model and storage purchase cost.
[0159] With the development of storage media, the proportion of all-flash storage is increasing day by day and gradually becoming the mainstream storage method. To give full play to the performance of all-flash storage media, it is necessary to optimize small I / O scenarios. I / O aggregation is one of the optimization methods. Through techniques such as write-time redirection, multiple random small I / Os are aggregated into sequential large I / Os to improve write bandwidth and reduce latency. This aggregation is usually achieved with the help of a cache, so that the small I / O write operation can return to the client after being completed in the cache, while the background performs processing processes such as I / O aggregation. The small I / Os of business objects are aggregated into internal objects (internal objects refer to objects in a distributed storage system that are aggregated from multiple small I / Os) in the cache and then written into the data pool. The business object maintains the metadata pointing to the location of the internal object, which is called the forward relationship, and the internal object also maintains the metadata pointing to the location of the business object, which is called the reverse relationship.
[0160] In related technologies, for the update of reverse metadata, usually the metadata of the internal object is first read and then the metadata is modified. However, the above method has the technical problem of low efficiency in recycling the internal object.
[0161] To solve the problems encountered in the use of the above algorithm, this application proposes a design method for distributed storage to manage reverse metadata. This solution designs a set of in-memory data structures to store the valid data length of internal objects with minimal memory resources, avoiding loading internal object metadata information from disk. Each time of flushing, directly update the valid data length of the old internal object in memory, and then synchronously flush the updated old internal object information to disk when the business object is flushed to disk; during garbage collection, also perform garbage collection on the internal objects in the data pool according to the valid data length of the internal objects in memory, and at the same time update the valid data length of the internal objects in memory and flush it to disk.
[0162] After the service restarts, the valid data length information of the internal objects in it is loaded into memory during flushing or garbage collection, avoiding loading during the service startup process, which affects the recovery efficiency of the failure scenario. The valid data length of the internal objects in memory is written to the kv database for flushing in the form of shard sharding, reducing the amount of data flushed each time.
[0163] Optionally, this method designs a set of in-memory data structures to store the valid data length of internal objects with minimal memory resources and is convenient for quickly searching and accessing internal objects.
[0164] Optionally, the internal object metadata information stored in memory resides in memory all the time, and there are two disk flushing times for it. One is to flush to disk together with the business object metadata after flushing is completed, and the other is to modify and flush during garbage collection.
[0165] Optionally, during the flushing process, the relevant old internal object metadata information can be directly read from memory and memory can be updated. During the garbage collection process, the internal objects that need to be recycled can be directly retrieved from memory, avoiding retrieval after reading from disk.
[0166] Optionally, after the service restarts, the internal object metadata information is loaded from disk into memory and resides in memory during flushing or garbage collection.
[0167] Optionally, the internal object metadata information can be stored in the kv database in the form of shard, and each time the metadata corresponding to the shard is processed, avoiding full-scale update.
[0168] In this embodiment, a bool type variable is added to each PG, and this type of variable can be used to represent whether the internal object metadata has been loaded into memory.
[0169] Optionally, add a bool variable: Assume that each storage server (i.e., PG) has a bool variable named `is_loaded`. When a PG starts, the initial value of `is_loaded` is `false`, indicating that the internal object metadata of this PG has not been loaded into memory yet. When the metadata is loaded, the value of `is_loaded` becomes `true`.
[0170] In this embodiment, a data structure can be designed to cache the metadata information of internal objects (such as the first metadata) in memory:
[0171] unorder_map<uint32_t, unordered_map<uint64_t, uint32_t>> agg_shards;
[0172] Among them, each shard is written into the kv database with the index as the key.
[0173] Optionally, the above `agg_shards` can be a two-layer hash table structure for caching the metadata information of internal objects in memory. Among them, the key of the first-layer hash table is a 32-bit unsigned integer (i.e., the index of the shard), and the value is the second-layer hash table. The key of the second-layer hash table is a 64-bit unsigned integer (internal object ID), and the value is a 32-bit unsigned integer (the valid data length of the internal object).
[0174] In this embodiment, each time the business object data is flushed down, the valid data length of the internal object in agg_shards is updated by obtaining the old forward relationship.
[0175] Optionally, assume that the business object, A, contains data segments `offset1, length1`, and these data segments are aggregated to form an internal object `agg_oid1`. When the business object, A, needs to modify and write new data segments `offset2, length2` and aggregate them into a new internal object `agg_oid2`, the system will update the information about the valid data length of `agg_oid1` in `agg_shards` and insert the information about `agg_oid2` at the same time. This is not deleting data, but updating the metadata to reflect the data status, that is, the new data is aggregated into a new internal object, and the valid length of the old data is updated.
[0176] In this embodiment, when updating Pginfo, the updated agg_shard can be synchronized to disk.
[0177] Optionally, whenever Pginfo is updated, such as after the storage server restarts or when the metadata needs to be refreshed, the system synchronously updates the data in `agg_shards` and writes this data to persistent storage, such as a KV database.
[0178] In this embodiment, during garbage collection, the internal objects to be recycled are directly checked through `agg_shards` in memory, and there is no need to list and retrieve the internal object metadata from disk anymore.
[0179] Optionally, during the garbage collection process, the valid data length information of the internal objects in `agg_shards` can be checked. If the valid data length of an internal object is 0, it means that the object is no longer referenced by any business object and can be safely recycled. Since `agg_shards` has been resident in memory, the system can directly find the unnecessary internal objects in memory without reading the metadata from the hard disk, thus greatly accelerating the garbage collection speed.
[0180] In the embodiment of the present application, the first metadata is pre-stored in memory, and this first metadata can be used to determine the data storage situation of historical internal objects. Therefore, when the business data of the initial business object changes, the initial business object changes to a target business object, then the initial metadata corresponding to the initial business object can be further determined, and the initial metadata is adjusted to obtain target cloud data. Further, based on the target metadata, it can be determined whether to recycle the historical internal objects, thereby solving the technical problem of low efficiency in recycling internal objects and achieving the technical effect of improving the efficiency of recycling internal objects.
[0181] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present application.
[0182] In this embodiment, a processing device for internal objects is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0183] Figure 7 is a structural block diagram of a processing device for internal objects according to an embodiment of the present application. As Figure 7 shown, the device includes:
[0184] A first determination unit 72, configured to determine, in response to an initial service object changing to a target service object, initial metadata corresponding to the initial service object from at least one first metadata stored in a memory, where the first metadata is used to at least characterize the data storage situation of historical internal objects associated with the initial service object, and the historical internal objects are used to store service data in the initial service object;
[0185] A change unit 74, configured to change the initial metadata to obtain target metadata;
[0186] A second determination unit 76, configured to determine a recycling policy for historical internal objects based on the target metadata, where the recycling policy is a rule indicating whether to recycle historical internal objects.
[0187] Through the above device, when the service data of the initial service object changes, the initial service object changes to a target service object. Then, the initial metadata corresponding to the initial service object can be further determined, and the initial metadata is adjusted to obtain target cloud data. Further, based on the target metadata, it can be determined whether to perform a recycling process on historical internal objects, thereby solving the technical problem of low efficiency in recycling internal objects and achieving the technical effect of improving the efficiency of recycling internal objects.
[0188] It should be noted that the above-mentioned respective modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: all the above modules are located in the same processor; or, the above-mentioned respective modules are separately located in different processors in any combination form.
[0189] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above method embodiments when running.
[0190] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0191] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0192] Optionally, Figure 8 is a block diagram of a computer system structure of an electronic device according to an embodiment of the present application. As Figure 8 shown, the computer system 800 includes a central processing unit 801 (CPU), which can perform various appropriate actions and processes according to a program stored in a read-only memory 802 (ROM) or a program loaded from a storage section 808 into a random access memory 803 (RAM). In the random access memory 803, various programs and data required for system operation are also stored. The central processing unit 801, the read-only memory 802, and the random access memory 803 are connected to each other via a bus 804. An input / output interface 805 (Input / Output interface, i.e., I / O interface) is also connected to the bus 804.
[0193] The following components are connected to the input / output interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a local area network card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disc, a magneto-optical disc, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0194] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0195] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0196] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0197] An embodiment of the present application also provides a computer program. The computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any one of the above method embodiments.
[0198] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0199] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.
[0200] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for processing an internal object, characterized in that: include: In response to the initial business object changing to the target business object, determining initial metadata corresponding to the initial business object from at least one first metadata stored in the memory, wherein the first metadata is used to at least characterize the storage state of data in a historical internal object associated with the initial business object, and the historical internal object is used to store the business data of the initial business object; Modifying the initial metadata to obtain target metadata; Based on the target metadata, a recycling policy of the historical internal object is determined, wherein the recycling policy is used to indicate a rule of whether to recycle the historical internal object.
2. The method according to claim 1, characterized in that The step of modifying the initial metadata to obtain target metadata includes: Determining the data volume corresponding to the business data; According to the data volume, the valid data length information in the initial metadata is modified, wherein the valid data length information is used to characterize the data storage volume in the historical internal object.
3. The method according to claim 2, characterized in that The determining, based on the target metadata, a recycling strategy for the historical internal object includes: In response to the valid data length information being smaller than a first target value, determining the recycling strategy to recycle the historical internal object.
4. The method according to claim 3, characterized in that The determining, based on the target metadata, a recycling strategy for the historical internal object includes: In response to the valid data length information being greater than or equal to the first target value, determining the recycling strategy as not requiring recycling processing of the historical internal object.
5. The method according to claim 1, characterized in that In response to the initial business object changing to the target business object, determining the initial metadata corresponding to the initial business object from at least one first metadata stored in the memory includes: In response to the initial business object changing to the target business object, determining a reverse relationship corresponding to the initial business object, wherein the reverse relationship is used to characterize a mapping relationship between the historical internal object and the initial business object; Based on the reverse relationship, determining the historical internal object mapped by the initial business object; Determining identity information of the historical internal object; The first metadata including the identity information in at least one of the first metadata is determined as the initial metadata.
6. The method according to claim 5, characterized in that The step of determining the first metadata including the identity information in at least one of the first metadata as the initial metadata includes: The database is searched according to the identity information to obtain the initial metadata, wherein the database is used to store the first metadata.
7. The method according to claim 6, characterized in that The method further comprises: When the memory is restarted, in response to the need to recycle the historical internal object or a change in the data storage situation of the historical internal object, the initial metadata is loaded into the memory.
8. The method according to claim 1, characterized in that The method further comprises: In response to a data storage instruction issued by the client, obtaining a plurality of to-be-stored data corresponding to the data storage instruction; Aggregate the multiple data to be stored to obtain a first internal object; Based on the data volume of the data to be stored, construct second metadata of the first internal object, wherein the second metadata is at least used to characterize the data storage situation of the first internal object; Constructing a mapping relationship between the data to be stored and the first internal object to obtain a first positive relationship; The first forward relationship and the second metadata are stored in the memory.
9. The method according to claim 8, characterized in that The method further comprises: Obtaining a data acquisition request sent by the client, wherein the data acquisition request is used to obtain the data to be acquired stored in the data pool; Determine a first business object corresponding to the data acquisition request; Acquire the first forward relationship corresponding to the first business object from the memory; Based on the first forward relationship, the identity information of the first internal object mapped to the target business object is determined; and according to the identity information, the data to be obtained is obtained from the first internal object deployed in the data pool.
10. The method according to claim 9, characterized in that The method further comprises: The data to be obtained is transmitted to the client, and in response to obtaining the data to be obtained from the first internal object, effective data length information of third metadata corresponding to the first internal object is modified, wherein the third metadata is stored in the memory.
11. The method according to claim 1, characterized in that: The method further comprises: In response to a change in the business data, performing aggregation processing on the changed business data to obtain a second internal object, wherein the second internal object is different from the historical internal object; Based on the changed data volume of the business data, construct fourth metadata of the second internal object, wherein the fourth metadata is at least used to characterize the data storage situation of the second internal object; A mapping relationship between the changed business data and the second internal object is constructed to obtain a second forward relationship; and the second forward relationship and the fourth metadata are stored in the memory.
12. A device for processing an internal object, characterized in that: include, A first determining unit is configured to determine, in response to the initial business object changing to a target business object, initial metadata corresponding to the initial business object from at least one first metadata stored in a memory, wherein the first metadata is used to at least characterize a data storage situation of a historical internal object associated with the initial business object, and the historical internal object is used to store business data in the initial business object; A modification unit, used for modifying the initial metadata to obtain target metadata; The second determining unit is used to determine a recycling policy for the historical internal object based on the target metadata, wherein the recycling policy is used to indicate a rule for whether to recycle the historical internal object.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 11 when executed by a processor.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 11 are implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Data processing method and electronic equipment
CN114895848A
Garbage recycling method, device and equipment
CN117785039A
Data management method and device and hard disk
CN117873383A
Battery data storage method and device and battery management system
CN118778890A
Universal middleware system built based on ZNS SSD and related method
CN119045753A