Data writing method and device based on distributed storage system, equipment and medium
By identifying data objects in a distributed storage system and allocating fixed storage space when writing for the first time, a mapping relationship is established, which solves the metadata redundancy and fragmentation problems in the encapsulation and mapping process and improves data writing efficiency.
Patent Information
- Application Number
- CN202510799075.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-10-17
AI Technical Summary
In a distributed storage system, the existing technology causes additional metadata redundancy and cumbersome space management problems during data writing due to the dual process of encapsulation and mapping, and easily leads to data fragmentation.
By identifying the data object and directly allocating fixed storage space when writing for the first time, a mapping relationship is established. When writing subsequently, the storage location is directly queried to write data, avoiding multiple space allocation and recycling operations.
It reduces metadata redundancy, avoids data fragmentation, and improves the data writing efficiency of distributed storage systems.
Smart Images

Figure CN120803345A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of storage, in particular to a data writing method and device based on a distributed storage system, equipment and medium. BACKGROUND
[0002] Distributed storage stores data on multiple nodes, so even if some nodes fail, data can still be recovered from other nodes, thereby improving data reliability and availability. Moreover, a distributed storage system disperses data to multiple nodes to achieve load balancing, providing higher read-write performance and throughput. The distributed storage system usually uses high-performance hardware devices, supports data caching and compression technology, and further improves system performance.
[0003] In some related technologies, the integrity and atomicity of data in a single object need to be ensured in a distributed storage system, which results in the bottom block device of the storage system being unable to be fully utilized; as shown in the prior art, the bottom block device needs to be encapsulated into an object during data writing, and then mapped, and the double processes of encapsulation and mapping will generate additional metadata, thereby causing redundancy in space management, and the multiple space allocation and recovery operations are cumbersome, and when rewriting data, the data fragmentation problem will also be caused. Figure 1 SUMMARY
[0004] The present application provides a data writing method and device based on a distributed storage system, equipment and medium, which first identifies a data object when writing data, and directly writes the data object to the allocated fixed space when writing for the first time. When subsequent writing is needed, the storage location of the previous data object is queried, and the data is directly written to the original storage location, thereby reducing the cumbersome problem of multiple space allocation and recovery operations.
[0005] The present application provides a data writing method based on a distributed storage system, the method comprising:
[0006] receiving data to be written, and identifying that the data to be written comprises a data object;
[0007] In response to confirming that the data object is first-time writing data, allocating an initial storage space for the data object;
[0008] configuring the initial storage space to be associated with a logical section in the distributed storage system to form a mapping relationship;
[0009] In response to performing a non-first-time writing operation on the initial storage space, querying the storage location of the logical section associated with the initial storage space in combination with the mapping relationship, and writing the non-first-time writing operation data to the storage location.
[0010] In one specific embodiment, receiving data to be written, and identifying that the data to be written comprises a data object, specifically comprises: receiving the data to be written from a storage disk in the distributed storage system, decoding the data to be written to obtain decoded information; if the decoded information comprises a preset keyword, confirming that the data to be written comprises the data object, and the data object is first write data.
[0011] In one specific embodiment, in response to confirming that the data object is first write data, configuring the initial storage space as a fixed capacity, and setting the fixed capacity as a minimum allocation unit; in response to performing the first write data operation, creating a logical container, and updating the logical container to the mapping relationship; performing the first write data object write operation to the logical container.
[0012] In one specific embodiment, the logical container comprises one or more physical storage blocks, and the method further comprises: configuring the logical segment to be mapped to the physical storage block, and setting the physical storage block as the storage location; setting the logical segment to have a first identifier, and setting the physical storage block to have a second identifier; configuring the logical segment and the physical storage block to be correspondingly associated, and the first identifier and the second identifier one-to-one correspond.
[0013] In one specific embodiment, before the method in response to performing the non-first write operation of the initial storage space, the method further comprises: judging whether the length of the non-first write operation data is less than the minimum allocation unit; if the length of the non-first write operation data is less than the minimum allocation unit, reading a target data block with a capacity of the minimum allocation unit size; and merging the non-first write operation data into the target data block.
[0014] In one specific embodiment, merging the non-first write operation data into the target data block specifically comprises: reading existing data in the target data block, combining the non-first write operation data to replace a to-be-modified part in the existing data as modified data; taking the modified data and the remaining data in the target data block as modified target data block data; and synchronizing the modified target data block data to the storage segment.
[0015] In one specific embodiment, after querying the storage location of the logical segment associated with the initial storage space in combination with the mapping relationship, the method further comprises: reading a reference flag in the logical segment on the storage location, and if the reference flag is a value greater than or equal to 1, confirming that the storage location is reusable; writing the non-first write operation data to the storage location, and setting the physical storage block unchanged.
[0016] The application further provides a data writing device based on a distributed storage system for implementing the method, and the device comprises:
[0017] An identification module is configured to receive data to be written and identify that the data to be written comprises a data object.
[0018] An allocation module is configured to allocate initial storage space for the data object in response to confirming that the data object is first-time writing data.
[0019] A configuration module is configured to configure the initial storage space to be associated with a storage segment in the distributed storage system to form a mapping relationship.
[0020] An execution module is configured to, in response to performing a non-first-time writing operation based on the initial storage space, query a storage location of the storage segment associated with the initial storage space in combination with the mapping relationship, and write data of the non-first-time writing operation to the storage location.
[0021] The application further provides an electronic device, which comprises a memory configured to store a computer program and a processor configured to implement the steps of any of the data writing methods based on a distributed storage system when executing the computer program.
[0022] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is configured to implement the steps of any of the data writing methods based on a distributed storage system when executed by a processor.
[0023] The application further provides a computer program product, which comprises a computer program, and the computer program is configured to implement the steps of any of the data writing methods based on a distributed storage system when executed by a processor.
[0024] In the application, a data object is identified when data is written, and then it is determined whether the data object is first-time writing data. If yes, initial storage space is allocated for the data object, and the initial storage space is associated with the distributed storage system to form a mapping relationship. When a re-writing operation is performed on the initial storage space, the mapping relationship is queried to obtain a storage location associated with the initial storage space, and data of the non-first-time writing operation is directly written to the storage location. The problem of additional metadata generated in the encapsulation and mapping process is solved, and the problem of data fragmentation is avoided. Meanwhile, in the application, non-first-time writing data can be performed at the original storage location, so that when subsequent data writing is performed, multiple space allocation and recycling are not required, and the data writing efficiency of the distributed storage system is improved. BRIEF DESCRIPTION OF DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort based on these drawings.
[0026] Figure 1 The data writing process based on the distributed storage system in the background art provided for the embodiments of the present application is shown in the figure.
[0027] Figure 2 The data writing process based on the distributed storage system provided for the embodiments of the present application is shown in the figure.
[0028] Figure 3 The data writing method based on the distributed storage system provided for the embodiments of the present application is shown in the figure.
[0029] Figure 4 The data writing device based on the distributed storage system provided for the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION
[0030] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present application.
[0031] It should be noted that, in the description of the present application, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the processes, methods, articles or devices comprising a series of elements not only include those elements, but also include other elements not explicitly listed, or further include the elements inherent to such processes, methods, articles or devices. The terms "first", "second" and the like in the present application are used to distinguish similar objects, not to describe a specific order or sequence.
[0032] In order to make the skilled in the art better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0033] The embodiments of the present application provide a data writing method based on a distributed storage system, as shown in the figure. Figure 2 and Figure 3As shown, the distributed storage system is applied to the distributed storage system in the embodiment, which is configured as Ceph. Ceph is an open source distributed file system, which can support block storage, file storage and object storage and can be applied to an open source cloud computing platform, such as Figure 2 As shown, the method comprises:
[0034] Step 1, receiving the to-be-written data, and identifying that the to-be-written data comprises a data object.
[0035] Specifically, the to-be-written data from the storage disk in the distributed storage system is received, and the to-be-written data is decoded to obtain decoded information. If the decoded information comprises a preset keyword, it is confirmed that the to-be-written data comprises a data object, and the data object is block storage data.
[0036] In the Ceph in the embodiment, an OSD and a plurality of storage disks for storing data are configured. The OSD refers to a process for managing the storage disks in the Ceph. One storage disk is configured with one OSD. The to-be-written data is received, information from the OSD is decoded to obtain decoded information, and the object type is confirmed according to the decoded information. For example, if the decoded information comprises a keyword “rbd-data”, the written object is RBD interface data, so that it is confirmed that the data object is block storage data.
[0037] Step 2, in response to confirming that the data object is first-time written data, an initial storage space is allocated for the data object.
[0038] Specifically, in response to confirming that the data object is first-time written data, the initial storage space is configured as a fixed capacity, and the fixed capacity is set as a minimum allocation unit. In response to performing the first-time written data operation, a logical container is created, and the first-time written data object is written into the logical container.
[0039] When it is confirmed that the data object is first-time written data, the OSD is used to allocate a corresponding initial storage space for the data object. The capacity of the initial storage space is configured by a configuration file when the distributed storage system is initialized, and the capacity of the initial storage space is fixedly set as a minimum allocation unit of the distributed storage system. Through the above setting, space management can be simplified, and other storage spaces are allocated according to the capacity of the above minimum allocation unit in subsequent space allocation.
[0040] Wherein, the Ceph includes Blob, Extent and Pextent, wherein the Blob is a logical container, the Extent is a logical segment, and the Pextent is a physical block, and a plurality of physical blocks can constitute a logical container; the logical container is used for managing the physical mapping of a plurality of Extents, the logical segment identifies a continuous data segment, that is, a data segment starting from a certain offset, and the logical segment can be mapped to a plurality of discontinuous physical blocks in the underlying.
[0041] In the execution of the write operation of the data object, two write modes are included, that is, COW write and RMW write, in the embodiment, when the data object is written for the first time, the pre-allocated fixed-size storage space is used, that is, the COW new write performed in the embodiment, and it is necessary to ensure that the data size of the COW new write is smaller than the initial storage space size, so as to ensure that the COW new write data can be written into the initial storage space.
[0042] Step 3, configure the initial storage space to be associated with the logical segment in the distributed storage system to form a mapping relationship.
[0043] Wherein, the logical container includes one or more physical storage blocks, the logical segment is configured to be mapped to the physical storage block, and the physical storage block is set as a storage location; the logical segment is set to have a first identifier, and the physical storage block is set to have a second identifier; the logical segment is configured to be correspondingly associated with the physical storage block, and the first identifier and the second identifier correspond to each other. Through the above setting, the initial storage space is correspondingly associated with the physical storage location in the distributed storage system.
[0044] After the logical container is created, the logical container is updated to the mapping relationship, so that the updated mapping relationship includes the newly created logical container, for subsequent data writing.
[0045] Step 4, in response to performing a non-first write operation based on the initial storage space, the storage location of the logical segment associated with the initial storage space is queried in combination with the mapping relationship, and the non-first write operation data is written to the storage location.
[0046] After the storage location of the logical segment associated with the initial storage space is queried in combination with the above mapping relationship, the reference flag in the logical segment on the storage location is read, and if the reference flag is greater than or equal to 1, it is confirmed that the storage location is reusable; the non-first write operation data is written to the storage location, and the physical storage block is not changed, and through the above method, it is confirmed that the storage location associated with the initial storage space is reusable, at this time, the logical container associated with the initial storage space is found and obtained, and the physical storage block associated with the logical container is obtained, and the data of the non-first write operation is written to the physical storage block, thereby realizing efficient use of the storage space without the need to perform redundant space allocation and recovery again.
[0047] Exemplarily, when performing the COW overwrite write in the embodiment, a logical container is queried and a reference flag value of the logical container is acquired. When the reference flag value of the logical container is 1, it is confirmed that the storage location associated with the initial storage space is reusable.
[0048] In a specific embodiment, after it is confirmed that the storage location is reusable, the non-first-write data is written into the storage location, and at this time, the physical storage block position of the storage location is not changed. Before the non-first-write data is written into the storage location, the following steps are further included: querying whether the physical storage block associated with the storage location is available, if available, setting the physical storage block not to be changed; if the physical storage block is not available, querying the physical storage block associated with the next number associated with the number of the unavailable physical storage block, and setting the physical storage block associated with the next number as the storage block associated with the storage location, so as to ensure that the storage location can realize smooth writing of data.
[0049] Exemplarily, the number of the current physical storage block associated with the storage location is set to “001”, the physical storage block associated with the next number associated with the current physical storage block is “002”, and when the current physical storage block is not available, the physical storage block associated with the next number is set as the physical storage block associated with the storage location.
[0050] In a specific embodiment, before the non-first-write operation method of the initial storage space is executed in response, it is judged whether the length of the non-first-write operation data is less than the minimum allocation unit; if the length of the non-first-write operation data is less than the minimum allocation unit, a target data block with a capacity of the minimum allocation unit size is read; and the non-first-write operation data is merged into the target data block.
[0051] That is, in the embodiment, first, the length of the non-first-write operation data is judged, if the length is less than the minimum allocation unit, RWM write is performed, at this time, a target data block with a minimum allocation unit size is read, and the non-first-write operation data is written into the target data block.
[0052] The non-first-write operation data is merged into the target data block, specifically including: reading the existing data in the target data block, replacing the to-be-modified part in the existing data with the modified data in combination with the non-first-write operation data; taking the modified data and the remaining data in the target data block as the modified target data block data; and synchronizing the modified target data block data to the storage section.
[0053] Exemplarily, when performing the RWM write, the target data block is read, the data is merged in the storage system, the existing data in the target data block is replaced, the remaining content is unchanged, the replaced data is directly written back to the original storage section, and the physical storage block associated with the storage section is set not to be changed, so as to reduce the write delay.
[0054] By identifying the data object when data is written in the present application, it is then judged whether the data object is first-time written data, if so, an initial storage space is allocated for it, and the initial storage space is associated with the distributed storage system to form a mapping relationship. When a re-write operation is performed on the initial storage space, the mapping relationship is queried to obtain the storage location associated with the initial storage space, and the data of the non-first-time write operation is directly written into the storage location, solving the problem of additional metadata generated in the encapsulation and mapping process, and avoiding the problem of data fragmentation. At the same time, the present application can write non-first-time data in the original storage location, so that subsequent data writing does not need to be allocated and recycled multiple times, improving the data writing efficiency of the distributed storage system.
[0055] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment.
[0056] The embodiments of the present application also provide a data writing device based on a distributed storage system, as shown in Figure 4 The device comprises:
[0057] An identification module is configured to receive to-be-written data and identify that the to-be-written data comprises a data object.
[0058] An allocation module is configured to allocate an initial storage space for the data object in response to confirming that the data object is first-time written data.
[0059] A configuration module is configured to configure the initial storage space to be associated with a storage segment in the distributed storage system to form a mapping relationship.
[0060] An execution module is configured to write data of a non-first-time write operation into a storage location of a storage segment associated with the initial storage space in combination with the mapping relationship in response to performing the non-first-time write operation based on the initial storage space.
[0061] In one specific embodiment, the identification module is further configured to receive to-be-written data from a storage disk in the distributed storage system, decode the to-be-written data to obtain decoded information, and confirm that the to-be-written data comprises a data object and the data object is block storage data if the decoded information comprises a preset keyword.
[0062] In one specific embodiment, the configuration module is further configured to, in response to confirming that the data object is first write data, configure the initial storage space as a fixed capacity, and set the fixed capacity as the minimum allocation unit; in response to performing the first write data operation, create a logical container, and update the logical container to the mapping relationship; and perform a data object write logical container operation of the first write data.
[0063] In one specific embodiment, the logical container includes one or more physical storage blocks, and the configuration module is further configured to configure the logical section to be mapped to the physical storage block, set the physical storage block as the storage location; set the logical section to have a first identifier, and set the physical storage block to have a second identifier; configure the logical section to be associated with the physical storage block in a one-to-one correspondence, and the first identifier and the second identifier are in a one-to-one correspondence.
[0064] In one specific embodiment, the execution module is further configured to determine whether the length of the non-first write operation data is less than the minimum allocation unit; if the length of the non-first write operation data is less than the minimum allocation unit, read a target data block with a capacity of the minimum allocation unit; and merge the non-first write operation data into the target data block.
[0065] In one specific embodiment, the execution module is further configured to read existing data in the target data block, replace a to-be-modified part in the existing data with modified data in combination with the non-first write operation data; take the modified data and the remaining data in the target data block as modified target data block data; and synchronize the modified target data block data to the storage section.
[0066] In one specific embodiment, the execution module is further configured to read a reference flag in the logical section on the storage location, and if the reference flag is a value greater than or equal to 1, confirm that the storage location is reusable; write the non-first write operation data to the storage location, and set the physical storage block not to be changed.
[0067] The description of the features of the embodiments of the data writing device based on the distributed storage system can be referred to the related description of the embodiments of the data writing method based on the distributed storage system, which will not be repeated here.
[0068] Embodiments of the present application also provide an electronic device including a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps of any one of the embodiments of the data writing method based on the distributed storage system.
[0069] In one specific embodiment, the steps of the embodiments of the data writing method based on the distributed storage system specifically include:
[0070] Step 101, receiving to-be-written data, and identifying that the to-be-written data includes a data object;
[0071] Step 102, in response to confirming that the data object is first write data, allocating initial storage space for the data object;
[0072] Step 103, configuring the initial storage space to be associated with a logical segment in the distributed storage system to form a mapping relationship;
[0073] Step 104, in response to performing a non-first write operation based on the initial storage space, querying the storage location of the logical segment associated with the initial storage space in combination with the mapping relationship, and writing the non-first write operation data to the storage location.
[0074] In one embodiment, the processor, when executing the computer program, also implements the following steps: receiving to-be-written data from a storage disk in the distributed storage system, decoding the to-be-written data to obtain decoded information; if the decoded information includes a preset keyword, it is confirmed that the to-be-written data includes a data object, and the data object is block storage data.
[0075] In one embodiment, the processor, when executing the computer program, also implements the following steps: in response to confirming that the data object is first write data, configuring the initial storage space to have a fixed capacity, and setting the fixed capacity as a minimum allocation unit; in response to performing a first write data operation, creating a logical container and updating the logical container to the mapping relationship; and performing a data object write operation of the first write data to the logical container.
[0076] In one embodiment, the processor, when executing the computer program, also implements the following steps: configuring the logical segment to be mapped to a physical storage block, setting the physical storage block as the storage location; setting the logical segment to have a first identifier and setting the physical storage block to have a second identifier; configuring the logical segment to be correspondingly associated with the physical storage block, and the first identifier and the second identifier one-to-one corresponding.
[0077] In one embodiment, the processor, when executing the computer program, also implements the following steps: determining whether the length of the non-first write operation data is less than the minimum allocation unit; if the length of the non-first write operation data is less than the minimum allocation unit, reading a target data block having a capacity of the minimum allocation unit; and merging the non-first write operation data into the target data block.
[0078] In one embodiment, the processor, when executing the computer program, also implements the following steps: reading existing data in the target data block, replacing a to-be-modified part in the existing data with modified data in combination with the non-first write operation data; taking the modified data and the remaining data in the target data block as modified target data block data; and synchronizing the modified target data block data to the storage segment.
[0079] In one embodiment, the processor, when executing the computer program, also implements the following steps: reading the reference flag in the logical section on the storage location, confirming that the storage location is reusable if the reference flag is a value greater than or equal to 1; and writing the non-first write operation data to the storage location and setting the physical storage block not to be changed.
[0080] Embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above data write methods based on a distributed storage system when running.
[0081] In one specific embodiment, the steps in the data write method based on a distributed storage system specifically include:
[0082] Step 201: receiving data to be written and identifying that the data to be written includes a data object;
[0083] Step 202: in response to confirming that the data object is first write data, allocating an initial storage space for the data object;
[0084] Step 203: configuring the initial storage space to be associated with a logical section in the distributed storage system to form a mapping relationship;
[0085] Step 204: in response to performing a non-first write operation based on the initial storage space, querying a storage location of the logical section associated with the initial storage space in combination with the mapping relationship, and writing non-first write operation data to the storage location.
[0086] In one embodiment, the computer program, when executed by the processor, also implements the following steps: receiving data to be written from a storage disk in the distributed storage system, decoding the data to be written to obtain decoded information, and confirming that the data to be written includes a data object and the data object is block storage data if the decoded information includes a preset keyword.
[0087] In one embodiment, the computer program, when executed by the processor, also implements the following steps: in response to confirming that the data object is first write data, configuring the initial storage space to have a fixed capacity and setting the fixed capacity as a minimum allocation unit; in response to performing a first write data operation, creating a logical container and updating the logical container to the mapping relationship; and performing a data object write logical container operation of the first write data.
[0088] In one embodiment, the computer program, when executed by the processor, also implements the following steps: configuring the logical section to be mapped to a physical storage block and setting the physical storage block as a storage location; setting the logical section to have a first identifier and setting the physical storage block to have a second identifier; configuring the logical section and the physical storage block to be correspondingly associated, and the first identifier and the second identifier to be in one-to-one correspondence.
[0089] In one embodiment, the computer program, when executed by the processor, further implements the following steps: determining whether the length of the non-first write operation data is less than the minimum allocation unit; if the length of the non-first write operation data is less than the minimum allocation unit, reading the target data block with a capacity of the minimum allocation unit size; and merging the non-first write operation data into the target data block.
[0090] In one embodiment, the computer program, when executed by the processor, further implements the following steps: reading the existing data in the target data block, replacing the to-be-modified part in the existing data with the modified data in combination with the non-first write operation data; taking the modified data and the remaining data in the target data block as the modified target data block data; and synchronizing the modified target data block data to the storage section.
[0091] In one embodiment, the computer program, when executed by the processor, further implements the following steps: reading the reference flag in the logical section on the storage location, and if the reference flag is a value greater than or equal to 1, confirming that the storage location is reusable; writing the non-first write operation data to the storage location, and setting the physical storage block unchanged.
[0092] In one example embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0093] Embodiments of the present application also provide a computer program product, which includes a computer program, and the computer program, when executed by a processor, implements the steps in any of the above data write methods based on a distributed storage system.
[0094] Embodiments of the present application also provide another computer program product, which includes a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps in any of the above data write methods based on a distributed storage system.
[0095] Those skilled in the art will further realize that the mere concepts, teachings, and embodiments described herein are merely meant to provide examples of the various aspects of the present application and that various modifications, equivalents and alternatives are intended to fall within the scope of the present application. Accordingly, the appended claims are intended to embrace all such alterations, modifications, and improvements as fall within the scope of the present application. As can be seen, the application provides a novel and improved method and apparatus for data write based on distributed storage system.
[0096] The above provides a detailed introduction to the data write method, device, equipment and medium based on the distributed storage system. The principle and implementation of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method and its core idea of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the principle of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A data writing method based on a distributed storage system, characterized in that: The method comprises: receiving data to be written, and identifying that the data to be written includes a data object; In response to confirming that the data object is data written for the first time, allocating initial storage space for the data object; Configuring the initial storage space to be associated with the logical segments in the distributed storage system to form a mapping relationship; In response to executing a non-first write operation based on the initial storage space, the storage location of the logical segment associated with the initial storage space is queried in combination with the mapping relationship, and the non-first write operation data is written into the storage location.
2. The data writing method based on the distributed storage system according to claim 1, characterized in that: Receiving data to be written and identifying that the data to be written includes a data object, specifically including: receiving the data to be written from the storage disk in the distributed storage system, and decoding the data to be written to obtain decoded information; If the decoded information includes the preset keyword, it is confirmed that the data to be written includes the data object, and the data object is block storage data.
3. The data writing method based on a distributed storage system according to claim 1 or 2, characterized in that: The method further comprises: In response to confirming that the data object is written for the first time, configuring the initial storage space to have a fixed capacity, and setting the fixed capacity as a minimum allocation unit; In response to executing the first data writing operation, a logical container is created, and the mapping relationship between the logical container and the logical container is updated; and an operation of writing the data object of the first data writing operation into the logical container is executed.
4. The data writing method based on the distributed storage system according to claim 3, characterized in that: The logical container includes one or more physical storage blocks, and the method further includes: Configuring the logical segment to be mapped to the physical storage block, and setting the physical storage block as the storage location; Setting the logical segment to have a first identifier and setting the physical storage block to have a second identifier; The logical segments are configured to be associated with the physical storage blocks, and the first identifiers are in one-to-one correspondence with the second identifiers.
5. The data writing method based on a distributed storage system according to claim 1 or 2, characterized in that: Before responding to executing the non-first write operation to the initial storage space, the method further includes: Determining whether the length of the non-first write operation data is less than the minimum allocation unit; If the length of the non-first write operation data is smaller than the minimum allocation unit, then reading a target data block with a capacity equal to the size of the minimum allocation unit; And the non-first write operation data is merged into the target data block.
6. The data writing method based on the distributed storage system according to claim 5, characterized in that: Merging the non-first write operation data into the target data block specifically includes: Reading existing data in the target data block, and replacing the to-be-modified portion of the existing data with the revised data in combination with the non-first write operation data; Using the corrected data and the remaining data in the target data block as corrected target data block data; Synchronize the modified target data block data to the storage segment.
7. The data writing method based on the distributed storage system according to claim 4, characterized in that: After querying the storage location of the logical segment associated with the initial storage space in combination with the mapping relationship, the method further includes: reading a reference flag in the logical segment at the storage location, and if the reference flag is a value greater than or equal to 1, confirming that the storage location is reusable; The non-first write operation data is written into the storage location, and the physical storage block is set to remain unchanged.
8. A data writing device based on a distributed storage system for implementing the method according to any one of claims 1 to 7, characterized in that: The device comprises: an identification module, configured to receive data to be written and identify that the data to be written includes a data object; an allocation module, configured to allocate initial storage space for the data object in response to confirming that the data object is data written for the first time; A configuration module, configured to configure the initial storage space to be associated with the storage segments in the distributed storage system to form a mapping relationship; An execution module is used to query the storage location of the storage segment associated with the initial storage space in response to executing a non-first write operation based on the initial storage space in combination with the mapping relationship, and write the non-first write operation data into the storage location.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data writing method based on a distributed storage system according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the data writing method based on the distributed storage system according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Universal large-size data storage method and system
CN121657941A