Unstructured protocol data mutual access method and device, electronic equipment and medium
By sharding unstructured data and creating object indexes, the problem of mutual access and interoperability between unstructured storage protocols is solved, efficient and stable data interaction is achieved, and data processing efficiency and system reliability are improved.
Patent Information
- Application Number
- CN202510396987.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The existing unstructured storage protocols cannot achieve efficient and stable mutual access and interoperability, resulting in limited flexibility and efficiency of data storage and interaction, and traditional solutions have problems such as high resource consumption and long-term consumption.
By sharding the unstructured data based on the first data storage protocol, a shard file is generated, and an object index is created for each shard file, including storage location and index identification information, the storage location of the target shard data is quickly positioned using the object index to realize data access under different protocols.
It realizes smooth interaction between storage systems under different protocols, reduces data search time, improves data reading speed and processing efficiency, reduces resource consumption, and improves system stability and scalability.
Smart Images

Figure CN120336262A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of storage technologies, and in particular, to a method and apparatus for mutual access of unstructured protocol data, an electronic device, and a medium. Background Art
[0002] The Simple Storage Service (S3) has a function of uploading in chunks to solve the problem of uploading large files, and can achieve resume from breakpoint and improve the upload efficiency. However, the Hadoop Distributed File System Protocol (HDFS) and the Network-Attached Storage Protocol (NAS) do not support this function. This results in the failure of reading and writing operations on files uploaded in chunks through the S3 protocol in the HDFS protocol or NAS protocol environment, making it impossible to truly achieve mutual access and interconnection among the three protocols, and restricting the flexibility and efficiency of data storage and interaction. In related technologies, usually at the completion stage of S3 chunked upload, the chunked file data is relocated and merged. However, this approach has problems such as high resource consumption and long time consumption. Therefore, how to achieve efficient and stable mutual access and interconnection among unstructured storage protocols has become an urgent problem to be solved. Summary of the Invention
[0003] The present disclosure provides a method and apparatus for mutual access of unstructured protocol data, an electronic device, and a medium. Its main purpose is to solve the problem that mutual access and interconnection among unstructured storage protocols cannot be efficient and stable.
[0004] According to a first aspect of the present disclosure, there is provided a method for mutual access of unstructured protocol data, including:
[0005] Performing chunking processing on unstructured data based on a first data storage protocol to generate at least one chunked file;
[0006] Creating an object index for each chunked file, where the object index includes the chunk layout information of the chunked file, and the chunk layout information includes at least the storage location information of the allocated file and the index identification information;
[0007] In response to a read instruction for target chunked data issued based on a second data storage protocol, querying the object index of the chunked file to determine the storage location information of the target chunked data;
[0008] Reading the target chunked data based on the storage location information.
[0009] Optionally, creating an object index for each chunked file includes:
[0010] Generate an object index corresponding to each shard file based on the storage location information and index identification information of each shard file. The storage location information includes: offset, data length, and storage path. The index identification information includes an index number;
[0011] Arrange the object indexes corresponding to at least one shard file in the order of the index numbers in the object index.
[0012] Optionally, after creating an object index for each shard file, the method further includes:
[0013] Associate and store at least one shard file and the object index to generate a large shard file, and store the shard identifier of the large shard file in the object index.
[0014] Optionally, before querying the object index of the shard file to determine the storage location information of the target shard data in response to a read instruction for the target shard data issued based on the second data storage protocol, the method further includes:
[0015] Parse the access request of the second data storage protocol to determine the target data length and target offset of the target shard data to be accessed;
[0016] Generate a read instruction for reading the target shard data based on the target data length and target offset;
[0017] Determine whether the target file pointed to by the read instruction is a large shard file generated according to the first data storage protocol; if the target file is a large shard file, find the storage location information of the target shard data from the object index.
[0018] Optionally, determining whether the target file pointed to by the read instruction is a large shard file generated according to the first data storage protocol includes:
[0019] Read the metadata of the target file and determine whether there is a shard identifier in the metadata.
[0020] Optionally, in response to a read instruction for the target shard data issued based on the second data storage protocol, querying the object index of the shard file to determine the storage location information of the target shard data includes:
[0021] Based on the read instruction, read the object index of the shard file to determine the target shard range where the target shard data is located;
[0022] Obtain the target storage path of each target shard file that overlaps with the target shard range, and calculate the target read range corresponding to each target shard file;
[0023] According to the target read range and the shard layout information of the target shard data, find the storage location information of the target shard data.
[0024] Optionally, based on the read instruction, read the object index of the shard file, and determine the target shard range where the target shard data is located, including:
[0025] Parse the read instruction to obtain the target data length and target offset of the target shard data;
[0026] Traverse each object index, determine the object index that overlaps with the target data length and target offset, and generate a target shard range based on the overlapping object index.
[0027] Optionally, calculate the target read range corresponding to each target shard file, including:
[0028] Based on the target offset of the read instruction, the offset and data length of the starting target shard file within the target shard range, calculate the target read range of the starting target shard file;
[0029] According to the target offset, target read length of the read instruction, and the offset of the ending target shard file within the target shard range, calculate the target read range of the ending target shard file.
[0030] Optionally, the method further includes:
[0031] If the target shard file is an overlapping target shard file that completely overlaps with the data read range, then use the data range of the overlapping target shard file as the target read range.
[0032] Optionally, perform sharding on unstructured data based on the first data storage protocol to generate at least one shard file, including:
[0033] Divide the unstructured data into at least one shard data according to the first data storage protocol;
[0034] Create at least one shard file, store the at least one shard data in the at least one shard file, and save one shard data in one shard file.
[0035] Optionally, divide the unstructured data into shards according to the first data storage protocol, including:
[0036] Divide the unstructured data into at least one shard data according to the shard size threshold;
[0037] Assign sequence numbers to the at least one shard data, and assign a unique shard sequence number to each shard data.
[0038] Optionally, before dividing the unstructured data into shards according to the first data storage protocol, the method further includes:
[0039] Send a sharded upload request for unstructured data to the storage system and generate a unique upload identifier according to the sharded upload semantics of the first data storage protocol.
[0040] Optionally, create at least one sharded file and store at least one sharded data in at least one sharded file, including:
[0041] Create at least one sharded file and its corresponding storage path according to the shard sequence number of at least one sharded data and the upload identifier;
[0042] Store at least one sharded data in its corresponding at least one sharded file according to the storage path respectively, and record the offset and data length of each sharded file respectively.
[0043] Optionally, the method further includes:
[0044] In response to a modification instruction for target sharded data issued based on the second data storage protocol, query the object index of the sharded file to determine the storage location information of the target sharded data;
[0045] Based on the storage location information, read the target sharded data from at least one sharded file for modification, and modify the object index according to the modification information.
[0046] Optionally, modifying the object index according to the modification information includes:
[0047] Obtain the modified offset and modified data length of the modified sharded file;
[0048] Modify the object index based on the modified offset and modified data length.
[0049] Optionally, the method further includes:
[0050] In response to a deletion instruction for target sharded data issued based on the second data storage protocol, query the object index of the sharded file to determine the storage location information of the target sharded data;
[0051] Based on the storage location information, read the target sharded data from at least one sharded file for deletion.
[0052] According to the second aspect of the present disclosure, there is provided a device for mutual access of unstructured protocol data, including:
[0053] A first generation unit for sharding unstructured data based on the first data storage protocol to generate at least one sharded file;
[0054] A creation unit, configured to create an object index for each sharded file, where the object index includes the sharding layout information of the sharded file, and the sharding layout information at least includes the storage location information and index identification information of the allocated file;
[0055] A first query unit, configured to query the object index of the sharded file to determine the storage location information of the target sharded data in response to a read instruction for the target sharded data issued based on the second data storage protocol;
[0056] A read unit, configured to read the target sharded data based on the storage location information.
[0057] Optionally, the creation unit includes:
[0058] A first generation module, configured to generate an object index corresponding to each sharded file based on the storage location information and index identification information of each sharded file, where the storage location information includes: offset, data length, and storage path, and the index identification information includes an index number;
[0059] An arrangement module, configured to arrange the object indexes corresponding to at least one sharded file in the order of the index numbers in the object index.
[0060] Optionally, the apparatus further includes:
[0061] A second generation unit, configured to, after creating an object index for each sharded file, associatively store at least one sharded file and the object index to generate a sharded large file, and store the sharding identifier of the sharded large file to the object index.
[0062] Optionally, the apparatus further includes:
[0063] A determination unit, configured to parse the access request of the second data storage protocol to determine the target data length and target offset of the target sharded data to be accessed before querying the object index of the sharded file to determine the storage location information of the target sharded data in response to a read instruction for the target sharded data issued based on the second data storage protocol;
[0064] A third generation unit, configured to generate a read instruction for reading the target sharded data based on the target data length and target offset;
[0065] A judgment unit, configured to judge whether the target file pointed to by the read instruction is a sharded large file generated according to the first data storage protocol; if the target file is a sharded large file, then find the storage location information of the target sharded data from the object index.
[0066] Optionally, the judgment unit is further configured to:
[0067] Read the metadata of the target file and judge whether there is a sharding identifier in the metadata.
[0068] Optionally, the first query unit includes:
[0069] A determination module, configured to read the object index of the shard file based on the read instruction, and determine the target shard range where the target shard data is located;
[0070] A calculation module, configured to obtain the target storage path of each target shard file that overlaps with the target shard range, and calculate the target read range corresponding to each target shard file;
[0071] A search module, configured to search for the storage location information of the target shard data according to the target read range and the shard layout information of the target shard data.
[0072] Optionally, the determination module is further configured to:
[0073] Parse the read instruction to obtain the target data length and target offset of the target shard data;
[0074] Traverse each object index, determine the object index that overlaps with the target data length and target offset, and generate a target shard range based on the overlapping object index.
[0075] Optionally, the calculation module is further configured to:
[0076] Based on the target offset of the read instruction, the offset and data length of the starting target shard file within the target shard range, calculate the target read range of the starting target shard file;
[0077] According to the target offset, target read length of the read instruction, and the offset of the terminating target shard file within the target shard range, calculate the target read range of the terminating target shard file.
[0078] Optionally, the apparatus is further configured to:
[0079] If the target shard file is an overlapping target shard file that completely overlaps with the data read range, use the data range of the overlapping target shard file as the target read range.
[0080] Optionally, the first generation unit includes:
[0081] A second generation module, configured to divide the unstructured data into shards according to the first data storage protocol, and generate at least one shard of data;
[0082] A storage module, configured to create at least one shard file, store at least one shard of data into at least one shard file, and save one shard of data into one shard file.
[0083] Optionally, the second generation module is further configured to:
[0084] Divide the unstructured data into at least one piece of data according to the piece size threshold;
[0085] Assign sequence numbers to at least one piece of data, and assign a unique piece sequence number to each piece of data.
[0086] Optionally, the device further includes:
[0087] A sending unit, configured to, before dividing the unstructured data into pieces according to the first data storage protocol, send a piece upload request for the unstructured data to the storage system and generate a unique upload identifier according to the piece upload semantics of the first data storage protocol.
[0088] Optionally, the storage module is further configured to:
[0089] Create at least one piece file and its corresponding storage path according to the piece sequence number and upload identifier of at least one piece of data;
[0090] Store at least one piece of data into its corresponding at least one piece file respectively according to the storage path, and record the offset and data length of each piece file respectively.
[0091] Optionally, the device further includes:
[0092] A second query unit, configured to, in response to a modification instruction for target piece data issued based on the second data storage protocol, query the object index of the piece file to determine the storage location information of the target piece data;
[0093] A modification unit, configured to, based on the storage location information, read the target piece data from at least one piece file for modification, and modify the object index according to the modification information.
[0094] Optionally, the modification unit includes:
[0095] An acquisition module, configured to acquire the modification offset and modification data length of the modified piece file;
[0096] A modification module, configured to modify the object index based on the modification offset and modification data length.
[0097] Optionally, the device further includes:
[0098] A third query unit, configured to, in response to a deletion instruction for target piece data issued based on the second data storage protocol, query the object index of the piece file to determine the storage location information of the target piece data;
[0099] A deletion unit, configured to, based on the storage location information, read the target piece data from at least one piece file for deletion.
[0100] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0101] at least one processor; and
[0102] a memory communicatively connected to the at least one processor; wherein,
[0103] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for unstructured protocol data mutual access described in the foregoing first aspect.
[0104] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method for unstructured protocol data mutual access described in the foregoing first aspect.
[0105] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method for unstructured protocol data mutual access described in the foregoing first aspect.
[0106] The present disclosure provides a method and apparatus for unstructured protocol data mutual access, an electronic device, and a medium, relating to the field of storage technologies. In the present disclosure, aiming at the problem of mutual access difficulties caused by semantic differences between unstructured storage protocols, data is processed in slices by means of a first data storage protocol, and data is read in combination with a second data storage protocol. Effective access to sliced data under different protocols is achieved through object indexing and slice layout information. This enables data to smoothly interact between storage systems of multiple protocols, breaking down protocol barriers. An object index containing key information such as storage location and index identifier is created for each sliced file. When a read instruction is received, the system can quickly locate the storage location of the target sliced data based on the index, without the need to comprehensively retrieve a large amount of data, reducing the data search time. Especially when dealing with massive unstructured data, this index-based fast positioning mechanism significantly improves the data reading speed, speeds up system response, and improves data processing efficiency. Compared with the traditional solution that requires merging large files after sliced uploading, it is possible to avoid resource consumption caused by large-scale data relocation and merging. The amount of data processing is reduced, the requirements for the I / O performance of storage devices and network bandwidth are lowered, the overall system load is alleviated, which helps to maintain the stable operation of the storage system and improve the reliability and scalability of the system. When dealing with large files, the required data can be obtained without waiting for the file merging operation to complete, which can improve convenience and fluency.
[0107] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0108] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0109] Figure 1 is a schematic flowchart of a method for mutual access of unstructured protocol data provided by an embodiment of the present disclosure;
[0110] Figure 2 is a schematic flowchart of another method for mutual access of unstructured protocol data provided by an embodiment of the present disclosure;
[0111] Figure 3 is a schematic structural diagram of a device for mutual access of unstructured protocol data provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0112] The following describes exemplary embodiments of the present disclosure with reference to the drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0113] The following describes a method, device, electronic device, and medium for mutual access of unstructured protocol data according to embodiments of the present disclosure with reference to the drawings.
[0114] Figure 1 is a schematic flowchart of a method for mutual access of unstructured protocol data provided by an embodiment of the present disclosure.
[0115] As Figure 1 shown, the method includes the following steps:
[0116] Step 101, fragment the unstructured data based on the first data storage protocol to generate at least one fragmented file.
[0117] In an embodiment of the present disclosure, step 101 aims to implement sharding processing of unstructured data based on the first data storage protocol, and then generate at least one sharded file. The "first data storage protocol" here can be the S3 object storage protocol, which has semantic features supporting sharded upload and can effectively solve the problems of resume upload and upload efficiency during the transmission of large files; it can also be other protocols that can achieve sharded storage, and the embodiments of the present disclosure do not limit this. "Unstructured data" refers to data with irregular or incomplete data structures, without a predefined data model, and is inconvenient to be represented by a database two-dimensional logical table, such as documents, pictures, videos, etc.
[0118] When large unstructured data needs to be stored and uploaded, according to the rules of the first data storage protocol, the data is split according to a specific sharding strategy. This sharding strategy usually considers factors such as the size of the data volume, network transmission conditions, and the performance of the storage system. For example, it can be sharded according to a preset fixed size, and a large video file can be split into multiple appropriately sized segments, and each segment is a sharded file.
[0119] After the data sharding is completed, a corresponding storage carrier is created for each sharded file, that is, at least one sharded file is generated. Each sharded file is the smallest unit for independently storing sharded data, and they together constitute the storage set of the original unstructured data. During the generation process, each sharded file is assigned specific identification information for subsequent identification, positioning, and management. For example, in the S3 protocol, the sharded file is associated with corresponding uploadid and other information for tracking and operation at each stage of sharded upload.
[0120] Step 102, create an object index for each sharded file, and the object index includes the sharding layout information of the sharded file, and the sharding layout information includes at least the storage location information of the allocated file and the index identification information.
[0121] In an embodiment of the present disclosure, step 102 focuses on creating an object index for each sharded file, and this operation aims to achieve efficient storage of unstructured data and multi-protocol interoperability. The "object index" is a key data structure for quickly locating and accessing sharded files. It is similar to the bibliographic index in a library, and through specific identification and description information, it helps the system quickly find the required sharded file. The object index is created for each sharded file, integrating the key information of the file to facilitate subsequent operations.
[0122] "Fragment layout information" is a core component of object indexing, which details important information such as the distribution and identity of fragment files in the storage system. Among them, "the storage location information of the allocated file" specifies the specific storage address of each fragment file in the storage medium. Taking a distributed storage system as an example, this can include, but is not limited to, the IP address of the storage node, the directory path where the file is located, etc., just like labeling the specific shelf location of each item in a large warehouse to ensure that the system can accurately find the corresponding fragment file. "Index identification information" is the unique identity of each fragment file. It can be a hash value calculated based on the file content, a sequential number generated according to certain rules, or other unique identifiers. This information is not only used to distinguish different fragment files but also helps the system quickly identify and match the required fragment files during multi-protocol interactions, ensuring the accuracy and efficiency of data operations.
[0123] In the actual process of creating an object index, the system will automatically collect and organize this fragment layout information after generating each fragment file. Taking fragment upload based on the S3 protocol as an example, during the fragment upload process, the system will record in real-time the storage location information of each fragment file, such as which bucket the fragment is stored in, the specific object path, etc., and generate a unique index identification information for it. Then, these information are integrated to form an object index and stored in a dedicated index database or metadata management module for subsequent quick query and use.
[0124] Step 103, in response to a read instruction for target fragment data issued based on the second data storage protocol, query the object index of the fragment file to determine the storage location information of the target fragment data.
[0125] In the embodiments of the present disclosure, for the target fragment data read instruction in the second data storage protocol environment, the present disclosure determines the storage location information of the target fragment data by means of an object index.
[0126] The "second data storage protocol" is a data storage protocol that does not support fragment upload or storage. For example, it can be a protocol such as the HDFS protocol or the NAS file protocol, which do not directly support the semantics of fragment upload but are widely used in the field of data storage. "Target fragment data" refers to a specific part of fragment data that a user or application expects to obtain, which is a subset of the overall unstructured data after fragment processing. The "read instruction" is an operation request sent by a user or application to the storage system to obtain the target fragment data, and this instruction contains identification information of the target data, such as key information like file name, fragment identifier, etc., so that the storage system can accurately locate the target fragment data to be read.
[0127] After the system receives a read instruction issued based on the second data storage protocol, a series of complex but orderly operations will be triggered. First, the system will parse the read instruction and extract the key identification information for locating the target shard data from the instruction. This information can include, but is not limited to, the file name, shard number, or other identifiers related to the target shard data. Subsequently, the system will perform a query operation in the object index of the shard file. The "object index" is a data structure created in step 102 for quickly locating shard files. It contains the shard layout information of each shard file, including the storage location information we need. This query process is similar to retrieving data in a database. The system matches the identification information in the read instruction with the data in the object index to find the object index entry corresponding to the target shard data. Once the corresponding object index entry is found, the system can obtain the storage location information of the target shard data from this entry. These storage location information details record the specific storage location of the target shard data in the storage system, which may include key information such as the physical address of the storage device and the logical storage path.
[0128] Step 104, read the target shard data based on the storage location information.
[0129] In the embodiments of the present disclosure, the "storage location information" is the key data determined by querying the object index of the shard file in the previous steps. It precisely indicates the storage location of the target shard data in the storage system. This information can include, but is not limited to, the specific identification of the storage device (such as the IP address of the storage server, the storage node number, etc.), the directory path in the file system (similar to the specific path identifier in the file tree structure for locating the folder containing the target shard data), and detailed information such as the offset and data length within the storage medium. These contents combined are like a detailed address guide, enabling the system to accurately find the physical or logical location of the target shard data in the storage system.
[0130] The "target shard data" refers to the specific data segment that a user or application expects to obtain. After unstructured data is sharded, a large number of shard data are formed, and the target shard data is the part selected for the read operation. It is a subset of the original unstructured data and contains the specific information required by the user.
[0131] After the system obtains the storage location information of the target shard data, it will initiate the read operation. First, the system will establish a connection with the corresponding storage device according to the storage device identifier in the storage location information. In a distributed storage environment, it may be necessary to communicate with the storage server through a network protocol (such as the TCP / IP protocol) to ensure access to the node storing the target shard data. Then, based on the directory path information, the system navigates to the specific location containing the target shard data in the file system of the storage device. This process is similar to finding a file by path in a local file system, except that in a distributed or complex storage environment, it may involve a more complex file system architecture and permission verification mechanism. Finally, according to the offset and data length information, the system accurately reads the target shard data from the storage medium. The offset determines the starting read position of the data in the storage medium, and the data length specifies the amount of data to be read. The system extracts the target shard data from the storage medium according to these parameters and transfers it back to the requestor (which can be a user device, an application server, etc.), completing the entire read operation process, thereby meeting the acquisition requirements of the user or application for the target shard data.
[0132] The present disclosure provides a method for mutual access of unstructured protocol data. In the present disclosure, in view of the mutual access problem caused by the semantic differences between unstructured storage protocols, the data is processed by means of sharding according to the first data storage protocol, and the data is read in combination with the second data storage protocol. Effective access to shard data under different protocols is achieved through object indexing and shard layout information. This enables data to smoothly interact between storage systems of multiple protocols, breaking down protocol barriers. An object index containing key information such as storage location and index identifier is created for each shard file. When a read instruction is received, the system can quickly locate the storage location of the target shard data based on the index, without the need to comprehensively retrieve a large amount of data, reducing the data search time. Especially when dealing with massive unstructured data, this index-based fast positioning mechanism significantly improves the data read speed, speeds up the system response, and improves the data processing efficiency. Compared with the traditional solution that requires merging large files after shard uploading, it can avoid the resource consumption caused by large-scale data relocation and merging. It reduces the amount of data processing, reduces the requirements for the I / O performance of storage devices and network bandwidth, reduces the overall system load, helps to maintain the stable operation of the storage system, and improves the reliability and scalability of the system. When dealing with large files, the required data can be obtained without waiting for the file merging operation to complete, which can improve the convenience and fluency.
[0133] To clearly illustrate the embodiments of the present disclosure, this embodiment provides a schematic flowchart of another method for mutual access of unstructured protocol data.
[0134] As Figure 2 shown, the method includes the following steps:
[0135] Step 201: According to the shard upload semantics of the first data storage protocol, send a shard upload request for unstructured data to the storage system and generate a unique upload identifier.
[0136] Specifically in Step 201, the "first data storage protocol" is a storage protocol that supports shard upload semantics, such as the S3 object storage protocol. "Shard upload semantics" means that the protocol allows data to be split into multiple segments for upload, which can effectively solve problems such as resume from breakpoint and improve upload efficiency for large file uploads. "Unstructured data" refers to data with irregular data structures and no predefined models, such as documents, pictures, videos, etc. When storing this type of data, according to the shard upload rules of the first data storage protocol, the system will construct a shard upload request.
[0137] During the process of sending this request, the storage system will generate a "unique upload identifier". This identifier is specially allocated by the storage system for this shard upload task. It is unique and is used to identify all shard data of this upload operation. Subsequently, during the upload process, each shard data is associated with this upload identifier, so that the storage system can manage and integrate the entire upload task.
[0138] Step 202: Divide the unstructured data into shards according to the first data storage protocol to generate at least one shard data.
[0139] As a specific implementation, dividing the unstructured data into shards according to the first data storage protocol includes: dividing the unstructured data into at least one shard data according to the shard size threshold; assigning serial numbers to at least one shard data, and each shard data is assigned a unique shard serial number.
[0140] Specifically in Step 202, when performing shard division, it will first operate according to the "shard size threshold". The "shard size threshold" is a pre-set data volume standard used to determine the upper limit of the size of each shard data. The system cuts the unstructured data into at least one "shard data" according to this threshold. For example, if the shard size threshold is set to 10MB, a 50MB video file will be divided into 5 shard data. After the division is completed, in order to facilitate the management and identification of these shard data, "serial number assignment" will be performed on them. Each "shard data" will be assigned a "unique shard serial number". This shard serial number is like the "ID card number" of each shard data. In subsequent upload, storage, and integration operations, the system can accurately find the corresponding shard data through this unique serial number, ensuring the accuracy and efficiency of the entire data processing process.
[0141] Step 203: Create at least one shard file, store at least one shard of data into at least one shard file, and save one shard of data into one shard file.
[0142] As a specific implementation, creating at least one shard file and storing at least one shard of data into at least one shard file includes: creating at least one shard file and its corresponding storage path according to the shard sequence number and upload identifier of at least one shard of data; storing at least one shard of data into its corresponding at least one shard file according to the storage path respectively, and recording the offset and data length of each shard file respectively.
[0143] Specifically in Step 203, the "shard sequence number" is a unique number assigned to each shard of data in Step 202, which is used to distinguish different shards of data and facilitate subsequent management and identification. The "upload identifier" is generated in Step 201 and is used to identify the shard upload task of the entire unstructured data, ensuring that all relevant shards of data belong to the same upload operation. Based on this information, the system starts to "create at least one shard file and its corresponding storage path". The storage path is like the "address" of the file in the storage system and is generated by combining the shard sequence number and the upload identifier. For example, a root directory may be created according to the upload identifier, and then subdirectories and file names are created under it according to the shard sequence number, so that each shard file has a unique storage path. After the storage path is created, "store at least one shard of data into its corresponding at least one shard file according to the storage path respectively". This step ensures that each shard of data can be accurately stored into the corresponding shard file. At the same time, "record the offset and data length of each shard file respectively". The "offset" refers to the starting storage position of the shard file in the storage medium, and the "data length" represents the size of the data stored in the shard file. Recording this information helps to quickly locate and read the data in the shard file subsequently, improving the data access efficiency.
[0144] Step 204: Generate an object index corresponding to each shard file based on the storage location information and index identification information of each shard file. The storage location information includes: offset, data length, and storage path. The index identification information includes an index number.
[0145] Specifically in step 204, "storage location information" includes several key elements: offset, data length, and storage path. "Offset" refers to the location identifier where the fragmented file starts to be stored in the storage medium, similar to the starting page number of a certain chapter in a book. Through it, the starting point of the fragmented file in the storage device can be accurately located. "Data length" represents the size of the amount of data stored in the fragmented file, similar to the length of the chapter, which stipulates the range of data contained in the fragmented file. "Storage path" is like the detailed address of the file in the storage system, used to uniquely determine the storage location of the fragmented file in the directory structure of the storage system. The "index number" in the "index identification information" is a unique number assigned to each fragmented file, similar to everyone's ID card number, used to accurately identify and distinguish different files among many fragmented files.
[0146] Based on this information, the system generates an "object index" for each fragmented file. The object index is a data structure that integrates the storage location information and index identification information of the fragmented file. By combining this information, the object index can help the system quickly locate and access specific fragmented files in subsequent processing. For example, when a certain fragmented file needs to be read, the system can quickly find the location of the fragmented file in the storage system and perform the read operation according to the information in the object index, greatly improving the efficiency and accuracy of data processing.
[0147] In step 205, arrange the object indexes corresponding to at least one fragmented file in the order of the index numbers in the object index.
[0148] Specifically in step 205, the "object index" is generated in step 204, which is a data structure used to quickly locate and manage fragmented files. It integrates the storage location information (such as offset, data length, storage path) and index identification information (mainly the index number) of each fragmented file. The "index number" is the unique identifier assigned to the object index of each fragmented file, similar to an identification code, used to distinguish different fragmented files among many object indexes. When executing step 205, the system reads the object indexes corresponding to each fragmented file and extracts the index numbers therein. Then, these object indexes are rearranged in the order of the index numbers. This arrangement helps to manage and access fragmented files more efficiently in the future. For example, when a batch of fragmented files need to be read or processed, the object indexes arranged in the order of the index numbers can allow the system to access each fragmented file in a certain logical order, avoiding chaotic searches and improving the efficiency and accuracy of data processing.
[0149] In step 206, associate and store at least one fragmented file and the object index to generate a large fragmented file, and store the fragmentation identifier of the large fragmented file in the object index.
[0150] Specifically in step 206, the "sharded file" is a file unit created in step 203 for storing sharded data. Each sharded file stores a part of the original unstructured data. The "object index" is generated in step 204, which contains the storage location information (offset, data length, storage path) and index identification information (index number) of the sharded file, and is used for quickly locating and managing the sharded file. "Associated storage" is to establish a logical connection between at least one sharded file and the corresponding object index to ensure that each sharded file can be accessed and managed through its object index. Through this association, the storage system can efficiently organize and search for data. The "sharded large file" is a logical file that integrates multiple sharded files. It is not a simple physical merge, but rather realizes the unified management of each sharded file through the object index. When generating the sharded large file, the system associates each sharded file with the corresponding object index to form a whole.
[0151] The "shard identifier" is the identification information used to uniquely identify each shard in the sharded large file. The system stores the shard identifier of the sharded large file in the object index. In this way, when operating on the sharded large file subsequently, relevant information of each shard, including the shard identifier, can be obtained through the object index, facilitating the accurate access and management of each shard in the sharded large file and improving the accuracy and efficiency of data processing.
[0152] In step 207, parse the access request of the second data storage protocol, determine the target data length and target offset of the target sharded data to be accessed; based on the target data length and target offset, generate a read instruction for reading the target sharded data.
[0153] Specifically in step 207, the "second data storage protocol" refers to another data storage interaction specification different from the first data storage protocol, such as the common HDFS protocol or NAS file protocol, etc., which is used to stipulate the access method of data in the storage system. The "access request" is an instruction for a user or application program to initiate data acquisition or operation to the storage system, and in this scenario, it is initiated based on the second data storage protocol. When the system receives this access request, it will "parse" it. During the parsing process, the system extracts two key parameters from the information carried by the access request: the "target data length" and the "target offset". The "target data length" represents the number of bytes of the target sharded data that the user expects to obtain, that is, the amount of data to be obtained in this read operation. The "target offset" specifies from which position of the target sharded data to start reading, similar to starting to read a specified length from a certain page of a book.
[0154] Next, the system generates a "read instruction" specifically for reading the target shard data based on the obtained "target data length" and "target offset". This read instruction contains clear operation requirements and parameters, which inform the storage system where to start reading and how much data to read. The storage system accurately obtains the target shard data from the corresponding shard file according to this instruction, thus meeting the data access requirements of users or applications.
[0155] Step 208: Determine whether the target file pointed to by the read instruction is a sharded large file generated according to the first data storage protocol.
[0156] If the target file is a sharded large file, find the storage location information of the target shard data from the object index.
[0157] As a specific implementation, determining whether the target file pointed to by the read instruction is a sharded large file generated according to the first data storage protocol includes: reading the metadata of the target file and determining whether there is a shard identifier in the metadata.
[0158] Specifically in step 208, the "read instruction" is generated in step 207 according to the access request of the second data storage protocol. It contains key information such as the target data length and target offset, and is used to instruct the storage system to read the target shard data. The "target file" is the data file that this read instruction wants to access. It may be an ordinary file or a sharded large file generated according to the first data storage protocol. When performing the judgment operation, the system will "read the metadata of the target file". "Metadata" is data about data, which contains various attribute information of the file, such as file type, size, creation time, etc. In this scenario, the focus is on whether there is a "shard identifier" in the metadata. The "shard identifier" is stored in the object index when generating the sharded large file in step 206, and is a special identifier used to uniquely identify each shard in the sharded large file. If a shard identifier is detected in the metadata of the target file, it can be determined that the target file is a sharded large file generated according to the first data storage protocol. If it is determined that the target file is a sharded large file, the system will find the storage location information of the target shard data from the "object index". The "object index" integrates the storage location information (offset, data length, storage path) of the shard file and the index identifier information (index number). Through it, the specific location of the target shard data in the storage system can be quickly located, preparing for subsequent data reading operations.
[0159] Step 209: Based on the read instruction, read the object index of the shard file to determine the target shard range where the target shard data is located.
[0160] As a specific implementation, based on the read instruction, the object index of the sharded file is read, and the target shard range where the target shard data is located is determined, including: parsing the read instruction to obtain the target data length and target offset of the target shard data; traversing each object index to determine the object index overlapping with the target data length and target offset, and generating the target shard range based on the overlapping object indexes.
[0161] Specifically in step 209, the system will "parse the read instruction". Through parsing, the system extracts the "target data length" and "target offset" from the read instruction. The "target data length" specifies the number of bytes of the target shard data that the user wants to obtain, and the "target offset" specifies from which position of the target shard data to start reading. Then, the system will "traverse each object index". The system checks each object index in turn, and compares the offset and data length of the sharded file recorded therein with the target data length and target offset obtained from the read instruction. When it is found that the offset and data length of the sharded file corresponding to an object index overlap with the target data length and target offset, it is determined that the object index is related to the target data.
[0162] Finally, the system will "generate the target shard range based on the overlapping object indexes". The sharded files corresponding to all the overlapping object indexes are determined as the target shard range. These sharded files together constitute the storage set of the target shard data that can meet the requirements of the read instruction, laying a foundation for accurately reading the target shard data subsequently.
[0163] Step 210, obtain the target storage path of each target sharded file overlapping with the target shard range, and calculate the target read range corresponding to each target sharded file.
[0164] As a specific implementation, calculating the target read range corresponding to each target sharded file includes: calculating the target read range of the starting target sharded file based on the target offset of the read instruction, the offset and data length of the starting target sharded file within the target shard range; calculating the target read range of the ending target sharded file according to the target offset of the read instruction, the target read length, and the offset of the ending target sharded file within the target shard range. If the target sharded file is an overlapping target sharded file that completely overlaps with the data read range, then the data range of the overlapping target sharded file is used as the target read range.
[0165] Specifically, in step 210, the "target shard range" is determined based on the read instruction in step 209. It is a set of shard file collections that contain the target shard data. The "target shard file" refers to each shard file within the target shard range, and these files store the content that may contain the target shard data. The "target storage path" is the specific storage location identifier of each target shard file in the storage system, similar to the "address" of a file in the storage system, through which the corresponding shard file can be accurately accessed. When performing this step, first, obtain the target storage path of each target shard file that overlaps with the target shard range. This step is to prepare for subsequent data reading to ensure that the system knows where to read the data. Next, calculate the target read range corresponding to each target shard file. Here, several key concepts are involved: The "read instruction" contains key information such as the target offset and the target data length, which is used to guide the data reading operation. The "target offset" specifies from which position in the target shard data to start reading, and the "target data length" specifies the amount of data to be read. The "starting target shard file" and the "ending target shard file" are respectively the first and the last shard files within the target shard range. When calculating the target read range of the starting target shard file, the system determines it based on the target offset of the read instruction, the offset of the starting target shard file itself, and its data length. The offset of the starting target shard file represents its starting position in the storage medium, and the data length represents the amount of data it contains. By comparing these values, determine from which position in the starting target shard file to start reading and how much data to read, thereby obtaining the target read range of the starting target shard file.
[0166] For the ending target shard file, the system calculates its target read range based on the target offset of the read instruction, the target read length (i.e., the target data length), and the offset of the ending target shard file. Through this information, determine the starting position and the read length of the data to be read in the ending target shard file. If a certain target shard file is an "overlapping target shard file", that is, its data range completely overlaps with the data read range required by the read instruction, then directly use the data range of this overlapping target shard file as the target read range. In this way, it is possible to efficiently and accurately determine the data part that each target shard file needs to read, providing precise guidance for subsequent data reading operations.
[0167] Step 211: According to the target read range and the shard layout information of the target shard data, find the storage location information of the target shard data.
[0168] Specifically in step 211, the "target reading range" defines the starting position and length of the data to be read in each target shard file. This range is determined based on information such as the target offset and target data length in the reading instruction, as well as the offset and data length of the target shard file itself. The "target shard data" is the specific data segment that the user hopes to obtain, which is part of the data after the original unstructured data has been sharded. The "shard layout information" includes the distribution of each shard file in the storage system and related attribute information, such as the offset, data length, and storage path of each shard file. This information is generated and stored in the object index in the previous steps, and is used to describe the storage status and location characteristics of the shard file.
[0169] When executing step 211, the system combines the target reading range with the shard layout information of the target shard data. The system uses the information such as the offset and data length of each shard file recorded in the shard layout information to compare and match with the target reading range. In this way, the system can accurately find the specific storage location where the target shard data is located among numerous shard files, so as to obtain the storage location information of the target shard data, including detailed information such as the storage device, storage path, and specific offset of the data in the storage medium. These storage location information will be used to guide the next data reading operation to ensure that the system can accurately read the target shard data required by the user from the storage system.
[0170] Step 212, read the target shard data based on the storage location information.
[0171] Specifically in step 212, when the system obtains the storage location information of the target shard data, it will initiate the reading operation. First, the system will establish a connection with the corresponding storage device according to the storage device identifier in the storage location information. In a distributed storage environment, it may be necessary to communicate with the storage server through a network protocol (such as the TCP / IP protocol) to ensure that the node storing the target shard data can be accessed. Then, based on the directory path information, the system navigates to the specific location containing the target shard data in the file system of the storage device. This process is similar to finding a file through a path in the local file system, except that in a distributed or complex storage environment, it may involve a more complex file system architecture and permission verification mechanism. Finally, according to the offset and data length information, the system accurately reads the target shard data from the storage medium. The offset determines the starting reading position of the data in the storage medium, and the data length specifies the amount of data to be read. The system extracts the target shard data from the storage medium according to these parameters and transfers it back to the request end (which can be a user device, an application server, etc.) to complete the entire reading operation process, thus meeting the acquisition requirements of the user or application program for the target shard data.
[0172] Furthermore, in other embodiments of the present disclosure, based on the data reading method of the above embodiments, not only can data reading be achieved, but also operations such as data modification and deletion can be performed. For the modification of target sharded data, the following methods can be used but are not limited to: in response to a modification instruction for target sharded data issued based on the second data storage protocol, query the object index of the sharded file to determine the storage location information of the target sharded data; based on the storage location information, read the target sharded data from at least one sharded file for modification, and modify the object index according to the modification information.
[0173] As a specific implementation manner, modifying the object index according to the modification information includes:
[0174] Obtain the modification offset and modification data length of the modified sharded file;
[0175] Based on the modification offset and the modification data length, modify the object index.
[0176] Specifically, the "second data storage protocol" is a set of rules for interacting with a storage system, which stipulates the ways and formats of data access and operations, such as common HDFS protocols, NAS file protocols, etc. "Target sharded data" refers to the sharded data that the user hopes to modify, which is a subset of the original unstructured data. The "modification instruction" is a request sent by the user or application program to the storage system to modify the target sharded data. This instruction contains clear modification intentions and necessary parameters to inform the system of key information such as the data content and location to be modified. When the system receives a modification instruction for target sharded data issued based on the second data storage protocol, a series of operations will be initiated. First, the system will "query the object index of the sharded file". The "object index" is a key data structure that integrates important information of each sharded file, including "storage location information". The storage location information covers the specific storage location of the sharded file in the storage system, such as the address of the storage device, the file directory path, the offset in the storage medium, and the data length, similar to the "detailed address" provided for each sharded file. By querying the object index, the system can quickly determine the storage location information of the target sharded data, laying a foundation for subsequent data reading and modification. Then, the system "reads the target sharded data from at least one sharded file for modification based on the storage location information". The system accurately finds the sharded file containing the target sharded data according to the obtained storage location information and reads the target sharded data therein. After that, the read data is modified according to the specific requirements in the modification instruction.
[0177] After the modification is completed, it is necessary to "modify the object index according to the modification information" to ensure that the information in the object index is consistent with the modified data. During this process, "obtain the modification offset and the length of the modified data of the modified shard file". The "modification offset" represents the starting position of the modified data in the shard file, and the "length of the modified data" represents the length of the modified data. These two parameters reflect the specific impact range of the modification operation on the data in the shard file.
[0178] Finally, "modify the object index based on the modification offset and the length of the modified data". The system uses the obtained modification offset and the length of the modified data to update the relevant information of the corresponding shard file in the object index. For example, update the offset and the length of the data of the shard file recorded in the object index to make it consistent with the actual situation after modification. In this way, when subsequent data reading or other operations are performed again, the system can accurately obtain the modified data according to the updated object index, ensuring the accuracy and consistency of data operations.
[0179] Furthermore, in other embodiments of the present disclosure, based on the data reading method of the above embodiments, not only can data reading be achieved, but also operations such as data modification and deletion can be performed. For the deletion of target shard data, the following methods can be used but are not limited to: in response to a deletion instruction for target shard data issued based on the second data storage protocol, query the object index of the shard file to determine the storage location information of the target shard data;
[0180] Based on the storage location information, read the target shard data from at least one shard file for deletion.
[0181] Specifically, the "second data storage protocol" is a protocol that stipulates data storage, access, and operation rules. For example, common HDFS (Hadoop Distributed File System) protocol or NAS (Network Attached Storage) file protocol, etc. It defines the specifications for interacting with the storage system. The "target shard data" refers to a specific part of the original unstructured data after sharding processing, and this part of the data is the object on which the user hopes to perform a deletion operation. The "deletion instruction" is the request information sent by the user or application program to the storage system for deleting the target shard data. This instruction contains necessary parameters such as the key identifier for identifying the target shard data, so that the storage system can accurately find the data to be deleted.
[0182] After the system receives a deletion instruction for target shard data issued based on the second data storage protocol, it will perform a series of ordered operations. First, the system will "query the object index of the shard file". The "object index" is a key data structure that integrates the storage location information, index identification information, etc. of each shard file. Among them, the "storage location information" records the specific storage location of the shard file in the storage system, covering details such as storage device identification, file path, offset in the storage medium, and data length, just like the "exact address" set for each shard file in the storage system. By querying the object index, the system can quickly determine the storage location information of the target shard data, providing a positioning basis for subsequent data deletion operations.
[0183] Next, the system "reads the target shard data from at least one shard file based on the storage location information for deletion". The system accurately locates one or more shard files containing the target shard data according to the obtained storage location information. Then, it reads the target shard data from these shards and performs a deletion operation in the storage system to remove the target shard data from the storage medium. Through this series of operations, the deletion of the target shard data is completed, ensuring that the data in the storage system meets the latest requirements of users and maintaining the accuracy and consistency of data management.
[0184] It should be noted that the embodiments of the present disclosure may include multiple steps. For the sake of description, these steps are numbered, but these numbers are not intended to limit the execution time slots and execution orders between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not make any limitations in this regard.
[0185] Corresponding to the above method for mutual access of unstructured protocol data, the present disclosure also proposes an apparatus for mutual access of unstructured protocol data. Since the apparatus embodiments of the present disclosure correspond to the above method embodiments, for the details not disclosed in the apparatus embodiments, reference may be made to the above method embodiments, and the present disclosure will not repeat them here.
[0186] Figure 3 Shown in the following is a schematic structural diagram of an apparatus for mutual access of unstructured protocol data provided by an embodiment of the present disclosure, as Figure 3 shown, including:
[0187] A first generation unit 31, configured to perform sharding processing on unstructured data based on a first data storage protocol to generate at least one shard file;
[0188] A creation unit 32, configured to create an object index for each shard file, where the object index includes the shard layout information of the shard file, and the shard layout information at least includes the storage location information and index identification information of the allocated file;
[0189] A first query unit 33, configured to query the object index of the shard file to determine the storage location information of the target shard data in response to a read instruction for the target shard data issued based on the second data storage protocol;
[0190] A read unit 34, configured to read the target shard data based on the storage location information.
[0191] The present disclosure provides a device for cross-accessing unstructured protocol data. In the present disclosure, aiming at the cross-access problem caused by the semantic differences between unstructured storage protocols, data is processed in shards by means of the first data storage protocol, and data is read in combination with the second data storage protocol. Effective access to shard data under different protocols is achieved through object indexes and shard layout information. This enables data to smoothly interact between storage systems of multiple protocols, breaking down protocol barriers. An object index containing key information such as storage location and index identifier is created for each shard file. When a read instruction is received, the system can quickly locate the storage location of the target shard data based on the index, without comprehensively retrieving a large amount of data, reducing the data search time. Especially when dealing with massive unstructured data, this index-based fast positioning mechanism significantly improves the data reading speed, speeds up system response, and improves data processing efficiency. Compared with the traditional solution that requires merging large files after shard uploading, it can avoid the resource consumption caused by large-scale data relocation and merging. It reduces the amount of data processing, reduces the requirements for the I / O performance of storage devices and network bandwidth, reduces the overall system load, helps maintain the stable operation of the storage system, and improves the reliability and scalability of the system. When dealing with large files, the required data can be obtained without waiting for the file merging operation to complete, which can improve convenience and fluency.
[0192] It should be noted that the foregoing explanation of the method embodiments also applies to the device of this embodiment, with the same principle, and will not be limited in this embodiment.
[0193] For the description of the features in the corresponding embodiment of the device for cross-accessing unstructured protocol data, reference can be made to the relevant description in the corresponding embodiment of the method for cross-accessing unstructured protocol data, which will not be elaborated here one by one.
[0194] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the foregoing method embodiments for cross-accessing unstructured protocol data.
[0195] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the foregoing method embodiments for cross-accessing unstructured protocol data when running.
[0196] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to: various media capable of storing computer programs such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs.
[0197] The embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the method embodiments of the above-mentioned unstructured protocol data mutual access.
[0198] The embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the method embodiments of the above-mentioned unstructured protocol data mutual access.
[0199] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be regarded as exceeding the scope of the present application.
[0200] The above has introduced in detail a method, apparatus, electronic device, and medium for unstructured protocol data mutual access provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for mutual access to unstructured protocol data, characterized in that The method includes: Performing sharding processing on unstructured data based on a first data storage protocol to generate at least one shard file; Creating an object index for each of the shard files, where the object index includes shard layout information of the shard file, and the shard layout information at least includes storage location information of the allocated file and index identification information; In response to a read instruction for target shard data issued based on a second data storage protocol, querying the object index of the shard file to determine the storage location information of the target shard data; Reading the target shard data based on the storage location information.
2. The method for mutual access of unstructured protocol data according to claim 1, characterized in that, The creating an object index for each of the shard files includes: Generating the object index corresponding to each of the shard files based on the storage location information and index identification information of each of the shard files, where the storage location information includes: offset, data length, and storage path, and the index identification information includes an index number; Arranging the object indexes corresponding to the at least one shard file in the order of the index numbers in the object index.
3. The method for mutual access to unstructured protocol data according to claim 1, wherein After creating an object index for each of the shard files, the method further includes: Associatively storing the at least one shard file and the object index to generate a sharded large file, and storing the shard identifier of the sharded large file to the object index.
4. The method for mutual access to unstructured protocol data according to claim 3, wherein Before, in response to a read instruction for target shard data issued based on a second data storage protocol, querying the object index of the shard file to determine the storage location information of the target shard data, the method further includes: Parsing the access request of the second data storage protocol to determine the target data length and target offset of the target shard data to be accessed; Generating the read instruction for reading the target shard data based on the target data length and the target offset; Determining whether the target file pointed to by the read instruction is the sharded large file generated according to the first data storage protocol; if the target file is the sharded large file, then searching for the storage location information of the target shard data from the object index.
5. The method for mutual access of unstructured protocol data according to claim 4, wherein The determining whether the target file pointed to by the read instruction is the sharded large file generated according to the first data storage protocol includes: Reading the metadata of the target file and determining whether there is a shard identifier in the metadata.
6. The method for mutual access to unstructured protocol data according to claim 1, characterized in that The in response to a read instruction for target shard data issued based on a second data storage protocol, querying the object index of the shard file to determine the storage location information of the target shard data includes: Based on the read instruction, reading the object index of the shard file to determine the target shard range where the target shard data is located; Obtaining the target storage path of each target shard file overlapping with the target shard range, and calculating the target read range corresponding to each of the target shard files; Searching for the storage location information of the target shard data according to the target read range and the shard layout information of the target shard data.
7. The method for mutual access of unstructured protocol data according to claim 6, wherein The based on the read instruction, reading the object index of the shard file to determine the target shard range where the target shard data is located includes: Parse the read instruction to obtain the target data length and target offset of the target shard data; Traverse each of the object indexes to determine the object indexes that overlap with the target data length and the target offset, and generate the target shard range based on the overlapping object indexes.
8. The method for mutual access of unstructured protocol data according to claim 6, wherein The calculating the target read range corresponding to each of the target shard files includes: Based on the target offset of the read instruction, the offset and data length of the starting target shard file within the target shard range, calculate the target read range of the starting target shard file; According to the target offset of the read instruction, the target read length, and the offset of the ending target shard file within the target shard range, calculate the target read range of the ending target shard file.
9. The method for mutual access of unstructured protocol data according to claim 8, wherein The method further includes: If the target shard file is an overlapping target shard file that completely overlaps with the data read range, use the data range of the overlapping target shard file as the target read range.
10. The method for mutual access of unstructured protocol data according to claim 1, wherein The fragmenting the unstructured data based on the first data storage protocol to generate at least one shard file includes: Fragment the unstructured data according to the first data storage protocol to generate at least one shard data; Create at least one shard file, store the at least one shard data into the at least one shard file, and store one shard data into one shard file.
11. The method for mutual access of unstructured protocol data according to claim 10, characterized in that, The fragmenting the unstructured data according to the first data storage protocol includes: Divide the unstructured data into the at least one shard data according to a shard size threshold; Assign sequence numbers to the at least one shard data, and assign a unique shard sequence number to each shard data.
12. The method for mutual access of unstructured protocol data according to claim 11, wherein Before fragmenting the unstructured data according to the first data storage protocol, the method further includes: According to the shard upload semantics of the first data storage protocol, send a shard upload request for the unstructured data to the storage system and generate a unique upload identifier.
13. The method for mutual access of unstructured protocol data according to claim 12, characterized in that, The creating at least one shard file and storing the at least one shard data into the at least one shard file includes: Create the at least one shard file and its corresponding storage path according to the shard sequence numbers of the at least one shard data and the upload identifier; Store the at least one shard data into its corresponding at least one shard file according to the storage path, and record the offset and data length of each shard file respectively.
14. The method for mutual access of unstructured protocol data according to any one of claims 1-13, characterized in that, The method further includes: In response to a modification instruction for the target shard data issued based on the second data storage protocol, query the object index of the shard file to determine the storage location information of the target shard data; Based on the storage location information, read the target shard data from the at least one shard file for modification, and modify the object index according to the modification information.
15. The method for mutual access of unstructured protocol data according to claim 14, wherein The modifying the object index according to the modification information includes: Obtain the modified offset and modified data length of the modified shard file; Modify the object index based on the modified offset and the modified data length.
16. The method for mutual access of unstructured protocol data according to any one of claims 1-13, characterized in that, The method further includes: In response to a deletion instruction for target shard data issued based on the second data storage protocol, query the object index of the shard file to determine the storage location information of the target shard data; Based on the storage location information, read the target shard data from the at least one shard file for deletion.
17. A device for mutual access to unstructured protocol data, characterized in that, The apparatus includes: A first generation unit, configured to perform sharding processing on unstructured data based on a first data storage protocol to generate at least one shard file; A creation unit, configured to create an object index for each of the shard files, where the object index includes shard layout information of the shard file, and the shard layout information at least includes storage location information of an allocated file and index identification information; A first query unit, configured to, in response to a read instruction for target shard data issued based on the second data storage protocol, query the object index of the shard file to determine the storage location information of the target shard data; A read unit, configured to read the target shard data based on the storage location information.
18. An electronic device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for mutual access of unstructured protocol data according to any one of claims 1-16.
19. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method for mutual access of unstructured protocol data according to any one of claims 1-16.
20. A computer program product, characterized in that, Comprising a computer program, where the computer program, when executed by a processor, implements the method for mutual access of unstructured protocol data according to any one of claims 1-16.