A Method and System for Storing and Retrieving Deep Learning Datasets
The hash algorithm generates unique file encoding of the deep learning data set, and uses encoding interception to improve retrieval efficiency, solving the problems of data duplication and inability to retrieve, and achieving efficient data storage and retrieval.
Patent Information
- Application Number
- CN202310791755.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-06-29
AI Technical Summary
The prior art is difficult to effectively manage and store deep learning data sets, especially unstructured data, which leads to the problem of data duplication and inability to retrieve, and thus wastes storage space.
The deep learning data set is traversed through the preset hash algorithm, and a unique file encoding for each data item is generated, and the retrieval efficiency is improved through encoding interception. If the encoded value of the data item is the same as the existing folder name, it is determined to be duplicate data and deleted.
It realizes the uniqueness and efficient storage of deep learning data sets, avoids data duplication and space waste, and solves the problem of data being unable to flow through the retrieval of label information.
Smart Images

Figure CN116795788B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method and system for storing and retrieving deep learning data sets. Background Art
[0002] With the substantial improvement of computing power, deep learning has developed rapidly and become the main growth point and driving force of artificial intelligence. In addition, the advent of the big data era has led to the evolution of deep learning from the original AI model development-centered to data-centered development, which has also greatly promoted the development efficiency and higher accuracy of AI algorithm models.
[0003] Different from traditional structured data, the common formats of deep learning data sets are unstructured data such as videos, audios, pictures, texts, etc., and the data volume is large. Traditional databases cannot store them. Currently, the management of data sets is mainly stored in disk media in the form of file naming. This not only makes data labels unable to be retrieved, resulting in data immobility, but also there will be a large amount of duplicate data between multiple compressed packages, wasting storage space. Summary of the Invention
[0004] The present invention discloses a method and system for storing and retrieving deep learning data sets, which solve the technical problems of data duplication and inability to be retrieved in the prior art, and reduce the space occupied by data storage.
[0005] To achieve the above object, the present invention discloses a method for storing and retrieving deep learning data sets, including:
[0006] Traverse each data item in the obtained deep learning data set according to a preset hash algorithm, and generate a first file code corresponding to each data item;
[0007] Perform encoding truncation on the first file code of each data item respectively to obtain an encoding value corresponding to each data item, and retrieve according to the encoding value to determine whether there is a first folder named after the encoding value;
[0008] If there is a first folder named after the encoding value, compare the first file code of the data item corresponding to the encoding value with the second file code of the data item existing in the first folder to determine whether there is a second file code that is the same as the first file code; if so, delete the data item corresponding to the first file code; if not, add the data item corresponding to the first file code to the first folder;
[0009] If there is no first folder named with the encoding value, add the data item corresponding to the encoding value to a preset folder, and name the preset folder according to the encoding value to obtain the first folder;
[0010] Generate a file address corresponding to each data item according to the first folder, and save the file encoding, dataset name, and annotation information corresponding to each data item in the deep learning dataset through a preset database for retrieval.
[0011] The present invention discloses a method for storing and retrieving a deep learning dataset, including traversing each data item in the deep learning dataset by using a preset hash algorithm to generate a file encoding corresponding to each data item one by one, ensuring the uniqueness of the data item. Then, after generating the file encoding corresponding to the data item, in order to improve the efficiency of saving the data item, several encodings are intercepted from the file encoding as the name of the folder where the data item is located, so that when saving the folder later, the retrieval result can be obtained by using part of the encoding for retrieval, improving the efficiency of file saving. And the folder is named by using the encoding value. If the encoding value of the newly added data item is the same as the name of the existing folder, it is determined that the data item is repeated, and the data item is deleted, avoiding data duplication and reducing the space resource occupancy rate.
[0012] As a preferred example, the method for storing and retrieving a deep learning dataset further includes:
[0013] Retrieve the database according to the preset annotation information and dataset name to obtain the file encoding of the data item to be retrieved, and generate a retrieval path according to the file encoding;
[0014] Retrieve several data items saved in the first folder according to the retrieval path to obtain the data item to be retrieved.
[0015] In the present invention, the database saves the annotation information, dataset name, and file encoding of the data item to be retrieved. After obtaining the annotation information and dataset name of the data item to be retrieved, the file encoding of the data item to be retrieved can be determined. According to the file encoding, the address of the folder where it is saved can be determined, and then the data item to be retrieved can be searched in the folder according to the address, solving the technical problem that the annotation information cannot be retrieved in the prior art, resulting in the inability of data to flow, and realizing the retrieval of data.
[0016] As a preferred example, when traversing each data item in the deep learning dataset obtained according to the preset hash algorithm and generating a first file encoding corresponding to each data item, it includes:
[0017] Traverse the obtained deep learning dataset in a loop to obtain a number of data items included in the deep learning dataset; the data items include videos, audios, pictures, and texts.
[0018] Perform hash calculation on each of the number of data items according to a preset hash algorithm to obtain the MD5 value corresponding to each data item.
[0019] The present invention first traverses the deep learning dataset in a loop to obtain all the data items included in the deep learning dataset, avoiding omission of data storage. Then, it uses a hash algorithm to generate the MD5 value corresponding to each data item. Based on the high confidentiality of the MD5 value, the accuracy of the data item corresponding to each data item is guaranteed, and the accuracy rate of file storage is improved.
[0020] As a preferred example, in encoding and intercepting the first file encoding of each data item to obtain the encoding value corresponding to each data item, and performing retrieval according to the encoding value, it includes:
[0021] According to a preset interception rule, perform encoding interception on the first file encoding corresponding to each data item through a string interception algorithm to obtain the encoding value corresponding to each data item.
[0022] Retrieve from a preset database according to the encoding value, and judge whether there is a folder named with the encoding value in the database.
[0023] After obtaining the file encoding, since the amount of data of the file encoding is large, in order to improve the file storage efficiency, a preset string interception algorithm is used to intercept a part of the encoding from the file encoding as the retrieval basis, thereby reducing the amount of data for data matching and improving the retrieval efficiency.
[0024] As a preferred example, in adding the data item corresponding to the encoding value to a preset folder and naming the preset folder according to the encoding value to obtain the first folder, it includes:
[0025] Move the data item into the folder according to a preset file moving method.
[0026] Rename the preset folder through a preset file renaming algorithm according to the encoding value corresponding to the data item to obtain the first folder named with the encoding value.
[0027] The present invention names the folder according to the encoded value, and at the same time saves the data item corresponding to the encoded value in the folder, ensuring that during the data storage process, duplicate data storage is avoided through the encoded value, thus avoiding waste of space resources.
[0028] As a preferred example, generating a retrieval path according to the file encoding includes:
[0029] Continuing to encode and intercept the obtained file encoding to obtain the encoded value corresponding to the file encoding, and then obtaining the storage path of the folder where the encoded value is located and the folder name of the folder according to the encoded value;
[0030] According to the storage path of the folder and the folder name of the folder, combining with the file encoding to generate the retrieval path of the data to be retrieved.
[0031] The present invention solves the technical problem that data cannot be retrieved using annotation information in the prior art by saving the annotation information, file encoding, and dataset name of each data item in the database, enabling data retrieval by inputting the annotation information.
[0032] On the other hand, the present invention discloses a deep learning dataset storage and retrieval system, including a file encoding module, an encoding interception module, a file saving module, and a data saving module;
[0033] The file encoding module is used to traverse each data item in the obtained deep learning dataset according to a preset hash algorithm and generate a first file encoding corresponding to each data item;
[0034] The encoding interception module is used to respectively perform encoding interception on the first file encoding of each data item to obtain the encoded value corresponding to each data item, and perform a retrieval according to the encoded value to determine whether there is a first folder named with the encoded value;
[0035] The file saving module is used to, if there is a first folder named with the encoded value, compare the first file encoding of the data item corresponding to the encoded value with the second file encoding of the data item existing in the first folder to determine whether there is a second file encoding identical to the first file encoding; if so, delete the data item corresponding to the first file encoding; if not, add the data item corresponding to the first file encoding to the first folder; if there is no first folder named with the encoded value, add the data item corresponding to the encoded value to a preset folder, and name the preset folder according to the encoded value to obtain the first folder;
[0036] The data storage module is used to generate a file address corresponding to each data item according to the first folder, and save the file code, dataset name, and annotation information corresponding to each data item in the deep learning dataset through a preset database for retrieval.
[0037] The present invention discloses a deep learning dataset storage and retrieval system, which includes traversing each data item in the deep learning dataset by using a preset hash algorithm to generate a file code corresponding to each data item one by one, ensuring the uniqueness of the data item. Then, after generating the file code corresponding to the data item, in order to improve the efficiency of data item storage, several codes are intercepted from the file code as the name of the folder where the data item is located, so that when saving the folder subsequently, the retrieval result can be obtained by using part of the code for retrieval, improving the efficiency of file storage. And the folder is named by using the code value. If the code value of the newly added data item is the same as the name of the existing folder, it is determined that the data item is repeated, and then the data item is deleted, avoiding data duplication and reducing the space resource occupancy rate.
[0038] As a preferred example, the deep learning dataset storage and retrieval system further includes a path module and a retrieval module;
[0039] The path module is used to retrieve the database according to the preset annotation information and dataset name, obtain the file code of the data item to be retrieved, and generate a retrieval path according to the file code;
[0040] The retrieval module is used to retrieve several data items saved in the first folder according to the retrieval path to obtain the data item to be retrieved.
[0041] In the present invention, the database saves the annotation information, dataset name, and file code of the data item to be retrieved. After obtaining the annotation information and dataset name of the data item to be retrieved, the file code of the data item to be retrieved can be determined. According to the file code, the address of the folder where it is saved can be determined, and then the data item to be retrieved can be obtained by searching in the folder according to the address, solving the technical problem that the annotation information cannot be retrieved in the prior art, resulting in data immobility, and realizing data retrieval.
[0042] As a preferred example, the file coding module includes a traversing unit and a coding unit;
[0043] The traversing unit is used to circularly traverse the obtained deep learning dataset to obtain several data items included in the deep learning dataset; the data items include videos, audios, pictures, and texts;
[0044] The encoding unit is used to perform hash calculation on each of the several data items according to a preset hash algorithm to obtain the MD5 value corresponding to each data item.
[0045] The present invention first traverses the deep learning dataset in a loop to obtain all the data items included in the deep learning dataset, avoiding omission of data storage. Then, it uses a hash algorithm to generate the MD5 value corresponding to each data item. Based on the high confidentiality of the MD5 value, the accuracy of the data item corresponding to each data item is ensured, and the accuracy of file storage is improved.
[0046] As a preferred example, the encoding truncation module includes a truncation unit and a retrieval unit;
[0047] The truncation unit is used to perform encoding truncation on the first file encoding corresponding to each data item according to a preset truncation rule through a string truncation algorithm to obtain the encoding value corresponding to each data item;
[0048] The retrieval unit is used to retrieve from a preset database according to the encoding value to determine whether there is a folder named with the encoding value in the database.
[0049] After obtaining the file encoding in the present invention, due to the large amount of data of the file encoding, in order to improve the file storage efficiency, a part of the encoding is intercepted from the file encoding by using a preset string truncation algorithm as the retrieval basis, thereby reducing the amount of data for data matching and improving the retrieval efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 : It is a schematic flowchart of a method for storing and retrieving a deep learning dataset provided by an embodiment of the present invention;
[0051] Figure 2 : It is another schematic flowchart of a method for storing and retrieving a deep learning dataset provided by an embodiment of the present invention;
[0052] Figure 3 : It is a schematic structural diagram of a system for storing and retrieving a deep learning dataset provided by an embodiment of the present invention;
[0053] Figure 4 : It is another schematic structural diagram of a system for storing and retrieving a deep learning dataset provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0055] Embodiment
[0056] An embodiment of the present invention provides a method for storing and retrieving a deep learning dataset. For the specific implementation process of the storage and retrieval method, please refer to Figure 1 , which mainly includes steps 101 to 105. The steps are as follows:
[0057] Step 101: Traverse each data item in the obtained deep learning dataset according to a preset hash algorithm, and generate a first file code corresponding to each data item.
[0058] In this embodiment, this step mainly includes: circularly traversing the obtained deep learning dataset to obtain several data items included in the deep learning dataset; the data items include videos, audios, pictures, and texts; performing hash calculation on each data item in the several data items according to a preset hash algorithm to obtain an MD5 value corresponding to each data item.
[0059] Exemplarily, in this embodiment, after obtaining the deep learning dataset to be stored, circularly traversing the deep learning dataset can obtain each data item included in the deep learning dataset. Multiple data items form a deep learning dataset. The data items include data such as videos, audios, pictures, and texts. For example, multiple pictures of cars form a car deep learning dataset. According to a preset hash algorithm, a file code corresponding to each data item is calculated. In this embodiment, preferably, the MD5 algorithm can be used for hashing to obtain a unique MD5 value corresponding to each data item.
[0060] In this embodiment, first, the deep learning dataset is circularly traversed to obtain all the data items included in the deep learning dataset, avoiding omission of data storage. Then, the MD5 value corresponding to each data item is generated by using the hash algorithm. Based on the high confidentiality of the MD5 value, the accuracy of each data item is ensured, and the accuracy of file saving is improved.
[0061] Step 102: Perform encoding truncation on the first file code of each data item respectively to obtain an encoding value corresponding to each data item, and perform retrieval according to the encoding value to determine whether there is a first folder named with the encoding value.
[0062] In this embodiment, this step mainly includes: according to a preset truncation rule, performing encoding truncation on the first file encoding corresponding to each data item through a string truncation algorithm to obtain the encoding value corresponding to each data item; retrieving from a preset database according to the encoding value to determine whether there is a first folder named with the encoding value in the database.
[0063] Exemplarily, in this embodiment, several bits of data are truncated from the file encoding corresponding to each data item for folder search. Preferably, the string truncation algorithm can be used to truncate the last two bits of the file encoding as the encoding value corresponding to the data item. Based on the above steps, the last two bits of the MD5 value are selected as the encoding value, and then whether there is a folder named with the encoding value is identified and retrieved according to the last two bits of the encoding. Preferably, before performing the folder search, when saving the data item, the mobile file API of the programming language can be called to put the MD5 values with the same last two bits into the same folder and name the folder with the last two bits of the MD5, forming the first folder.
[0064] After obtaining the file encoding in this embodiment, since the amount of data based on the file encoding is large, in order to improve the file storage efficiency, a preset string truncation algorithm is used to truncate a part of the encoding from the file encoding as the retrieval basis, thereby reducing the amount of data for data matching and improving the retrieval efficiency.
[0065] Step 103: If there is a first folder named with the encoding value, compare the first file encoding of the data item corresponding to the encoding value with the second file encoding of the data item existing in the first folder to determine whether there is a second file encoding that is the same as the first file encoding; if there is, delete the data item corresponding to the first file encoding; if not, add the data item corresponding to the first file encoding to the first folder.
[0066] Exemplarily, in this embodiment, according to the encoding value corresponding to each data item, first retrieve whether there is a folder named with the encoding value. In this embodiment, the retrieval is performed based on the last two bits of the MD5 value. If there is a folder named with the last two bits of the MD5 value, extract the MD5 values corresponding to several data items saved in the folder, and then compare the MD5 value of the data item to be stored with the MD5 values corresponding to several data items saved in the folder respectively to determine whether there is the same MD5 value. If there is, it means that the data is repeated, and the data to be stored is deleted. If not, the data item can be added to the folder.
[0067] Step 104: If there is no first folder named with the encoding value, add the data item corresponding to the encoding value to a preset folder, and name the preset folder according to the encoding value to obtain the first folder.
[0068] In this embodiment, the step mainly includes: moving the data item into the folder according to a preset file moving method; renaming the preset folder according to the encoding value corresponding to the data item through a preset file renaming algorithm to obtain a first folder named with the encoding value.
[0069] Exemplarily, in this embodiment, if there is no folder named with the encoding value, in this embodiment, based on the last two digits of the MD5 value for retrieval, if there is no folder named with the last two digits of the MD5 value, save the data to be stored into a preset folder, and rename the preset folder according to the last two digits of the MD5 value of the data to be stored.
[0070] In this embodiment, the folder is named according to the encoding value, and at the same time, the data item corresponding to the encoding value is saved in the folder, which ensures that during the data storage process, duplicate data is avoided from being saved through the encoding value, and waste of space resources is avoided.
[0071] Step 105: Generate a file address corresponding to each data item according to the first folder, and save the file encoding, dataset name, and annotation information corresponding to each data item in the deep learning dataset through a preset database for retrieval.
[0072] Exemplarily, in this embodiment, after the data item is saved to the folder, record the name of the folder where the data item is located, the path of the folder, the MD5 value corresponding to the data item, the annotation information of the data item itself, and the name of the dataset to which it belongs, and save the recorded data to the database, so that the database stores information such as the name of the dataset corresponding to the data item, annotation information, and the path of the folder, for subsequent retrieval.
[0073] Refer to Figure 2 , the implementation process of a deep learning dataset storage and retrieval method provided in this embodiment further includes steps 201 to 202, and the steps are as follows:
[0074] Step 201: Retrieve the database according to the preset annotation information and dataset name to obtain the file encoding of the data item to be retrieved, and generate a retrieval path according to the file encoding.
[0075] In this embodiment, this step mainly includes: continuously encoding and intercepting the obtained file encoding to obtain the encoding value corresponding to the file encoding, and then obtaining the storage path of the folder where the encoding value is located and the folder name of the folder according to the encoding value; generating the retrieval path of the data to be retrieved according to the storage path of the folder and the folder name of the folder in combination with the file encoding.
[0076] Exemplarily, in this embodiment, according to the specified annotation information and dataset name, the database is retrieved to obtain the MD5 values of all data items, forming a path set. In this embodiment, since the annotation information, dataset name, and MD5 value information corresponding to the data items are stored in the database, the MD5 value information corresponding to them can be obtained according to the annotation information and dataset name. Then, since the last two digits of the MD5 value are used as the folder name, all data items with the same last two digits of the MD5 value will be placed in the same folder. Therefore, after obtaining the MD5 value, the last two digits of the MD5 value are intercepted to further obtain the folder name, and then the path of the data item and the default storage path of the data + folder name + MD5 value set can be obtained.
[0077] In this embodiment, by storing the annotation information, file encoding, and dataset name of each data item in the database, data retrieval can be performed by inputting the annotation information, solving the technical problem that data cannot be retrieved using annotation information in the prior art.
[0078] Step 202: Retrieve a number of data items saved in the first folder according to the retrieval path to obtain the data item to be retrieved.
[0079] Exemplarily, in this embodiment, the API for reading files in the programming language is used to loop through the retrieval path to access and read the data items corresponding to the retrieval path.
[0080] In addition, this embodiment also provides a deep learning dataset storage and retrieval system. For the specific structural composition of the storage and retrieval system, please refer to Figure 3 , including a file encoding module 301, an encoding interception module 302, a file saving module 303, and a data saving module 304.
[0081] The file encoding module 301 is used to traverse each data item in the obtained deep learning dataset according to a preset hash algorithm and generate a first file encoding corresponding to each data item.
[0082] The encoding truncation module 302 is used to perform encoding truncation on the first file encoding of each data item, obtain the encoding value corresponding to each data item, and perform a search based on the encoding value to determine whether there is a first folder named with the encoding value.
[0083] The file saving module 303 is used to, if there is a first folder named with the encoding value, compare the first file encoding of the data item corresponding to the encoding value with the second file encoding of the data items existing in the first folder to determine whether there is a second file encoding that is the same as the first file encoding; if so, delete the data item corresponding to the first file encoding; if not, add the data item corresponding to the first file encoding to the first folder; if there is no first folder named with the encoding value, add the data item corresponding to the encoding value to a preset folder, and name the preset folder according to the encoding value to obtain the first folder.
[0084] The data saving module 304 is used to generate the file address corresponding to each data item according to the first folder, and save the file encoding, dataset name, and annotation information corresponding to each data item in the deep learning dataset through a preset database for retrieval.
[0085] In this embodiment, the file encoding module 301 includes a traversal unit and an encoding unit.
[0086] The traversal unit is used to loop through the obtained deep learning dataset to obtain a number of data items included in the deep learning dataset; the data items include videos, audios, pictures, and texts.
[0087] The encoding unit is used to perform a hash calculation on each data item in the number of data items according to a preset hash algorithm to obtain the MD5 value corresponding to each data item.
[0088] In this embodiment, the encoding truncation module 302 includes a truncation unit and a retrieval unit.
[0089] The truncation unit is used to perform encoding truncation on the first file encoding corresponding to each data item according to a preset truncation rule through a string truncation algorithm to obtain the encoding value corresponding to each data item.
[0090] The retrieval unit is used to perform a search in a preset database according to the encoding value to determine whether there is a folder named with the encoding value in the database.
[0091] Refer to Figure 4, the structural composition of a deep learning dataset storage and retrieval system provided in this embodiment further includes a path module 401 and a retrieval module 402.
[0092] The path module 401 is used to retrieve the database according to the preset annotation information and dataset name, obtain the file code of the data item to be retrieved, and generate a retrieval path according to the file code.
[0093] The retrieval module 402 is used to retrieve a number of data items saved in the first folder according to the retrieval path, and obtain the data item to be retrieved.
[0094] The present invention discloses a deep learning dataset storage and retrieval method and system, including traversing each data item in the deep learning dataset by using a preset hash algorithm to generate a file code corresponding to each data item one by one, ensuring the uniqueness of the data item. Then, after generating the file code corresponding to the data item, in order to improve the efficiency of data item storage, several codes are intercepted from the file code as the name of the folder where the data item is located, so that when saving the folder later, the retrieval result can be obtained by using part of the code for retrieval, improving the efficiency of file storage. And the folder is named by using the code value. If the code value of the newly added data item is the same as the name of the existing folder, it is determined that the data item is repeated, and then the data item is deleted, avoiding data duplication and reducing the space resource occupancy rate. Then, the database is used to save the annotation information, dataset name and file code of the data item to be retrieved. After obtaining the annotation information and dataset name of the data item to be retrieved, the file code of the data item to be retrieved can be determined. According to the file code, the address of the folder where it is saved can be determined, and then the data item to be retrieved can be obtained by searching in the folder according to the address, solving the technical problem that the annotation information cannot be retrieved in the prior art and causing data immobility, and realizing data retrieval.
[0095] The above specific embodiments have further detailed the purpose, technical solution and beneficial effects of the present invention. It should be understood that the above is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. In particular, it is pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for storing and retrieving deep learning datasets, characterized in that, Including: Traverse each data item in the deep learning dataset obtained according to a preset hash algorithm, and generate a first file code corresponding to each data item; Perform code truncation on the first file code of each data item respectively to obtain a code value corresponding to each data item, and perform a search based on the code value to determine whether there is a first folder named with the code value; If there is a first folder named with the code value, compare the first file code of the data item corresponding to the code value with the second file code of the data item existing in the first folder to determine whether there is a second file code that is the same as the first file code; if so, delete the data item corresponding to the first file code; if not, add the data item corresponding to the first file code to the first folder; If there is no first folder named with the code value, add the data item corresponding to the code value to a preset folder, and name the preset folder according to the code value to obtain the first folder; Generate a file address corresponding to each data item according to the first folder, and save the file code, dataset name, and annotation information corresponding to each data item in the deep learning dataset through a preset database for retrieval.
2. The method for storing and retrieving a deep learning dataset according to claim 1, wherein Also including: Retrieve the database according to the preset annotation information and dataset name to obtain the file code of the data item to be retrieved, and generate a retrieval path according to the file code; Retrieve several data items saved in the first folder according to the retrieval path to obtain the data item to be retrieved.
3. The method for storing and retrieving a deep learning dataset according to claim 1, wherein The step of traversing each data item in the deep learning dataset obtained according to a preset hash algorithm and generating a first file code corresponding to each data item includes: Circularly traverse the obtained deep learning dataset to obtain several data items included in the deep learning dataset; the data items include videos, audios, pictures, and texts; Perform hash calculation on each data item in the several data items according to a preset hash algorithm to obtain an MD5 value corresponding to each data item.
4. A method for storing and retrieving a deep learning dataset according to claim 1, characterized in that, The step of performing code truncation on the first file code of each data item respectively to obtain a code value corresponding to each data item and performing a search based on the code value includes: According to a preset truncation rule, perform code truncation on the first file code corresponding to each data item through a string truncation algorithm to obtain a code value corresponding to each data item; Retrieve from a preset database according to the code value to determine whether there is a folder named with the code value in the database.
5. A method for storing and retrieving a deep learning dataset according to claim 1, characterized in that, The step of adding the data item corresponding to the code value to a preset folder and naming the preset folder according to the code value to obtain the first folder includes: Move the data item into the folder according to a preset file moving method; Rename the preset folder according to the encoding value corresponding to the data item through a preset file renaming algorithm to obtain a first folder named with the encoding value.
6. The method for storing and retrieving a deep learning dataset according to claim 2, wherein The generating the retrieval path according to the file encoding includes: Perform encoding truncation on the obtained file encoding to obtain the encoding value corresponding to the file encoding, and then obtain the storage path of the folder where the encoding value is located and the folder name of the folder according to the encoding value; Generate the retrieval path of the data to be retrieved according to the storage path of the folder and the folder name of the folder in combination with the file encoding.
7. A deep learning dataset storage and retrieval system, characterized in that, It includes a file encoding module, an encoding truncation module, a file saving module, and a data saving module; The file encoding module is used to traverse each data item in the obtained deep learning dataset according to a preset hash algorithm and generate a first file encoding corresponding to each data item; The encoding truncation module is used to perform encoding truncation on the first file encoding of each data item respectively to obtain the encoding value corresponding to each data item respectively, and perform a retrieval according to the encoding value to determine whether there is a first folder named with the encoding value; The file saving module is used to, if there is a first folder named with the encoding value, compare the first file encoding of the data item corresponding to the encoding value with the second file encoding of the data item existing in the first folder to determine whether there is a second file encoding that is the same as the first file encoding; if it exists, delete the data item corresponding to the first file encoding; if it does not exist, add the data item corresponding to the first file encoding to the first folder; if there is no first folder named with the encoding value, add the data item corresponding to the encoding value to a preset folder and name the preset folder according to the encoding value to obtain the first folder; The data saving module is used to generate a file address corresponding to each data item according to the first folder and save the file encoding, dataset name, and annotation information corresponding to each data item in the deep learning dataset through a preset database for retrieval.
8. The deep learning dataset storage and retrieval system according to claim 7, characterized in that, It further includes a path module and a retrieval module; The path module is used to retrieve the database according to the preset annotation information and dataset name to obtain the file encoding of the data item to be retrieved, and generate a retrieval path according to the file encoding; The retrieval module is used to retrieve several data items saved in the first folder according to the retrieval path to obtain the data item to be retrieved.
9. A deep learning dataset storage and retrieval system according to claim 7, characterized in that, The file encoding module includes a traversal unit and an encoding unit; The traversal unit is used to loop through the obtained deep learning dataset to obtain several data items included in the deep learning dataset; the data items include videos, audios, pictures, and texts; The encoding unit is used to perform a hash calculation on each data item in the several data items according to a preset hash algorithm to obtain an MD5 value corresponding to each data item respectively.
10. A deep learning dataset storage and retrieval system according to claim 7, characterized in that, The encoding truncation module includes a truncation unit and a retrieval unit; The truncation unit is used to perform encoding truncation on the first file encoding corresponding to each data item through a string truncation algorithm according to a preset truncation rule, so as to obtain the encoding value corresponding to each data item; The retrieval unit is used to retrieve from a preset database according to the encoding value to determine whether there is a folder named with the encoding value in the database.
Citation Information
Patent Citations
Data Deduplication Using Chunk Files
CN107111460A
A file storage management method based on an education system and an electronic device
CN109408462A