A data management method, apparatus, device, and storage medium
By generating hierarchical change tables and full scale tables in the data search and analysis engine and searching based on these tables, the problem of inefficiency of traditional metadata retrieval methods in massive data scenarios is solved, and efficient data retrieval and management are achieved.
Patent Information
- Application Number
- CN202210606084.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-31
AI Technical Summary
In massive data scenarios, traditional metadata retrieval methods such as the find command are inefficient when there are many files, complex directory levels or complex query conditions, and cannot meet customer needs.
By receiving data change information uploaded by the metadata service cluster in the data search analysis engine, generating hierarchical change tables and full scale tables, and searching in these tables according to user's search requests, to improve data retrieval efficiency.
This method significantly improves the efficiency of data retrieval, especially when the file size is over 100 million or the directory level is complex, reducing the search time and meeting customers' needs for efficient data management.
Smart Images

Figure CN114968926B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of distributed storage clusters, and particularly relates to a data management method, apparatus, device, and storage medium. Background Art
[0002] Currently, in the scenario of massive data, users need to manage the massive data based on metadata. For example, to obtain a list of all picture files with the suffix jpeg (Joint Photographic Experts Group, an international image compression standard), a list of files with a file size greater than 10 GBytes (gigabytes), or a list of files created before a given date. Only after the user can quickly obtain the files that meet the conditions can the corresponding data be efficiently managed.
[0003] The process of finding target files through the find command on traditional Linux (operating system) is a simple metadata retrieval. However, this method is less efficient in scenarios with a large number of files, deep directory hierarchies, or complex query conditions, and cannot meet customer needs. For example, when the number of files reaches hundreds of millions, using the find command to search for files generally takes more than an hour, and it will take even longer if the directory hierarchy is complex. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a data management method, apparatus, device, and storage medium, which can improve the data retrieval efficiency and the competitiveness of the product. The specific solutions are as follows:
[0005] In a first aspect, the present application discloses a data management method, which is applied to a data search and analysis engine and includes:
[0006] Receiving data change information uploaded by a metadata service cluster;
[0007] Generating a hierarchical change table based on the data change information and generating a full-scale table based on all the data information stored in the data search and analysis engine;
[0008] Receiving a retrieval request sent by a user, and retrieving in the hierarchical change table according to a preset retrieval method based on the retrieval request to obtain a retrieval result;
[0009] Retrieving in the full-scale table based on the retrieval result to obtain target information corresponding to the retrieval request;
[0010] Returning the target information to a web page for the user to view.
[0011] Optionally, the receiving data change information regularly uploaded by the metadata service cluster includes:
[0012] Receive the metadata regularly uploaded by the metadata service cluster and the current path corresponding to the metadata; wherein, the metadata is the data with changes monitored by the metadata service cluster, and the form of uploading the current path is in the form of an inode number.
[0013] Optionally, generating a hierarchical change table based on the data change information includes:
[0014] Generate a hierarchical change table based on the metadata, the current path, the historical path, and the change time; wherein, each change of each metadata corresponds to a change item in the hierarchical change table.
[0015] Optionally, generating a full-scale table based on all the data information stored in the data search and analysis engine includes:
[0016] Generate a full-scale table based on the metadata stored in the data search and analysis engine and the corresponding current path;
[0017] When receiving the data change information uploaded by the metadata service cluster each time, update the full-scale table using the data change information.
[0018] Optionally, receiving a retrieval request sent by a user and retrieving in the hierarchical change table according to a preset retrieval method based on the retrieval request to obtain a retrieval result includes:
[0019] Receive a retrieval request sent by a user, and determine whether the retrieval condition in the retrieval request contains a path; wherein, the retrieval condition includes the path, file name, file size, and operation time;
[0020] If the retrieval condition contains the path, retrieve in the hierarchical change table according to the retrieval request and the preset retrieval method to obtain a retrieval result;
[0021] Correspondingly, retrieving in the full-scale table based on the retrieval result to obtain target information corresponding to the retrieval request includes:
[0022] If the path is not included in the retrieval condition, directly retrieve in the full-scale table based on the retrieval condition to obtain target information corresponding to the retrieval request.
[0023] Optionally, retrieving in the hierarchical change table according to a preset retrieval method based on the retrieval request to obtain a retrieval result includes:
[0024] Obtain the target inode number of the data to be retrieved in the retrieval request;
[0025] Retrieve item by item in the hierarchical change table using the target index node number;
[0026] If there is a directory to be moved in, retrieve it item by item again. If there is no such directory to be moved in or out, end the retrieval;
[0027] Generate a target retrieval formula based on the target index node number, the index node number corresponding to the directory to be moved in, and the index node number corresponding to the directory to be moved out.
[0028] Optionally, retrieving in the full scale table based on the retrieval result to obtain target information corresponding to the retrieval request includes:
[0029] Retrieve in the full scale table based on the target retrieval formula to obtain path information in the form of the corresponding index node number;
[0030] Find the directory name of the parent directory layer by layer based on the path information in the form of the index node number;
[0031] Concatenate the directory names of the parent directories layer by layer to obtain the target information.
[0032] In a second aspect, the present application discloses a data management device, which is applied to a data search and analysis engine and includes:
[0033] An information receiving module, configured to receive data change information uploaded by a metadata service cluster;
[0034] A first table generation module, configured to generate a hierarchical change table based on the data change information;
[0035] A second table generation module, configured to generate a full scale table based on all data information stored in the data search and analysis engine;
[0036] A request receiving module, configured to receive a retrieval request sent by a user;
[0037] A first retrieval module, configured to retrieve in the hierarchical change table based on the retrieval request according to a preset retrieval method to obtain a retrieval result;
[0038] A second retrieval module, configured to retrieve in the full scale table based on the retrieval result to obtain target information corresponding to the retrieval request;
[0039] An information viewing module, configured to return the target information to a web page for the user to view.
[0040] In a third aspect, the present application discloses an electronic device, including:
[0041] A memory, configured to store a computer program;
[0042] A processor for executing the computer program to implement the steps of the data management method as disclosed above.
[0043] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the data management method as disclosed above is implemented.
[0044] It can be seen that the present application provides a data management method, including: receiving data change information uploaded by a metadata service cluster; generating a hierarchical change table based on the data change information and generating a full-scale table based on all data information stored in the data search and analysis engine; receiving a retrieval request sent by a user, and performing a retrieval in the hierarchical change table according to a preset retrieval method based on the retrieval request to obtain a retrieval result; performing a retrieval in the full-scale table based on the retrieval result to obtain target information corresponding to the retrieval request; and returning the target information to a web page for the user to view. Thus, it can be seen that in the present application, the data search and analysis engine constructs a hierarchical change table through the received data change information and generates a full-scale table based on all data information, retrieves the target information in the hierarchical change table and the full-scale table according to the user's retrieval request, and improves the data retrieval efficiency and the competitiveness of the product. Description of the Drawings
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0046] Figure 1 It is a flowchart of a data management method disclosed in the present application;
[0047] Figure 2 It is a schematic diagram of a specific data management method disclosed in the present application;
[0048] Figure 3 It is a schematic diagram of a hierarchical change table disclosed in the present application;
[0049] Figure 4 It is a schematic diagram of a full-scale table disclosed in the present application;
[0050] Figure 5 It is a schematic diagram of data movement disclosed in the present application;
[0051] Figure 6 It is a flowchart of a specific data management method disclosed in the present application;
[0052] Figure 7 A specific data management method flowchart disclosed in this application;
[0053] Figure 8 A specific data management method flowchart disclosed in this application;
[0054] Figure 9 A retrieval schematic diagram disclosed in this application;
[0055] Figure 10 A retrieval result conversion schematic diagram disclosed in this application;
[0056] Figure 11 A data movement path schematic diagram disclosed in this application;
[0057] Figure 12 A data movement result schematic diagram disclosed in this application;
[0058] Figure 13 A schematic diagram of the structure of the data management device provided by this application;
[0059] Figure 14 A structural diagram of an electronic device provided by this application. Detailed implementation manners
[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0061] Currently, the process of finding target files through the find command on traditional Linux is a simple metadata retrieval. However, this method is less efficient in scenarios with a large number of files, deep directory hierarchies, or complex query conditions, and cannot meet customer needs. For example, when the number of files reaches hundreds of millions, using the find command to search for files generally takes more than an hour, and it will take even longer if the directory hierarchy is complex. Therefore, this application provides a data management method that can improve the retrieval efficiency of data and the competitiveness of products.
[0062] An embodiment of the present invention discloses a data management method. Refer to Figure 1 As shown, it is applied to a data search and analysis engine. The method includes:
[0063] Step S11: Receive data change information uploaded by the metadata service cluster.
[0064] In this embodiment, the data search and analysis engine (Elasticsearch) receives the data change information regularly uploaded by the metadata service cluster. Specifically, it receives the metadata regularly uploaded by the metadata service cluster and the current path corresponding to the metadata. The metadata is the data with changes monitored by the metadata service cluster, and the current path is uploaded in the form of an index node number. It can be understood that the metadata service (Metadata Server, MDS) is used to maintain file metadata and process different metadata requests from clients. Multiple MDSs form a metadata service cluster, and each MDS is responsible for a different subtree of the entire system file tree to form a distributed metadata service cluster. As Figure 2 shown, each data search and analysis engine receives the data change information uploaded by a metadata service. Each metadata service monitors the metadata on a subtree of the entire system file tree. When it monitors that the metadata has changed, it puts the changed metadata information into a staging area, and then regularly uploads all the changed metadata information during this time period to the corresponding data search and analysis engine at a preset time interval. It should be noted that the MDS is responsible for the client metadata IO (Input / Output). The entire file system file tree is responsible for by multiple MDSs. Each MDS reports the data change information of the files it is responsible for to Elasticsearch. The data change information includes the metadata of the file itself and the current path, and the current path is reported in the form of an inode (index node) number.
[0065] It can be understood that Elasticsearch (data search and analysis engine) is a distributed storage database and also a RESTful (application or design that meets constraints and principles) style search and data analysis engine. It encapsulates Lucene (an open-source library for full-text retrieval and search), provides a set of simple and consistent RESTful APIs (Application Program Interface) to help with storage and retrieval, and can achieve sub-minute retrieval efficiency for tens of billions of files.
[0066] Step S12: Generate a hierarchical change table based on the data change information and generate a full-scale table based on all the data information stored in the data search and analysis engine.
[0067] In this embodiment, after receiving the data change information, a hierarchical change table is generated based on the data change information, and a full-scale table is generated based on all the data information stored in the data search and analysis engine. It can be understood that the design of the above database table is the core point. The metadata reporting to elasticsearch and retrieving from elasticsearch both rely on this data structure. File operations are very complex and involve multiple operations. Therefore, a full-scenario solution needs to be designed for the database table design.
[0068] Specifically, a hierarchical change table is generated based on the metadata, the current path, the historical path, and the change time; a full-scale table is generated based on the metadata stored in the data search and analysis engine and the corresponding current path. For example, as Figure 3 shown, the hierarchical change table is used to record the directories in the file system whose hierarchy has changed; as Figure 4 shown, the full-scale table is used to record all files and directories in the file system. Specifically, the hierarchical change table contains id, ino (index node), name, last historical path (old_parent_path), current path (parent_path), and change time (mtime). If the path of a directory has changed multiple times, a record (i.e., a change item in the hierarchical change table) will be generated in the hierarchical change table every time the path of the above directory changes, that is, each change information of the directory that has changed is recorded in the hierarchical change table, so as to obtain each step change information of all changed directories.
[0069] Furthermore, a full-scale table is generated based on all the metadata stored in the data search and analysis engine and the corresponding current path. The full-scale table contains id, name, data size (size), parent directory index node (parent_ino), and current path (parent_path). It should be noted that whenever the data change information uploaded by the metadata service cluster is received, the corresponding metadata is differentiated in the full-scale table using the data change information, and the path information in the full-scale table is updated using the path information in the data change information, that is, the path information in the full-scale table is always the latest path information. It should be noted that id and parent_ino are unique and unchanging information.
[0070] It can be understood that in a distributed file system, files and directories are organized in a tree structure. When transferred to ES (Elasticsearch), the storage changes from a tree structure to a flattened structure. The purpose of metadata retrieval users is to locate files. Therefore, the most critical metadata is the absolute path of the file. However, both the distributed file system and the ordinary file system organize files in a tree structure. The metadata of files and directories themselves do not record path information. The hierarchical relationship is established through the tree structure, and the file system, especially the directory hierarchy, can change. For example, the directory name changes or the directory is MV (moved) to another directory, but the metadata of the subdirectories and files under this directory do not change. Therefore, it is not possible to simply store the metadata of files and the file path information in elasticsearch. Especially when the directory name changes or the directory hierarchy changes, the paths of all files under it will change. For example, a change in the directory hierarchy under the root directory will cause the paths of all files in the entire system to change, and billions or even tens of billions of data need to be updated, which is not easy to implement and has low efficiency. As Figure 5 shown, when the file file2_2_1 is moved to the directory dir2_1, the directory dir2_1 contains the original file file2_1_1 and the newly moved file file2_2_1. If the directory dir2 is moved to the directory dir3_1_1, the paths of the directory dir2 and the directories dir2_1 and dir2_2 contained in dir2 will all change.
[0071] Furthermore, to solve the problems of directory renaming in the file tree and a large number of document modifications when MV to different levels, by designing two tables, namely the full-scale table and the hierarchical change table, and the paths are stored in the form of inode numbers. Through these two tables, all operations on the metadata of the file system can be reported to ES, and it can make the update simple, improve the retrieval speed, and enhance the accuracy of retrieval results.
[0072] It should be noted that in distributed unstructured storage, unstructured storage includes three forms: files, objects, and big data. The corresponding protocols are NAS (Network Attached Storage), NFS (Network File System) / CIFS (a newly proposed protocol), S3 / Swift (a programming language), and HDFS (Hadoop Distributed File System). The above three forms are uniformly undertaken by the distributed file system, that is, unified with files as the base. Requests for files, objects, and big data are all converted into file requests. From the bottom layer, it is a distributed file storage system.
[0073] Step S13: Receive the retrieval request sent by the user, and perform a retrieval in the hierarchical change table according to the preset retrieval method based on the retrieval request to obtain a retrieval result.
[0074] In this embodiment, receive the retrieval request sent by the user and perform a retrieval in the hierarchical change table according to the preset retrieval method based on the retrieval request to obtain a retrieval result. It can be understood that the retrieval request sent by the user contains the user's retrieval conditions. Obtain the retrieval conditions in the retrieval request. When the retrieval condition is a path, perform a retrieval in the hierarchical change table according to the preset retrieval method to obtain a retrieval result. The user selects different retrieval conditions, constructs a retrieval statement, and then sends a retrieval request to elasticsearch. It should be noted that the retrieval result obtained after the retrieval in the hierarchical change table is a retrieval formula.
[0075] Step S14: Perform a retrieval in the full-scale table based on the retrieval result to obtain the target information corresponding to the retrieval request.
[0076] In this embodiment, after performing a retrieval in the hierarchical change table and obtaining a retrieval result, perform a retrieval in the full-scale table based on the retrieval result to obtain the target information corresponding to the retrieval request. Specifically, perform a retrieval in the full-scale table according to the retrieval formula obtained after the retrieval in the hierarchical change table, and the target information corresponding to the retrieval request obtained is the target path information.
[0077] It can be understood that elasticsearch provides rich clients, and the retrieval function from elasticsearch can be implemented through various clients. The core is to strictly construct the retrieval statement according to the database table design, and then execute the retrieval operation to obtain the retrieval result.
[0078] Step S15: Return the target information to the web page for the user to view.
[0079] In this embodiment, after obtaining the target information, the target information is returned to the web page so that the user can directly view the current search results on the web page. It can be understood that this solution mainly includes three processes: database table design, metadata reporting, and metadata retrieval. This solution proposes an implementation of distributed unstructured storage metadata retrieval based on elasticsearch. The distributed unstructured storage is based on a distributed file system, uses elasticsearch to store the metadata of the file system. When the client performs metadata I / O, each MDS actively reports the metadata changes it is responsible for to elasticsearch. When the user retrieves a file or directory, it directly retrieves from elasticsearch, so as to achieve efficient retrieval in the scenario of massive data, quickly retrieve the file list that meets the user's requirements, and assist the user in data management.
[0080] It can be seen that this application provides a data management method, including: receiving data change information uploaded by a metadata service cluster; generating a hierarchical change table based on the data change information and generating a full-scale table based on all the data information stored in the data search analysis engine; receiving a retrieval request sent by a user, and performing a retrieval in the hierarchical change table according to a preset retrieval method based on the retrieval request to obtain a retrieval result; performing a retrieval in the full-scale table based on the retrieval result to obtain target information corresponding to the retrieval request; returning the target information to the web page for the user to view. Thus, in this application, the data search analysis engine constructs a hierarchical change table through the received data change information and generates a full-scale table based on all the data information, and retrieves the target information in the hierarchical change table and the full-scale table according to the user's retrieval request, improving the data retrieval efficiency and the competitiveness of the product.
[0081] See Figure 6 As shown, this embodiment of the present invention discloses a data management method. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution.
[0082] Step S21: Receive data change information uploaded by a metadata service cluster.
[0083] Step S22: Generate a hierarchical change table based on the data change information and generate a full-scale table based on all the data information stored in the data search analysis engine.
[0084] Step S23: Receive a retrieval request sent by a user, and determine whether the retrieval condition in the retrieval request includes a path.
[0085] In this embodiment, a retrieval request sent by a user is received, and it is determined whether the retrieval condition in the retrieval request includes a path. Specifically, as Figure 7As shown, first, it is determined whether the search condition in the received search request sent by the user includes the search for a path. Since the search condition can include a path, a file name, a file size, and an operation time, if the search condition is not set to include a path, a search request can be directly constructed according to the current search condition and directly retrieved from the full-scale table to obtain the target result. If the search condition is set to include a path, the following steps are performed.
[0086] Step S24: If the search condition includes the path, obtain the target index node number of the data to be retrieved in the search request.
[0087] In this embodiment, if the search condition is set to include a path, obtain the target index node number of the data to be retrieved in the search request. For example Figure 7 if the directory to be retrieved in the search request is directory A, obtain the index node number of directory A. It can be understood that if the path is not included in the search condition, a direct search is performed in the full-scale table based on the search condition to obtain the target information corresponding to the search request.
[0088] Step S25: Use the target index node number to traverse and search item by item in the hierarchical change table.
[0089] In this embodiment, after obtaining the target index node number of the data to be retrieved in the search request, the target index node number is used to traverse and search item by item in the hierarchical change table. For example, old_parent_path and parent_path are sequentially searched in the hierarchical change table to obtain a document containing the index node number of directory A. It can be understood that if the search result is empty, the path search condition is converted, that is, converted to an ino containing A, and then a new search request is constructed and a search is performed in the full-scale table
[0090] Step S26: If there is a directory to be moved in, perform a traversal search item by item on the directory to be moved in again. If there is no directory to be moved in and no directory to be moved out, end the search.
[0091] In this embodiment, after obtaining a document containing the index node number of directory A, it is determined whether there is a directory to be moved in and / or a directory to be moved out in the above document. If there is a directory to be moved in, perform a traversal search item by item on the directory to be moved in again. If there is still a directory containing the index node number of directory A in the directory to be moved in, it is again determined whether there is a directory to be moved in and / or a directory to be moved out in this directory, and the search is ended until there is no directory to be moved in and / or a directory to be moved out.
[0092] Step S27: Generate a target search formula based on the target index node number, the index node number corresponding to the directory to be moved in, and the index node number corresponding to the directory to be moved out.
[0093] In this embodiment, a target retrieval formula is generated based on the target index node number, the index node number corresponding to the directory to be moved in, and the index node number corresponding to the directory to be moved out. It can be understood that the target retrieval formula includes the directory to be moved in and the directory to be moved out.
[0094] Step S28: Retrieve in the full-scale table based on the target retrieval formula to obtain target information corresponding to the retrieval request.
[0095] In this embodiment, the target retrieval formula is used to retrieve in the full-scale table to obtain target information corresponding to the retrieval request. For example, when the retrieval condition in the retrieval request sent by the user is set to a path, the obtained target information is path information. When the retrieval condition is set to other conditions, corresponding retrieval results will be obtained.
[0096] Specifically, for example, if the retrieval condition in the retrieval request sent by the received user is to retrieve all files in the retrieval path dir3_1, then retrieve dir3_1 from the hierarchical change table (inspurfs_mv index), and the results dir2 moved in and dir3_1_1 moved out can be obtained; then retrieve dir2 from the hierarchical change table, and the results dir2_2 moved out and dir1 moved in can be obtained; then retrieve dir1 from the hierarchical change table, and the result is empty to end the retrieval; then the final target retrieval formula is obtained: dir3_1 + dir2 - dir3_1_1 - dir2_2 + dir1, that is, retrieve "including dir3_1 or dir2 or dir1" and "not including dir3_1_1 and dir2_2", and after the retrieval is completed, the path is spliced and filtered and returned.
[0097] Step S29: Return the target information to the web page for the user to view.
[0098] For the specific content of the above steps S21, S22, and S29, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.
[0099] It can be seen that in the embodiment of the present application, data change information uploaded by the metadata service cluster is received; a hierarchical change table is generated based on the data change information, and a full-scale table is generated based on all the data information stored in the data search and analysis engine; a retrieval request sent by a user is received, and it is determined whether a retrieval condition in the retrieval request includes a path; if the retrieval condition includes the path, a target index node number of data to be retrieved in the retrieval request is obtained; the target index node number is used to traverse and retrieve item by item in the hierarchical change table; if there is a directory to be moved in, the directory to be moved in is traversed and retrieved item by item again, and if there is no directory to be moved in and no directory to be moved out, the retrieval ends; a target retrieval formula is generated based on the target index node number, the index node number corresponding to the directory to be moved in, and the index node number corresponding to the directory to be moved out; the target retrieval formula is used to perform a retrieval in the full-scale table to obtain target information corresponding to the retrieval request; and the target information is returned to the web page for the user to view, improving the retrieval efficiency of data and the competitiveness of the product.
[0100] See Figure 8 As shown, an embodiment of the present invention discloses a data management method. Compared with the previous embodiment, this embodiment further describes and optimizes the technical solution.
[0101] Step S31: Receive data change information uploaded by the metadata service cluster.
[0102] Step S32: Generate a hierarchical change table based on the data change information and generate a full-scale table based on all the data information stored in the data search and analysis engine.
[0103] Step S33: Receive a retrieval request sent by a user, and perform a retrieval in the hierarchical change table according to a preset retrieval method based on the retrieval request to obtain a retrieval result.
[0104] Step S34: Perform a retrieval in the full-scale table based on the target retrieval formula to obtain path information in the form of the corresponding index node number.
[0105] In this embodiment, a search is performed in the full table based on the target search formula to obtain the corresponding path information in the form of the index node number. It can be understood that since the search is performed by the index node number during the search process, the path information obtained after the search is completed is displayed in the form of the index node number, that is, when the target information of the path corresponding to the search request sent by the user is obtained, since the path stored in elasticsearch is unreadable, it must be converted by encapsulating the elasticsearch client, that is, converting the current unreadable path into a readable format, that is, the full path in the form of a directory name and a file name, where the path splicing is to find the directory name of the parent directory inode from elasticsearch, and then splice it out layer by layer.
[0106] Step S35: searching the directory name of the parent directory layer by layer based on the path information in the form of the index node number.
[0107] In this embodiment, the directory name of the parent directory is searched layer by layer based on the path information in the form of the index node number. Figure 9 As shown in the figure, the parent_path in the figure is in the form of an index node number. The parent_path of dir3_1_1 is / 10000000000 / 10000000022, so it can be obtained from Figure 4 In the full table shown, it is found that the name of the id number 10000000022 is dir4. According to the parent_path of dir4, that is, 10000000000, it can be obtained that the name of the id number 10000000000 is share, and the search ends.
[0108] Step S36: concatenate the directory names of the parent directory layer by layer to obtain target information.
[0109] In this embodiment, the directory name of the parent directory is concatenated layer by layer to obtain the target information. It is understandable that the retrieved name (i.e., the directory name of the parent directory) is concatenated layer by layer in order to obtain the final target information. It is understandable that after Elasticsearch responds to the retrieval results, it performs path concatenation conversion through the retrieval tool, returns the key metadata information to the user, and performs result path concatenation and other processing according to the hierarchical change table and the full table, to assist users in massive data management and efficiently develop data value. For example Figure 10 As shown, the names retrieved by dir3_1_1 are concatenated to obtain the parent_path of / share / dir4.
[0110] In a specific tree structure, such as Figure 11As shown, the first step is to move dir2_2 into dir4, the second step is to move dir1 into dir2_1, the third step is to move dir2 into dir3_1, and the fourth step is to move dir3_1_1 into dir4. After the above steps, we can get the following Figure 12 The tree structure diagram shown.
[0111] Step S37: Return the target information to the web page for the user to view.
[0112] For the specific contents of the above steps S31, S32, and S36, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.
[0113] It can be seen that the embodiment of the present application receives data change information uploaded by a metadata service cluster; generates a hierarchical change table based on the data change information and generates a full table based on all data information stored in the data search and analysis engine; receives a search request sent by a user, and searches the hierarchical change table according to a preset search method based on the search request to obtain a search result; searches the full table based on the target search formula to obtain the corresponding path information in the form of the index node number; searches for the directory name of the parent directory layer by layer based on the path information in the form of the index node number; concatenates the directory names of the parent directory layer by layer to obtain the target information; and returns the target information to a web page for user viewing, thereby improving data retrieval efficiency and product competitiveness.
[0114] See also Figure 13 As shown, the embodiment of the present application also discloses a data management device, which is applied to a data search and analysis engine, including:
[0115] An information receiving module 11 is used to receive data change information uploaded by the metadata service cluster;
[0116] A first table generating module 12, configured to generate a level change table based on the data change information;
[0117] A second table generating module 13, used to generate a full table based on all data information stored in the data search and analysis engine;
[0118] The request receiving module 14 is used to receive the search request sent by the user;
[0119] A first search module 15, configured to search the level change table according to a preset search method based on the search request to obtain a search result;
[0120] A second retrieval module 16, configured to search the full table based on the retrieval result to obtain target information corresponding to the retrieval request;
[0121] An information viewing module 17 for returning the target information to a web page for a user to view.
[0122] It can be seen that the present application provides a data management method, including: receiving data change information uploaded by a metadata service cluster; generating a hierarchical change table based on the data change information and generating a full-scale table based on all data information stored in the data search and analysis engine; receiving a retrieval request sent by a user, and performing a retrieval in the hierarchical change table according to a preset retrieval method based on the retrieval request to obtain a retrieval result; performing a retrieval in the full-scale table based on the retrieval result to obtain target information corresponding to the retrieval request; returning the target information to a web page for a user to view. Thus, in the present application, the data search and analysis engine constructs a hierarchical change table through the received data change information and generates a full-scale table based on all data information, retrieves the target information in the hierarchical change table and the full-scale table according to the user's retrieval request, improving the data retrieval efficiency and the competitiveness of the product.
[0123] In some specific embodiments, the information receiving module 11 specifically includes:
[0124] An information receiving unit for receiving metadata regularly uploaded by a metadata service cluster and the current path corresponding to the metadata; wherein, the metadata is data with changes monitored by the metadata service cluster, and the current path is uploaded in the form of an index node number.
[0125] In some specific embodiments, the first table generating module 12 specifically includes:
[0126] A hierarchical change table generating unit for generating a hierarchical change table based on the metadata, the current path, the historical path, and the change time; wherein, each change of each metadata corresponds to a change item in the hierarchical change table.
[0127] In some specific embodiments, the second table generating module 13 specifically includes:
[0128] A full-scale table generating unit for generating a full-scale table based on the metadata stored in the data search and analysis engine and the corresponding current path;
[0129] A full-scale table updating unit for updating the full-scale table with the data change information each time the data change information uploaded by the metadata service cluster is received.
[0130] In some specific embodiments, the request receiving module 14 specifically includes:
[0131] A request receiving unit for receiving a retrieval request sent by a user;
[0132] The judging unit is used to judge whether the search condition in the search request includes a path; wherein the search condition includes the path, file name, file size and operation time.
[0133] In some specific embodiments, the first retrieval module 15 specifically includes:
[0134] a target index node number acquisition unit, configured to acquire the target index node number of the data to be retrieved in the retrieval request if the retrieval condition includes the path;
[0135] A first traversal and retrieval unit, used for traversing and retrieving one by one in the level change table by using the target index node number;
[0136] A second traversal retrieval unit, configured to traverse and retrieve the moved-in directory again one by one if there is a moved-in directory;
[0137] A search ending unit, used for ending the search if the directory moved in and the directory moved out do not exist;
[0138] The target search formula generating unit is used to generate a target search formula based on the target index node number, the index node number corresponding to the moved-in directory, and the index node number corresponding to the moved-out directory.
[0139] In some specific embodiments, the second retrieval module 16 specifically includes:
[0140] A path information acquisition unit, configured to search the full table based on the target search formula to obtain the corresponding path information in the form of the index node number if the search condition does not include the path;
[0141] A directory search unit, used for searching the directory name of the parent directory layer by layer based on the path information in the form of the index node number;
[0142] The target information acquisition unit is used to concatenate the directory names of the parent directories layer by layer to obtain the target information.
[0143] In some specific embodiments, the information viewing module 17 specifically includes:
[0144] The information viewing unit is used to return the target information to the web page for the user to view.
[0145] Furthermore, an embodiment of the present application also provides an electronic device. Figure 14 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.
[0146] Figure 14 This is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the data management method disclosed in any of the foregoing embodiments. Additionally, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0147] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.
[0148] In addition, as a carrier for resource storage, the memory 22 can be a read-only memory, a random access memory, a disk, or an optical disc, etc., and the resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0149] Among them, the operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, and it can be Windows Server, Netware, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of completing the data management method executed by the electronic device 20 disclosed in any of the foregoing embodiments, may further include a computer program capable of completing other specific tasks.
[0150] Furthermore, an embodiment of the present application also discloses a storage medium in which a computer program is stored, and when the computer program is loaded and executed by a processor, the steps of the data management method disclosed in any of the foregoing embodiments are implemented.
[0151] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0152] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0153] The above has introduced in detail a data management method, apparatus, device and storage medium provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A data management method, characterized in that, applied to a data search and analysis engine, including: Receiving data change information uploaded by a metadata service cluster; Generating a hierarchical change table based on the data change information and generating a full-scale table based on all the data information stored in the data search and analysis engine; the hierarchical change table is a table for recording directories in the file system where the hierarchy has changed, and the full-scale table is a table for recording all files and all directories in the file system; Receiving a retrieval request sent by a user, and determining whether the retrieval condition in the retrieval request contains a path; wherein, the retrieval condition includes any one or more of the path, file name, file size, and operation time; If the retrieval condition contains the path, retrieving in the hierarchical change table according to a preset retrieval method based on the retrieval request to obtain a target retrieval formula, and retrieving in the full-scale table based on the target retrieval formula to obtain target information corresponding to the retrieval request; If the retrieval condition does not contain the path, directly retrieving in the full-scale table based on the retrieval condition to obtain the target information; Returning the target information to a web page for the user to view.
2. The data management method according to claim 1, characterized in that, The receiving data change information uploaded by the metadata service cluster includes: Receiving metadata regularly uploaded by the metadata service cluster and the current path corresponding to the metadata; wherein, the metadata is data whose changes are monitored by the metadata service cluster, and the current path is uploaded in the form of an inode number.
3. The data management method according to claim 2, characterized in that, The generating a hierarchical change table based on the data change information includes: Generating a hierarchical change table based on the metadata, the current path, the historical path, and the change time; wherein, each change of each metadata corresponds to a change item in the hierarchical change table.
4. The data management method according to claim 2, characterized in that, The generating a full-scale table based on all the data information stored in the data search and analysis engine includes: Generating a full-scale table based on the metadata stored in the data search and analysis engine and the corresponding current path; When receiving the data change information uploaded by the metadata service cluster each time, updating the full-scale table using the data change information.
5. The data management method according to any one of claims 2 to 4, characterized in that, The retrieving in the hierarchical change table according to a preset retrieval method based on the retrieval request to obtain a target retrieval formula includes: Obtaining the target inode number of the data to be retrieved in the retrieval request; Traversing and retrieving item by item in the hierarchical change table using the target inode number; If there is a directory that has been moved in, traversing and retrieving it item by item again, if there is no such directory that has been moved in, ending the retrieval; Generate a target search formula based on the target index node number, the index node number corresponding to the moved-in directory, and the index node number corresponding to the retrieved moved-out directory.
6. The data management method according to claim 5, wherein, the retrieving in the full-scale table based on the target search formula to obtain target information corresponding to the retrieval request includes: retrieving in the full-scale table based on the target search formula to obtain path information in the form of the corresponding index node number; layer-by-layer searching for the directory names of the parent directories based on the path information in the form of the index node number; layer-by-layer concatenating the directory names of the parent directories to obtain the target information.
7. A data management device, wherein, applied to a data search and analysis engine, includes: an information receiving module, configured to receive data change information uploaded by a metadata service cluster; a first table generating module, configured to generate a hierarchical change table based on the data change information; the hierarchical change table is a table for recording directories in the file system where the hierarchy has changed; a second table generating module, configured to generate a full-scale table based on all data information stored in the data search and analysis engine; the full-scale table is a table for recording all files and all directories in the file system; a request receiving module, configured to receive a retrieval request sent by a user and determine whether the retrieval condition in the retrieval request includes a path; wherein, the retrieval condition includes any one or more of the path, file name, file size, and operation time; a first retrieval module, configured to, if the retrieval condition includes the path, retrieve in the hierarchical change table based on the retrieval request according to a preset retrieval method to obtain a target search formula, and retrieve in the full-scale table based on the target search formula to obtain target information corresponding to the retrieval request; a second retrieval module, configured to, if the path is not included in the retrieval condition, directly retrieve in the full-scale table based on the retrieval condition to obtain the target information; an information viewing module, configured to return the target information to a web page for the user to view.
8. An electronic device, wherein, includes: a memory, configured to store a computer program; a processor, configured to execute the computer program to implement the steps of the data management method according to any one of claims 1 to 6.
9. A computer-readable storage medium, wherein, used to store a computer program; wherein, when the computer program is executed by a processor, it implements the data management method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Directory-level referral method for parallel NFS with multiple metadata servers
US20140188953A1
Metadata management method, system and medium
US20220027326A1