Query method, device and equipment

By using prefix compression indexing technology in data query, the query statements and indexes are directly matched, the noise problem caused by full-text search is solved, the query accuracy and efficiency are improved, and the utilization of storage resources is optimized.

CN120336261APending Publication Date: 2025-07-18HUAWEI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410065191.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing data query method uses full-text search methods to cause a lot of noise in the query results, which is low in accuracy, making it difficult to effectively manage and query large-scale metadata.

Method used

The prefix compression index technology is adopted to directly locate the query target's identification through the matching of the query statement and the prefix compression index, reducing the inverted table interception operation, and improving the query accuracy.

Benefits of technology

Through prefix compression indexing technology, the noise during the query process is reduced, the accuracy and efficiency of data query are improved, and the utilization of storage resources is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336261A_ABST
    Figure CN120336261A_ABST
Patent Text Reader

Abstract

According to the query method, device and equipment, in the application, a query request is received, the query request is used for obtaining a query target from a file system, a query statement comprises data in the query target or a path of the query target, and the query statement comprises a wildcard character; and querying the prefix compression index based on the query statement to determine an identifier of the query target. The prefix compression index records the identifier of the file or the identifier of the catalogue, and the file or the catalogue meets any one of the following conditions: the path of the file or the catalogue comprises a catalogue file combination, and the file or the catalogue comprises a character combination. And obtaining an identifier of a query target from the prefix compression index according to the query statement, and obtaining and feeding back the query target to the user by utilizing the identifier of the query target. Compared with the prior art, the prefix compression index does not correspond to a single word any more, but is combined data, a certain specific prefix compression index can be positioned by directly utilizing the query statement, the identifier of the query target is obtained, the possibility of introducing noise in the query process is reduced, and the query accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a query method, apparatus, and device. Background Art

[0002] With the improvement of the storage capacity in computing devices, the data stored in the computing devices has increased significantly, and correspondingly, the scale of the metadata that needs to be managed will also increase accordingly. The increase in metadata also makes the difficulty of data query increase.

[0003] Currently, the common data query method is to use the full-text search method. The full-text search usually constructs an inverted index table in units of words, and each word corresponds to an inverted index table. The inverted index table record of any word contains one or more files that contain this word. When querying a file, the user can provide the statement that the file needs to contain to indicate querying the file that contains this statement. In order to query this file, the statement can be segmented to obtain one or more words, and then the inverted index tables corresponding to each word are found, and the intersection of the inverted index tables corresponding to each word is obtained to obtain the file that contains this one or more words. This file is the target to be queried. Although this method can improve the file query efficiency to a certain extent, there may be a large amount of noise in the query result, and the accuracy is relatively low. Summary of the Invention

[0004] A query method, apparatus, and device provided by an embodiment of this application are used to improve query accuracy.

[0005] In a first aspect, an embodiment of this application further provides a query method, and this query method can be executed by a query device. The query device can receive a query request, and the query device can be triggered by a user. The query request is used to obtain a query target from a file system, and the query target can be a target file or a target directory. The query request carries a query statement, and the query statement includes the data in the query target or the path of the query target, and the query statement includes a wildcard character.

[0006] After receiving the query request, the query device queries a prefix compression index based on the query statement to determine the identifier of the query target. The prefix compression index is an index provided by an embodiment of this application. For any prefix compression index, it corresponds to a directory file combination or a character combination. The directory file combination includes at least one file name and at least one directory name, or includes at least two directory names, and the character combination includes at least two characters.

[0007] The prefix compression index records the identifier of a file or a directory, and the file or directory satisfies any of the following conditions: the path of the file contains a directory file combination, the path of the directory contains a directory file combination, the file contains a character combination, or the directory contains a character combination.

[0008] The query statement can be processed to form a directory file combination or a character combination. According to this query statement, the query device can obtain the identifier of the query target from the prefix-compressed index, that is, the identifier of the file recorded in the prefix-compressed index includes the identifier of the target file, or the identifier of the directory recorded in the prefix-compressed index includes the identifier of the target directory. Then, the query target is obtained and fed back to the user using the identifier of the query target.

[0009] Through the above method, the index queried by the query device corresponds to combined data instead of a single word. When the query device determines the query target using the prefix-compressed index, it directly locates to one or more specific prefix-compressed indexes using the query statement, and then obtains the identifier of the query target, without the need to perform the operation of intersecting multiple inverted lists, reducing the possibility of introducing noise during the query process and improving the query accuracy.

[0010] In a possible implementation, the query device can also obtain first configuration information, which indicates the types of information included in the metadata. The types of information included in the metadata include some or all of the following: the identifier of the file, the path of the file, the identifier of the directory, the path of the directory; the query device can store the metadata of the file or directory based on the first configuration information.

[0011] Through the above method, the query device can adjust the types of information included in the metadata according to the indication of the first configuration information, ensuring that certain specific types of metadata required can be stored, while for other types of metadata not involved in the first configuration information, they can be not stored, which can effectively utilize the storage space and increase the utilization rate of storage resources.

[0012] In a possible implementation, the query device can also obtain second configuration information, which indicates to enable the prefix-compressed index. When it is determined that the prefix-compressed index needs to be started according to the second configuration information, the query device can construct the prefix-compressed index according to the second configuration information and the stored metadata of the file or directory.

[0013] Through the above method, the query device constructs the prefix-compressed index only when it is determined that the prefix-compressed index needs to be started. The construction of the index is relatively flexible, and it can be selected to construct or not construct the prefix-compressed index according to actual needs.

[0014] In a possible implementation, the second configuration information can also indicate the maximum level of the prefix-compressed index. The maximum level of the prefix-compressed index describes the total number of directory names and file names in the directory file combination, or describes the total number of characters in the character combination.

[0015] Through the above method, since a prefix compression index corresponds to a directory file combination or a character combination, the number of directory file combinations or character combinations determines the number of prefix compression indexes. The maximum level of the prefix compression index can effectively restrict the number of constructed prefix compression indexes, avoiding the prefix compression indexes from occupying too much storage space.

[0016] In a possible implementation, the prefix compression index includes at least one index entry.

[0017] For any index entry, when the index entry includes the identifier of a file, the index entry further includes the position of the file name of the file in the path of the file. Optionally, the index entry further includes the positions of other file names or directory names in the directory file combination in the path of the file.

[0018] When the index entry includes the identifier of a directory, the index entry further includes the position of the directory name of the directory in the path of the directory. Optionally, it further includes the positions of other file names or directory names in the directory file combination in the path of the directory.

[0019] Through the above method, the structure of the prefix compression index is relatively simple, facilitating the query device to accurately determine the identifier of the query target therefrom.

[0020] In a possible implementation, the prefix compression index includes at least one index entry.

[0021] For any index entry, when the index entry includes the identifier of a file, the index entry further includes the positions of one or more characters in the character combination in the file.

[0022] Through the above method, the prefix compression index has a relatively simple structure, facilitating the query device to locate the identifier of the query target therefrom.

[0023] In a possible implementation, in the scenario of querying a file using the path of the file or querying the path of a directory using the path of the directory, when the query device queries the prefix compression index to determine the identifier of the query target, the query device can determine the target prefix compression index from the prefix compression index according to the query statement, and the query statement includes the target directory file combination corresponding to the target prefix compression index.

[0024] After that, the query device determines the first target index entry from the target prefix compression index according to the query statement.

[0025] If the first target index includes the identifier of the target file, the identifier of the query target is the identifier of the target file. For the file names or directory names included in the target directory file combination, the first target index entry and the query statement satisfy some or all of the following:

[0026] The first position information included in the first target index entry is consistent with the position of the file name in the query statement, and the first position information is the position of the file name in the path of the target file.

[0027] The second position information included in the first target index entry is consistent with the position of the directory name in the query statement, and the second position information is the position of the directory name in the path of the target file.

[0028] If the first target index includes the identifier of the target directory, the identifier of the query target is the identifier of the target directory. For the file name or directory name included in the target directory file combination, the first target index entry and the query statement satisfy some or all of the following:

[0029] The third position information included in the first target index entry is consistent with the position of the file name in the query statement, and the first position information is the position of the file name in the path of the target directory.

[0030] The fourth position information included in the first target index entry is consistent with the position of the directory name in the query statement, and the fourth position information is the position of the directory name in the path of the target directory.

[0031] Through the above method, the target prefix compression index can be accurately determined using the query statement, and the first target index entry can be accurately determined using the positions of the file name and / or directory name in the query statement. During the entire query process, there is no need to perform intersection operations between prefix compression indexes, reducing the possible noise in the query process and ensuring the accuracy of the query.

[0032] In a possible implementation manner, in the scenario of querying a file using the data in the file, when the query device queries the prefix compression index to determine the identifier of the query target, the query device can determine the target prefix compression index from the prefix compression index according to the query statement, and the query statement includes the target character combination corresponding to the target prefix compression index.

[0033] After that, the query device determines the first target index entry from the target prefix compression index according to the query statement.

[0034] The first target index includes the identifier of the target file, and the identifier of the query target is the identifier of the target file. For the characters included in the target character combination, the first target index entry and the query statement satisfy:

[0035] The position information included in the first target index entry is consistent with the position of the character in the query statement, and the position information is the position of the character in the target file.

[0036] Through the above method, the target prefix compression index can be accurately determined using the query statement, and the first target index item can be accurately determined using the positions of the characters in the query statement. During the entire query process, there is no need to perform intersection operations between prefix compression indexes, reducing the possible noise in the query process and ensuring the accuracy of the query.

[0037] In a possible implementation, in addition to using the prefix compression index to determine the identifier of the query target, the query device can also use the inverted index to determine the query target. The specific process is as follows:

[0038] The query device can query the inverted index based on the query statement to determine the inverted indexes corresponding to each word in the query statement.

[0039] The query device determines the candidate targets that are included in all the inverted indexes corresponding to each word according to the inverted indexes corresponding to each word.

[0040] The query device determines the query target from the candidate targets according to the positions of each word in the query statement and the positions of the words recorded in the inverted indexes corresponding to each word in the candidate targets.

[0041] The query device feeds back the query target to the user.

[0042] Through the above method, when the query device determines the query target using the inverted index, in addition to performing intersection operations on the inverted index, it also determines the query target from the candidate targets using the positions of each word in the query statement. That is, after performing intersection operations on the inverted index, a further screening step is carried out to ensure that the query target can be accurately located eventually, so as to ensure that the correct query target can be fed back to the user.

[0043] In a possible implementation, for a word in the query statement, the query target satisfies:

[0044] The position of the word in the query statement is the same as the position of the word recorded in the inverted index corresponding to the word in the query target.

[0045] Through the above method, the positions of the words in the query target are consistent with those recorded in the inverted index, ensuring that the determined query target is the correct file or directory that the user needs to query.

[0046] In a possible implementation, before the search device queries the prefix compression index based on the query statement of the query target to determine the identifier of the query target, further screening can be performed to ensure that the query statement satisfies some or all of the following:

[0047] 1. The frequency of occurrence of a word in the query statement is not greater than the maximum frequency of occurrence of the word.

[0048] 2. The length of the query statement is not greater than the maximum length of the data in the file or the maximum length of the path.

[0049] 3. The position of the word in the query statement is the same as the preset position of the word.

[0050] Through the above method, screening is performed before querying the prefix compression index, which can reject some unprocessable query requests in advance and ensure the query efficiency.

[0051] In a second aspect, an embodiment of the present application further provides a query method, which can be executed by a query device. The beneficial effects can be referred to the relevant descriptions in the first aspect and will not be elaborated here.

[0052] The query device can query the inverted list based on the query statement to determine the inverted list corresponding to each word in the query statement. The query device determines the candidate targets included in all the inverted lists corresponding to each word according to the inverted lists corresponding to each word. The query device determines the query target from the candidate targets according to the position of each word in the query statement and the position of the word recorded in the inverted list corresponding to each word in the candidate target. The query device feeds back the query target to the user.

[0053] In a possible implementation manner, for a word in the query statement, the query target satisfies:

[0054] The position of the word in the query statement is the same as the position of the word recorded in the query target in the inverted list corresponding to the word.

[0055] In a possible implementation manner, the query device receives third configuration information, which indicates that the prefix compression index is not enabled. The query device constructs an inverted list based on the metadata of the file or directory. When it is determined that the prefix compression index is not enabled, the inverted list is constructed to ensure that subsequent queries of files or directories can be performed using the inverted list.

[0056] In a third aspect, an embodiment of the present application further provides a query device, which has the functions to implement the behaviors in the method examples of the first aspect or the second aspect. The beneficial effects can be referred to the descriptions in the first aspect and will not be elaborated here. The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In a possible design, the structure of the device includes a receiving module, an indexing module, and a feedback module. Optionally, it further includes a data processing module. These modules can execute the corresponding functions in the method examples of the first aspect. For specific details, refer to the detailed descriptions in the method examples and will not be elaborated here.

[0057] Fourthly, the present application also provides a computing device, which includes a processor and a memory, and may further include a communication interface. The processor executes the program instructions in the memory to execute the method provided by the first aspect or any possible implementation manner of the first aspect. Or the processor executes the program instructions in the memory to execute the method provided by the second aspect or any possible implementation manner of the second aspect. The memory is coupled to the processor and stores the necessary computer program instructions and data for determining the anomaly detection process.

[0058] The communication interface is used to communicate with other devices, such as obtaining a query request, first configuration information, second configuration information, third configuration information, feedback query target, etc.

[0059] Fifthly, the present application also provides a computing device, which includes an acceleration device and a processor. Optionally, it may further include a memory and a communication interface. The processor cooperates with the acceleration device to execute the method provided by the first aspect or any possible implementation manner of the first aspect. Or the acceleration device executes the method provided by the first aspect or any possible implementation manner of the first aspect. The memory is coupled to the processor and stores some computer program instructions and data required during the query process.

[0060] The communication interface is used to communicate with other devices, such as obtaining a query request, first configuration information, second configuration information, third configuration information, feedback query target, etc.

[0061] Or;

[0062] The processor and the acceleration device are configured to execute the method provided by the second aspect or any possible implementation manner of the second aspect. Or the acceleration device executes the method provided by the second aspect or any possible implementation manner of the second aspect.

[0063] The communication interface is used to communicate with other devices, such as obtaining a query request, third configuration information, feedback query target, etc.

[0064] Sixthly, the present application provides a computing device system, which includes at least one computing device. Each computing device includes a memory and a processor. The processor of at least one computing device is used to access the code in the memory to execute the method provided by the first aspect or any possible implementation manner of the first aspect, or the processor of at least one computing device is used to access the code in the memory to execute the method provided by the second aspect or any possible implementation manner of the second aspect.

[0065] Seventh aspect, the present application provides a computer-readable storage medium. When the computer-readable storage medium is executed by a computing device, the computing device executes the method provided in the foregoing first aspect or any possible implementation manner of the first aspect, or executes the method provided in the foregoing second aspect or any possible implementation manner of the second aspect. Computer program instructions are stored in the storage medium. The storage medium includes but is not limited to volatile memories such as random access memories, and non-volatile memories such as flash memories, hard disk drives (HDDs), and solid state drives (SSDs).

[0066] Eighth aspect, the present application provides a computing device program product. The computing device program product includes computer program instructions. When executed by a computing device, the computing device executes the method provided in the foregoing first aspect or any possible implementation manner of the first aspect, or executes the method provided in the foregoing second aspect or any possible implementation manner of the second aspect. The computer program product can be a software installation package. In the case where it is necessary to use the method provided in the foregoing first aspect or any possible implementation manner of the first aspect, or it is necessary to use the method provided in the foregoing second aspect or any possible implementation manner of the second aspect, the computer program product can be downloaded and executed on the computing device.

[0067] Ninth aspect, the present application further provides a computer chip. The chip is connected to a memory. The chip is used to read and execute the computer program instructions stored in the memory, and execute the method in the foregoing first aspect and each possible implementation manner of the first aspect, or execute the method in the foregoing second aspect and each possible implementation manner of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a schematic structural diagram of a query system provided by the present application;

[0069] Figure 2 It is a flowchart of a query method provided by the present application;

[0070] Figures 3A to 3D It is a schematic diagram of a configuration interface of information provided by the present application;

[0071] Figure 4 It is a schematic diagram of a computing device provided by the present application;

[0072] Figure 5 It is a schematic structural diagram of a query device provided by the present application;

[0073] Figures 6 to 7 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0074] Before describing a query method, device, and apparatus provided in an embodiment of the present application, some concepts related to the embodiment of the present application are clarified first.

[0075] (1), File system.

[0076] A file system is a structured form of data storage and organization. The file system organizes the data in a computing device using the concept of "files". Data for the same purpose is organized into different types of files according to the structural forms required by different application programs. Different suffixes are usually used to refer to different types, and a memorable name, that is, a "file name", is configured for each file. When the number of files is large, these files are grouped according to a certain division method, and each group of files is placed in the same directory (or called a folder). Moreover, in addition to files, a directory can also have a lower-level directory (referred to as a subdirectory or subfolder) below it. All files and directories form a tree structure. This tree structure has a special name: file system. There are many types of file systems. Common ones include FAT / FAT32 / NTFS of Windows, EXT2 / EXT3 / EXT4 / XFS / BtrFS, etc. of Linux. For easy searching, starting from the root node, going down through the directories level by level until the file itself, the directory names, subdirectory names, and file names are concatenated with special characters (for example, Windows / DOS uses "\", and Unix-like systems use " / "). Such a string of characters is called a file path. For example, " / etc / systemd / system.conf" in Linux or "C:\Windows\System32\taskmgr.exe" in Windows. The file path is the unique identifier for accessing a specific file. For example, D:\data\file.exe under Windows is the path of a file, which represents the file file.exe in the data directory under the D partition. Similar to files, a directory can be understood as a special "file", and a directory also has a directory path, which is the information for locating the directory. In the embodiment of the present application, the special characters that connect the directory names, subdirectory names, and file names in the directory path or file path are called delimiters. The embodiment of the present application does not limit the specific form of the delimiter.

[0077] (2), Inverted index.

[0078] An inverted list is also known as an inverted index or reverse index. An inverted list can be used in scenarios where records are retrieved based on the values of attributes. Here, a record can be understood as a carrier of one or more attributes. Depending on the scenario of retrieving records, the meanings represented by the record and the attributes can vary. For example, in the scenario of retrieving files based on the data in the files, the record can be a file, and the attribute can be the data in the file (such as one or more words or statements in the file). Another example is that in the scenario of retrieving directories based on directory paths or retrieving files based on file paths, the record can be understood as a directory path or a file path, and the attribute can be the directory name, sub-directory name, or file name in the directory path or file path.

[0079] In an inverted list, an index is established using the value of the attribute as the keyword, and the inverted list can indicate the records that contain the value of the attribute. The embodiments of the present application do not limit the specific presentation form of the inverted list.

[0080] The following lists two presentation forms of the inverted list:

[0081] 1) Bitmap.

[0082] A bitmap includes multiple bits, each bit corresponding to a record (such as in a file system scenario, a record can be understood as a file or a directory). The value of each bitmap represents whether the corresponding record contains the value of the attribute.

[0083] For example, in the scenario of querying files based on the data in the files, bitmaps can be established for some data in the files, and different data can correspond to different bitmaps. For any bitmap corresponding to data, the bitmap includes multiple bits, and the number of bits can be the same as the number of files in the file system. Each bit corresponds to a file, and the value of the bit represents whether the corresponding file contains the data. If the value of the bit is 1, it represents that the corresponding file contains the data; if the value of the bit is 0, it represents that the corresponding file does not contain the data.

[0084] Another example is that in the scenario of retrieving directories based on directory paths or retrieving files based on file paths, bitmaps can be established for some directory names, sub-directory names, or file names, and different directory names (sub-directory names, or file names) can correspond to different bitmaps. For any bitmap corresponding to a directory name, the bitmap includes multiple bits, and the number of bits can be the same as the total number of files and directories in the file system. Each bit corresponds to a file or a directory, and the value of the bit represents whether the path of the corresponding file or the path of the directory contains the directory name. If the value of the bit is 1, it represents that the path of the corresponding file or the path of the directory contains the directory name; if the value of the bit is 0, it represents that the path of the corresponding file or the path of the directory does not contain the directory name.

[0085] 2) Skip list.

[0086] The skip list includes information related to "records" containing the values of this attribute. That is, the skip list only carries information related to records containing the values of this attribute, and does not carry information related to records that do not contain the values of this attribute. The skip list established for the values of this attribute includes multiple index entries, and the number of index entries is the same as the number of records containing the values of this attribute. Each index entry may include the identification information of the record, and optionally, each index entry may also include the position of this attribute in the record.

[0087] For example, in the scenario of querying files using the data in the files, skip lists can be established for the data in some files, and different data can correspond to different skip lists. For any skip list corresponding to data, the skip list records the file identifier of the file, and optionally, also records the position of the data in the file. Another example is that in the scenario of finding directories using directory paths or finding files using file paths, skip lists can be established for some directory names, sub-directory names, or file names, and different directory names (sub-directory names, or file names) can correspond to different skip lists. For any skip list corresponding to a directory name, the skip list records the directory identifier, and optionally, can also record the position of the directory name in one or more directory paths. Or the skip list records the file identifier, and optionally, can also record the position in one or more file paths of the directory name. Among them, the directory identifier and the file identifier are a kind of information for identifying directories and files, and are the information required to obtain the directory and the file. In practical applications, the directory identifier and the file identifier can also be replaced by other information required to obtain the directory and the file.

[0088] The following takes the inverted index applicable to the scenario of finding directories using directory paths or finding files using file paths as an example to illustrate the structure of the skip list.

[0089] Suppose there are four files in the file system, and the file paths of the four files are respectively:

[0090] File 1: / abc / bcd / def / efg.

[0091] File 2: / abc / def / efg / bcd.

[0092] File 3: / def / efg / abc / bcd.

[0093] File 4: / def / efg / abc / bcd.

[0094] Then, there can be four inverted lists, and the four inverted lists are respectively:

[0095] Skip list 1 (the corresponding directory name or file name is abc): 1 / 1, 2 / 1, 3 / 3, 4 / 2.

[0096] Skip list 2 (the corresponding directory name or file name is bcd): 1 / 2, 2 / 4, 3 / 4, 4 / 3.

[0097] Skip list 3 (the corresponding directory name or file name is def): 1 / 3, 2 / 2, 3 / 1, 4 / 1.

[0098] Skip list 4 (the corresponding directory name or file name is efg): 1 / 4, 2 / 3, 3 / 2, 4 / 4.

[0099] Among them, in each index item of the skip list, the file identifier and the position of the directory name or file name corresponding to the skip list in the file path are separated by " / ". Before " / ", it is the file name, and after " / " is the position of the directory name or file name corresponding to the skip list in the file path.

[0100] It should be noted that the specific presentation form of the index items in the skip list in the foregoing examples is only for illustration, and the embodiments of the present application do not limit the presentation form of the index items. Any form that can record the file identifier and the position of the directory name or file name corresponding to the skip list in the file path is applicable to the embodiments of the present application.

[0101] (3), Prefix compression index, levels of prefix compression index.

[0102] The embodiments of the present application provide an index constructed based on paths (such as file paths and directory paths). For the convenience of description, this index is called a prefix compression index.

[0103] The prefix compression index can be used in scenarios where records are searched using values of multiple attributes. For the description of attributes and records, reference can be made to the foregoing content, which will not be elaborated here.

[0104] Different prefix compression indexes can be set for different combinations of attributes, and the combination of attributes includes values of multiple attributes. The prefix compression index of an attribute combination indicates the records containing the attribute combination. The embodiments of the present application do not limit the specific presentation form of the prefix compression index.

[0105] Similar to the inverted list, the prefix compression index can also be presented in the form of a bitmap or a skip list:

[0106] 1), Bitmap.

[0107] The bitmap includes multiple bits, each bit corresponding to a record (in the file system scenario, a record can be understood as a file or a directory), and the value of each bitmap represents whether the corresponding record contains the combination of attributes.

[0108] For example, in a scenario of querying a file using the data in the file, a bitmap can be established for some characters in the file (here, the characters are used to represent certain specific data included in the file or individual characters formed after word segmentation of the data in the file, such as words, characters, symbols, etc. included in the file). Different character combinations can correspond to different bitmaps, where a character combination contains multiple characters. For any bitmap corresponding to a character combination, the bitmap contains multiple bit positions, and the number of bit positions can be the same as the number of files in the file system. Each bit position corresponds to a file, and the value of the bit position represents whether the corresponding file contains the character combination. If the value of the bit position is 1, it represents that the corresponding file contains the character combination. If the value of the bit position is 0, it represents that the corresponding file does not contain the character combination.

[0109] It should be noted that the embodiments of the present application do not limit the specific form of any character in the character combination. For example, the character can be a letter, an English word, a Chinese character, a Chinese word, a punctuation mark, a Greek symbol, etc.

[0110] For another example, in a scenario of finding a directory using a directory path or finding a file using a file path, a bitmap can be established for some directory names, sub - directory names, or file names. Different directory - file combinations can correspond to different bitmaps, where a directory - file combination can include multiple directory names or can include at least one directory name and at least one file name.

[0111] For any bitmap corresponding to a directory - file combination, the bitmap contains multiple bit positions, and the number of bit positions can be the same as the total number of files and directories in the file system. Each bit position corresponds to a file or a directory, and the value of the bit position represents whether the path of the corresponding file or directory contains the directory name. If the value of the bit position is 1, it represents that the path of the corresponding file or directory contains the directory name. If the value of the bit position is 0, it represents that the corresponding file or directory does not contain the directory name.

[0112] 2) Skip list.

[0113] The skip list includes relevant information of the "record" containing this attribute combination. That is, the skip list only carries relevant information of the "records" containing this attribute combination and does not carry relevant information of the "records" that do not contain the attribute combination. The skip list established for this attribute combination includes multiple index entries, and the number of index entries is the same as the number of "records" containing this attribute combination. For example, each index entry can include the identification information of the record. Optionally, each index entry can also include the positions of some or all of the attributes in this attribute combination in the record.

[0114] For example, in a scenario where files are queried using data in the files, skip lists can be established for some character combinations in the files, and different character combinations can correspond to different skip lists. For any skip list corresponding to a character combination, the file identifier of the file is recorded in the skip list. Optionally, the positions of one or more characters in the character combination in the file are also recorded. Another example is that in a scenario where a directory is searched using a directory path or a file is searched using a file path, skip lists can be established for directory file combinations, and different directory file combinations can correspond to different skip lists. For any skip list corresponding to a directory file combination, the directory identifier and the positions of some or all of the directory file combination in one or more directory paths are recorded in the skip list, or the file identifier and the positions of some or all of the directory file combination in one or more file paths are recorded in the skip list. Among them, the directory identifier and the file identifier are a kind of information for identifying directories and files, which are the information required to obtain the directories and files. In practical applications, the directory identifier and the file identifier can also be replaced by other information required to obtain the directories and files.

[0115] Note: A subdirectory is a special "directory" and is a relative concept used in the description of the directory structure. A subdirectory refers to the next-level "directory" included in the directory.

[0116] The number of levels of the prefix compression index describes the number of corresponding attributes. The embodiments of the present application do not limit the value-taking method of the number of levels of the prefix compression index. For example, the number of levels of the prefix compression index is equal to the number of attributes recorded in the prefix compression index. Another example is that the number of levels of the prefix compression index is equal to the number of attributes recorded in the prefix compression index minus one. When the number of levels of the prefix compression index is equal to 0, the prefix compression index is the inverted list mentioned above.

[0117] Taking the prefix compression index applicable to the scenario of searching for a directory using a directory path or searching for a file using a file path as an example below, the structure and number of levels of the prefix compression index presented in the form of a skip list will be described.

[0118] Still taking the four files mentioned above as an example.

[0119] Assume that the number of levels of the prefix compression index is equal to the number of attributes recorded in the prefix compression index minus one. Then, the maximum number of levels of the prefix compression index is three. The " / ** / " in the directory name or file name corresponding to the prefix compression index is used to represent any character. The structure of the index item in the prefix compression index is similar to the structure of the index item of the inverted list mentioned above. The following is only an example where only the position of the last directory name in the directory file combination corresponding to the prefix compression index in the file path is recorded in the index item. Similarly, the embodiments of the present application do not limit the specific structure of the index item in the prefix compression index.

[0120] There are multiple first-level prefix compressed indexes, and two of them are listed below:

[0121] First-level prefix compressed index 1 (the corresponding directory file combination is abc / * / bcd): 1 / 2, 2 / 4.

[0122] First-level prefix compressed index 2 (the corresponding directory file combination is def / * / bcd)): 3 / 4, 4 / 3.

[0123] There are multiple second-level prefix compressed indexes, and one of them is listed below:

[0124] Second-level prefix compressed index 1 (the corresponding directory file combination is abc / bcd / * / efg): 1 / 4, 2 / 3.

[0125] There is one third-level prefix compressed index, and one of them is listed below:

[0126] Third-level prefix compressed index 1 (the corresponding directory file combination is abc / bcd / def / efg): 1 / 4.

[0127] As can be seen from the foregoing description, actually without considering the storage space occupied by the prefix compressed index, the number of the foregoing compressed indexes of different levels is related to the number of directory file combinations that can be constructed. If the storage space occupied by the prefix compressed index is considered, the number of the foregoing compressed indexes of different levels can be reduced according to actual needs. For example, the directory file combination corresponding to the foregoing compressed index must include the root directory, so that the prefix compressed index corresponding to the directory file combination that does not need to be constructed and does not include the root directory can be excluded.

[0128] (4), Metadata.

[0129] Metadata, also known as mediation data and relay data, is data about data, mainly information describing the properties of data, such as the address of the data, the modification record of the data, the size of the data, the creation date of the data, etc.

[0130] Here, taking the storage and organization of data in the form of a file system as an example, the metadata involved in this file system includes file metadata, directory metadata, etc.

[0131] The metadata of a file describes the properties of the data in the file, such as the file name, the path of the file (which can also be understood as the address of the file), the modification record of the data, the size of the file, the file identifier, etc. The metadata of a directory describes the attribute information of the directory, such as the directory name, the modification record of the directory, the directory path, the identification (ID) of the directory, the creation time of the directory, the modification time of the directory, the access time, the group to which it belongs, etc.

[0132] In the embodiments of the present application, users are allowed to define the types of information included in the metadata, and the query device can organize and store the metadata according to the types of information defined by the users.

[0133] (5) Word segmentation.

[0134] When performing file or directory queries, it is usually necessary to first perform word segmentation on the information related to the query target carried in the query request (in the embodiments of the present application, the information related to the query target is called a query statement), that is, to divide the query statement carried in the query request into words, and convert the query statement into one or more words. Among them, the situation of converting the query statement into one word is a special "word segmentation" situation, usually when the number of words in the information related to the query target carried in the query request is small (such as only including one character) and cannot be further divided.

[0135] The query statement includes but is not limited to: the data in the query target, and the path of the query target.

[0136] Usually, the path of the query target includes delimiters, and each directory name or file name after being separated by the delimiter can be understood as a word segment. For the data in the query target carried in the query request (which may include wildcards), it is necessary to perform word segmentation on the data in the query target. For example, if the query statement carried in the query request is "apple tree", then "apple tree" can be word-segmented and converted into "apple", "tree", "ping", "guo", etc.

[0137] It should be noted that in the embodiments of the present application, there is no special distinction between: the words formed after word segmentation and characters. The words formed after word segmentation can be understood as a type of character, and one of the manifestation forms of a character can be manifested as a word. In the embodiments of the present application, the meanings represented by words and characters are the same. Usually, in the description of the process of word segmentation of data or querying of an inverted index, the term "word" is used, and in the description of the foregoing compressed index, the term "character" is used.

[0138] As Figure 1 shown, it is a schematic structural diagram of a query system provided by the embodiments of the present application. The query system includes a query device 100 and a storage device 200.

[0139] The query device 100 can interact with users, can store the metadata of data (such as the metadata of a file or a directory) in the storage device 200 according to the instructions of the users, and can also process the query requests triggered by the users, and query files or directories from the storage device 200 based on a pre-constructed inverted index and / or prefix compression index.

[0140] Function 1. The query device 100 has a data processing function.

[0141] The data processing function of the query device 100 is mainly manifested in the processing of metadata and the construction of indexes.

[0142] 1) Processing of metadata.

[0143] The query device 100 faces the user and allows the user to configure the information types included in the metadata according to their own needs. After the query device 100 obtains the information types included in the metadata configured by the user, it processes the metadata of the data, and the information types included in the processed metadata are the same as those included in the metadata configured by the user. The query device 100 can store the processed metadata in the storage device 200.

[0144] 2) Construction of indexes.

[0145] After processing the metadata, the query device 100 can also construct indexes based on the metadata. In the embodiments of the present application, the query device 100 can construct two types of indexes, one is an inverted index, and the other is a prefix compression index.

[0146] From the foregoing descriptions of the inverted index and the prefix compression index, it can be seen that the inverted index and the prefix compression index are applicable to different query scenarios. The content included in the inverted index in different query scenarios is different, and the content included in the prefix compression index in different query scenarios is also different.

[0147] When constructing indexes, the query device 100 can construct both an inverted index and a prefix compression index applicable to the scenario of finding files using the data in the file, and an inverted index and a prefix compression index applicable to the scenario of finding directories using the directory path or finding files using the file path.

[0148] Among them, the characters in the file corresponding to the inverted index and the prefix compression index applicable to the scenario of finding files using the data in the file can be different or not completely the same. For example, the inverted index only corresponds to a certain character in the file, and the prefix compression index corresponds to a character combination, and the character combination includes at least two characters.

[0149] Among them, the file names or directory names corresponding to the inverted index and the prefix compression index applicable to the scenario of finding directories using the directory path or finding files using the file path can be different or not completely the same. For example, the inverted index only corresponds to a directory name or a file name in the file system, and the prefix compression index corresponds to a directory file name combination, and the directory file name combination includes at least two directory names in the file system or at least one directory name and at least one file name in the file system.

[0150] When building an index, the query device 100 can build an inverted index applicable to the scenario of finding files using the data in the files, and build a prefix compression index applicable to the scenario of finding directories using directory paths or finding files using file paths.

[0151] In addition, when building a prefix compression index, the query device 100 can build a prefix compression index in the form of a bitmap or a skip list respectively, or build a prefix compression index in the form of a bitmap or a skip list. Similarly, when building an inverted list, the query device 100 can build an inverted list in the form of a bitmap or a skip list respectively, or build an inverted list in the form of a bitmap or a skip list.

[0152] Function 2: The query device 100 has a data query function.

[0153] The query device 100 can process a query request triggered by a user, and this query request is used to request to query a target directory or a target file. For the convenience of description, the target directory or target file requested by the query request is called the query target. The query request may carry the data in the query target or the path of the query target.

[0154] It should be noted that in the embodiments of the present application, the query device 100 supports fuzzy queries, that is, the query request may carry the data in the query target or the path of the query target may contain wildcards. Among them, the wildcards can replace any characters, and the embodiments of the present application do not limit the specific form of the wildcards. For example, the wildcards can be "?", "*", spaces, etc. In other words, the data carried in the query request in the query target may be incomplete data, and the path of the query target carried in the query request may be an inaccurate path. Therefore, in the embodiments of the present application, the query device 100 can find the query target using incomplete data and inaccurate paths. For the case where the query request may carry an accurate path, the query device 100 can directly find the query target according to the accurate path. For the case where the query request may carry the complete data in the file (that is, the data that does not include wildcards), the query device 100 can directly query the index according to the complete data without segmenting the data, determine the identifier of the query target, and then find the query target.

[0155] After receiving the query request, the query device 100 can find the query target based on the built index. When the query request carries the data in the query target, the query device 100 can query the inverted index or the prefix compression index based on the data in the query target, and then determine the path of the query target, and then use the path to obtain the query target. Among them, the inverted index or the prefix compression index is an index built by the query device 100 applicable to the scenario of finding files using the data in the files.

[0156] When the query request carries the path of the query target (such as a path containing wildcards), the query device 100 can query the inverted index or the prefix compression index based on the path of the query target, and then determine the identifier of the query target, and further obtain the query target. Among them, the inverted index or the prefix compression index is an index constructed by the query device 100 and applicable to scenarios of finding files using file paths or querying directories using directory paths.

[0157] In addition, since the query device 100 can construct the inverted index or the prefix compression index in the form of a bitmap and a skip list, the embodiments of the present application will take the construction of an inverted list in the form of a bitmap and a skip list as an example for illustration.

[0158] When the inverted list is an index applicable to the scenario of finding files using the data in the file, the query device 100 queries the inverted list based on the data in the query target. The query device 100 can query the bitmap and the skip list synchronously; it can also only query the bitmap or the skip list. For example, when the shortest length of the data in the query target corresponding to the inverted list (that is, the shortest length of the inverted list corresponding to the words after word segmentation of the data in the query target) is greater than the first threshold, the query device 100 queries the bitmap. When the shortest length of the data in the query target corresponding to the inverted list is not greater than the first threshold, the query device 100 queries the skip list. The length of the inverted list describes the total number of identifiers of files and directories recorded in the inverted list.

[0159] When the inverted list is an index applicable to the scenario of finding files using file paths or querying directories using directory paths, the query device 100 queries the inverted list based on the path of the query target. The query device 100 can query the bitmap and the skip list synchronously; it can also only query the bitmap or the skip list. For example, when the shortest length of the path of the query target corresponding to the inverted list is greater than the second threshold, the query device 100 queries the bitmap. When the shortest length of the path of the query target corresponding to the inverted list is not greater than the second threshold, the query device 100 queries the skip list.

[0160] Before performing data query, the query device 100 can also perform pre-query filtering. The so-called "pre-query filtering" means determining whether the query request can be processed or determining whether there is a possibility of finding the query target before data query.

[0161] In the embodiments of the present application, the query device 100 can perform "pre-query filtering" from some or all of the following aspects:

[0162] Aspect 1: Word frequency.

[0163] For any term, there is always a maximum frequency of occurrence of the term in the file and a maximum frequency of occurrence of the term in the file path (or directory path). For example, the maximum number of occurrences of "apple tree" in a file is 10, that is, in any of the multiple files in the file system, "apple tree" appears at most 10 times. Another example, the maximum number of occurrences of "New File 1" in the file path (or directory path) is 2, that is, in the path of any file (or directory) in the file system, "New File 1" appears at most twice.

[0164] The query device 100 can filter the query request based on the word frequency. If it is determined that any word in the data of the query target carried in the query request or the path of the query target is greater than the maximum frequency of occurrence of the word, the query request is rejected and the query request is no longer processed. Among them, the maximum frequency of occurrence of the word is the maximum number of occurrences of the word in a single file in the file system (applicable to the scenario of finding files using the data in the file), and the maximum frequency of occurrence of the word is the maximum number of occurrences of the word segmentation in a single file path or directory path in the file system (applicable to the scenario of finding directories using the directory path or finding files using the file path).

[0165] Aspect two: The length of the data in the file or the length of the directory path (or file path).

[0166] For the data in any file in the file system (such as a statement in a file), there is always a maximum value for the length of the data. For the path of any file in the file system or the path of any directory, there is also always a maximum value for the length of the directory path (or file path).

[0167] If the length of the data or path in the query target carried in the query request exceeds the corresponding maximum value, then the query request can be rejected.

[0168] Aspect three: Word position.

[0169] For a certain word in the data in any file in the file system (such as a statement in a file), the position of the word in the data usually appears at a fixed position. For example, the words "le" and "ma" always appear at the end of the data. For the path of any file in the file system or the path of any directory, a certain directory name or file name in the directory path (or file path) usually appears at a fixed position. For example, the root directory always appears at the beginning of the directory path (or file path).

[0170] If the word position in the data of the query target carried in the query request is different from the preset position of the word, the query request can be rejected. If the word position in the path of the query target carried in the query request is different from the preset position of the word, the query request can be rejected. Among them, the preset position of the word is the position of the word in the data of the file in the file system (applicable to the scenario of finding a file using the data in the file), and the preset position of the word is the position of the word segmentation in the file path or directory path in the file system (applicable to the scenario of finding a directory using the directory path or finding a file using the file path).

[0171] The embodiments of the present application do not limit the specific form of the query device 100. The query device 100 can be a hardware device. The query device 100 can be a single computing device or a cluster including multiple computing devices. The query device 100 can also be a certain hardware component in the computing device, such as a processor in the computing device (such as a central processing unit (CPU), a data processing unit (DPU)), an offloading card, an acceleration card, etc. The query device 100 can be a software device. The query device 100 can be software for managing files, such as file management software, database management software, etc.

[0172] The storage device 200 is used to store data and metadata. The embodiments of the present application do not limit the specific type of the storage device 200. The storage device 200 can store data or metadata under the instruction of the query device 100, and can also transmit data to the query device 100 under the instruction of the query device 100. For example, when the query device 100 finds a file or a directory, it can execute the storage device 200 to transmit the file or directory to the query device 100 according to the identifier of the file or the description of the directory. The storage device 200 can be a computing device or a cluster of computing devices with storage functions. The storage device 200 can also be a component with storage functions in the computing device, such as a hard disk, a magnetic disk, etc. Any device with a data storage function is applicable to the embodiments of the present application.

[0173] The embodiments of the present application do not limit the deployment methods of the storage device 200 and the query device 100. The storage device 200 and the query device 100 can be deployed in the same computing device; for example, the query device 100 can be a processor in the computing device (such as a CPU, a DPU), and the storage device 200 is a hard disk in the computing device. The storage device 200 and the query device 100 can be deployed in different computing devices; for example, the storage device 200 and the query device 100 can be deployed in a storage system. The query device 100 is a computing node in the storage system (with data computing functions and undertaking the computing tasks in the storage system), and the storage device 200 is a storage node in the storage system (with data storage functions).

[0174] The following is combined with Figure 2 The query method provided by the embodiment of the present application is described. The query method includes two parts. One part is the metadata storage and index construction process, which can be specifically referred to in steps 201 to 204. The other part is the data query process, which can be specifically referred to in steps 205 to 213.

[0175] In this part, two methods can be used for data query. One method is to use the prefix compression index to implement data query, which can be specifically referred to in steps 207 to 209. The other method is to use the inverted list to implement data query, which can be specifically referred to in steps 210 to 213. These two methods are applicable to both the scenario of finding files using the data in the file and the scenario of finding directories using the directory path or finding files using the file path. In Figure 2 the shown embodiment, taking the use of the inverted list for data query in the scenario of finding files using the data in the file and the use of the prefix compression index for data query in the scenario of finding directories using the directory path or finding files using the file path as an example for description.

[0176] The method of using the prefix compression index for data query in the scenario of finding files using the data in the file is similar to the method of using the prefix compression index for data query in the scenario of finding directories using the directory path or finding files using the file path. The difference is only that the information corresponding to the prefix compression index in these two different scenarios is different. The information corresponding to the former scenario (the scenario of finding files using the data in the file) is the character combination, and the information corresponding to the latter scenario (the scenario of finding directories using the directory path or finding files using the file path) is the directory file combination. The principle of its data query is similar and will not be elaborated here.

[0177] In the scenario of finding directories using the directory path or finding files using the file path, the method of using the inverted list for data query is similar to the method of using the inverted list for data query in the scenario of finding files using the data in the file. The difference is only that the information corresponding to the inverted list in these two different scenarios is different. The information corresponding to the former scenario (the scenario of finding files using the data in the file) is the data, and the information corresponding to the latter scenario (the scenario of finding directories using the directory path or finding files using the file path) is a directory or a file. The principle of its data query is similar and will not be elaborated here.

[0178] Step 201: The query device 100 obtains the first configuration information provided by the user. The first configuration information indicates the types of information included in the metadata.

[0179] Metadata is used to describe data, and the types of information in metadata can cover various information that can characterize the attributes of the data. In actual applications, the more types of information included in the metadata, the larger the storage space occupied. From the perspective of saving storage space, some relatively important information can be included in the metadata. In addition, from the application scenario of the data itself, some information characterizing the data attributes is relatively important, while some information characterizing the data attributes can be ignored. For example, in a database scenario, information such as the modification record of the data, the address of the data, the data table name, or the file name is relatively important, and the number of times the data is modified can often be ignored.

[0180] In an embodiment of the present application, the query device 100 provides an interface for a user to configure the types of information included in the metadata. Through this interface, the user can transmit the first configuration information to the query device 100.

[0181] It should be noted that the interface provided by the query device 100 to the user refers to the function provided by the query device 100 to the user. The embodiment of the present application does not limit the specific manifestation form of this interface. This interface can be manifested as a dedicated command for providing the first configuration information, or this interface can also be manifested as a visual interface for the user.

[0182] As Figure 3A shown, it is a schematic diagram of a configuration interface provided by an embodiment of the present application. In this configuration interface, the user can check or enter the types of information included in the metadata. The user can respectively configure the types of information included in the metadata of the file and the metadata of the directory.

[0183] Step 202: The query device 100 obtains the second configuration information provided by the user. This second configuration information indicates whether to enable the prefix compression index and the maximum level of the prefix compression index.

[0184] The existence of the prefix compression index can achieve efficient data query, but the prefix compression index will occupy some storage space, and the larger the level of the prefix compression index, the higher the data query efficiency, and the larger the storage space occupied. In an embodiment of the present application, the user is allowed to select whether to enable the prefix compression index or not according to their actual needs. Further, if the user selects to start the prefix compression index, the user is also allowed to configure the maximum level of the prefix compression index.

[0185] In an embodiment of the present application, the query device 100 provides an interface for a user to configure the prefix compression index. Through this interface, the user can transmit the second configuration information to the query device 100.

[0186] Similar to the interface for the types of information included in the aforementioned provided configuration metadata, the interface for the configuration prefix compression index only describes a configuration function provided by the query device 100 for the user. The embodiments of the present application do not limit the specific manifestation form of this interface.

[0187] As Figure 3B shown, it is a schematic diagram of a configuration interface provided by an embodiment of the present application. In this configuration interface, the user can select whether to enable the prefix compression index and the maximum number of levels of the prefix compression index.

[0188] The embodiments of the present application do not limit the execution order of step 201 and step 202. Step 201 can be executed first and then step 202, or step 202 can be executed first and then step 201. Of course, in actual applications, the query device 100 can also execute step 201 and step 202 synchronously. The query device 100 allows the user to complete the configuration of the metadata information type and the prefix compression index at one time, that is, the user provides the first configuration information and the second configuration information synchronously.

[0189] As Figure 3C shown, it is a schematic diagram of a configuration interface provided by an embodiment of the present application. In this configuration interface, the user can not only check the types of information included in the metadata, but also select whether to enable the prefix compression index and the maximum number of levels of the prefix compression index.

[0190] In the foregoing description, the example is given that the user provides the first configuration information and the second configuration information. The query device 100 also allows the user to complete the configuration of other information. For example, the query device 100 allows the user to complete the resource configuration for data query. Among them, the resource for data query indicates the maximum resources that the query device 100 can occupy to implement data query. The data query resources include, but are not limited to: the number of processors, the number of processor cores, the size of the memory, the cluster settings (used to indicate the cluster or nodes in the cluster where the query device 100 needs to be deployed), and the backup settings (such as the data backup method).

[0191] In the embodiments of the present application, step 201 or step 202 is an optional step. The types of information included in the metadata can be pre-stored in the query device 100; the query device 100 can also default to start the prefix compression index, and the maximum number of levels of the prefix compression index preset is saved in the query device 100. The query device 100 can execute the subsequent steps by using the types of information included in the pre-stored metadata and / or the maximum number of levels of the prefix compression index.

[0192] Step 203: The query device 100 obtains the metadata of the data and saves the user's metadata in the storage device 200 based on the first configuration information.

[0193] The query device 100 provides a data interface for users. Through this data interface, users can transmit data to the query device 100. After the query device 100 obtains data through this data interface and generates metadata of the data, it can save the metadata according to the first configuration information, that is, retain the information types indicated by the first configuration information in the metadata and delete the information types indicated by the first configuration information.

[0194] Step 204: The query device 100 constructs an index for the user's metadata according to the second configuration information.

[0195] If the second configuration information indicates to enable the prefix compression index, the query device 100 can construct a prefix compression index for the metadata. If the second configuration information also indicates the maximum level of the prefix compression index, then the level of the prefix compression index constructed by the query device 100 is not greater than the maximum level of the prefix compression index indicated by the second configuration information. Optionally, the query device 100 can also construct an inverted index for the user's metadata.

[0196] If the second configuration information indicates not to enable the prefix compression index, the query device 100 can construct an inverted index for the user's metadata.

[0197] When the query device 100 executes step 204, the constructed prefix compression index can be a prefix compression index applicable to the scenario of querying files using the data in the files, or a prefix compression index applicable to the scenario of finding directories using directory paths or finding files using file paths, or can include a prefix compression index applicable to the scenario of querying files using the data in the files and a prefix compression index applicable to the scenario of finding directories using directory paths or finding files using file paths. In Figure 2 In the illustrated embodiment, an example is given where the prefix compression index constructed by the query device 100 is a prefix compression index applicable to the scenario of finding directories using directory paths or finding files using file paths.

[0198] In addition, the embodiments of the present application do not limit the existence form of the prefix compression index constructed by the query device 100. The query device 100 can construct prefix compression indexes existing in the form of skip lists and bit tables respectively, or can only construct a prefix compression index existing in the form of a skip list or a bit table.

[0199] Similarly, when the query device 100 executes step 204, the inverted index constructed may be an inverted index applicable to the scenario of querying files using the data in the files, or an inverted index applicable to the scenario of finding directories using directory paths or finding files using file paths. It may also be an inverted index that includes both an inverted index applicable to the scenario of querying files using the data in the files and an inverted index applicable to the scenario of finding directories using directory paths or finding files using file paths. In Figure 2 In the illustrated embodiment, an example is given where the inverted index constructed by the query device 100 is an inverted index applicable to the scenario of querying files using the data in the files.

[0200] In addition, the embodiments of the present application do not limit the existence form of the inverted list constructed by the query device 100. Here, it is assumed that the query device 100 constructs inverted lists in the form of skip lists and bitmaps respectively.

[0201] So far, the preparatory operations before the query device 100 executes data query have been completed. After constructing the index, the query device 100 can receive and process query requests.

[0202] Step 205: The query device 100 receives a query request triggered by the user. The query request carries a query statement, which may be the data in the query target or the path of the query target.

[0203] The embodiments of the present application do not limit the manner in which the query device 100 executes step 205. For example, the user can send the query request to the query device 100 through a computing device deployed on the user side. Alternatively, the user can directly interact with the query device 100, and the user can directly trigger the query device 100 to generate a query request.

[0204] As Figure 3D shown, a visual query interface provided by the query device 100 to the user in the embodiments of the present application is shown. In this query interface, the user can choose to query files using the data in the files and type the data in the query target at the corresponding position. The user can also choose to query the target file (or target directory) using the target file path (or target directory path) and type the path of the query target at the corresponding position.

[0205] The query device 100 supports fuzzy queries, that is, when the user types the data in the query target or the path of the query target, wildcards can be entered to replace any characters.

[0206] Step 206: The query device 100 performs pre-query filtering on the query request. This step 206 is an optional step. In actual applications, the query device 100 may also not perform pre-query filtering and directly execute the subsequent steps.

[0207] The query device 100 can perform pre-query filtering on the query request from some or all of the following parts.

[0208] Filter 1: Filter the query request based on word frequency.

[0209] If the data in the query target is carried in the query request, the data in the query target can be segmented to obtain one or more words. The maximum occurrence frequency of each word is stored in the query device 100, and the maximum occurrence frequency of any word is the maximum number of times the word appears in a single file in the file system.

[0210] For one of the words, the query device 100 determines whether the occurrence frequency of the word in the data of the query target carried in the query request is greater than the maximum occurrence frequency of the segmented word. If so, the query request is rejected; otherwise, the query request passes this filter and the query request can be processed.

[0211] If the path of the query target is carried in the query request, the path of the query target is segmented by delimiters to form each word. The maximum occurrence frequency of each word is stored in the query device 100, and the maximum occurrence frequency of any word is the maximum number of times the segmented word appears in a single file path or directory path in the file system.

[0212] For one of the words, the query device 100 determines whether the occurrence frequency of the word in the data of the query target carried in the query request is greater than the maximum occurrence frequency of the word. If so, the query request is rejected; otherwise, the query request passes this filter and the query request can be processed.

[0213] Filter 2: Filter the query request based on the length of the data in the file or the length of the directory path (or file path).

[0214] If the data in the query target is carried in the query request, the maximum length of the data is stored in the query device 100, and the maximum length of the data is the maximum value of the length of the data in a single file in the file system.

[0215] The query device 100 determines whether the length of the data of the query target carried in the query request is greater than the maximum length of the data. If so, the query request is rejected; otherwise, the query request passes this filter and the query request can be processed.

[0216] If the path of the query target is carried in the query request, the maximum length of the path is stored in the query device 100, and the maximum length of the path is the maximum value of the length of a single file path or directory path in the file system.

[0217] The query device 100 determines whether the path length of the query target carried in the query request is greater than the maximum path length. If so, the query request is rejected. Otherwise, the query request passes through this filter and the query request can be processed.

[0218] Filter three: Filter the query request based on the word position.

[0219] If the query request carries the data in the query target, the data in the query target can be segmented to obtain one or more words. Some or all of the preset positions of these words are stored in the query device 100. The preset position of any word is the position of the word in the data included in the file in the file system.

[0220] For one of the words, the query device 100 determines whether the position of the word in the data of the query target carried in the query request is different from the preset position of the word segmentation. If so, the query request is rejected. Otherwise, the query request passes through this filter and the query request can be processed.

[0221] If the query request carries the path of the query target, the path of the query target is segmented by delimiters to form each word. Some or all of the preset positions of these words are stored in the query device 100. The preset position of any word is the position of the word segmentation in the file path or directory path of the file system.

[0222] For any word, the query device 100 determines whether the position of the word in the data of the query target carried in the query request is different from the preset position of the word. If so, the query request is rejected. Otherwise, the query request passes through this filter and the query request can be processed.

[0223] When the query request passes through the query pre-filter, the subsequent steps can be executed. Otherwise, the query device 100 rejects the query request, and the query device 100 also informs the user that the query target cannot be queried.

[0224] By passing through the query pre-filter, some unprocessable query requests can be rejected in advance, avoiding the need to execute subsequent steps for these query requests, and effectively improving the data query efficiency.

[0225] Taking the query request carrying the path of the query target as an example below, the query device 100 processes the query request based on the prefix compression index will be described.

[0226] Step 207: The query device 100 queries the prefix compression index according to the path of the query target carried in the query request to determine the target prefix compression index. The path of the query target includes the target directory file combination corresponding to the target prefix compression index. That is, the path of the query target includes the directory name and file name in the target directory file combination.

[0227] Since the query device 100 has previously constructed one or more prefix compression indexes, the query device 100 needs to first locate the target prefix compression index that records the query target from the one or more prefix compression indexes.

[0228] Taking the specific example described for the prefix compression index before as an example, assume that the path of the query target carried in the query request is / abc / bcd / * / . Then, the query device 100 can query the prefix compression index (existing in the form of a skip list) corresponding to the directory file combination of abc and bcd, that is, the first-level prefix compression index 1.

[0229] Step 208: The query device 100 determines a first target index entry from the target prefix compression index according to the path of the query target, and the first target index entry records the identifier of the query target.

[0230] If the first target index entry includes the identifier of the target file, then some position information included in the first target index entry is consistent with the position of the directory name or file name in the query statement.

[0231] Specifically, for a file name in the target directory file combination (the file name can be the file name of the target file or the file name of other files except the target file), if the first target index entry includes the position of the file name in the path of the target file, then the position of the file name in the path of the target file included in the first target index entry is consistent with the position of the file name in the query statement. Taking the path of the query target as / abc / bcd / * / as an example, if the first target index entry includes the position of the file name bcd in the path of the target file, then the position of the file name abc in the path of the target file needs to be in the second place, which is consistent with the position of the file name bcd in the query statement.

[0232] For a directory name in the target directory file combination, if the first target index entry includes the position of the directory name in the path of the target file, then the position of the directory name in the path of the target file included in the first target index entry is consistent with the position of the directory name in the query statement. Taking the path of the query target as / abc / bcd / * / as an example, if the first target index entry includes the position of the directory name abc in the path of the target file, then the position of the directory name abc in the path of the target file needs to be in the first place, which is consistent with the position of the directory name abc in the query statement.

[0233] If the first target index entry includes the identifier of the target directory, then some position information included in the first target index entry is consistent with the position of the directory name or file name in the query statement.

[0234] Specifically, for a file name in the target directory file combination, if the first target index entry includes the position of the file name in the path of the target directory, then the position of the file name in the path of the target directory included in the first target index entry is the same as the position of the file name in the look-up statement.

[0235] For a directory name in the target directory file combination (the directory name can be the file name of the target directory or the directory name of other directories except the target directory), if the first target index entry includes the position of the directory name in the path of the target directory, then the position of the directory name in the path of the target directory included in the first target index entry is the same as the position of the target name in the look-up statement.

[0236] Therefore, each file or directory path recorded in the target prefix compression index determined in step 207 contains each file name or directory name in the path of the query target. The query device 100 needs to further analyze the target prefix compression index to determine the query target from each file or directory recorded in the target prefix compression index.

[0237] The query device 100 determines the first target index entry according to the position of the file name or directory name in the path of the query target in the path and the position of at least one file name or directory name included in the corresponding directory file combination recorded in the target prefix compression index in the file path.

[0238] Here, taking the position of a directory name included in the corresponding directory file combination recorded in the target prefix compression index in the file path as an example, the query device 100 can find an index entry in the target prefix index where the position of the recorded directory name in the file path is the same as the position of the directory name in the path of the query target, and this index entry is the first target index entry.

[0239] Step 209: The query device 100 obtains the query target according to the first target index entry and feeds back the query target to the user.

[0240] After determining the first target index entry, the query device 100 can determine the identifier of the query target from the first target index entry, and the query device 100 can obtain the query target from the storage device 200 according to the identifier of the query target.

[0241] Next, taking the query request carrying the data in the query target as an example, the process of the query device 100 processing the query request based on the inverted index is described.

[0242] Step 210: The query device 100 performs word segmentation on the data in the query target, and converts the data in the query target into one or more words. The embodiments of the present application do not limit the way in which the query device 100 performs word segmentation on the data in the query target. For example, the query device 100 may use the n-gram algorithm to perform word segmentation on the data in the query target, where n is the number of words obtained after word segmentation of the data in the query target.

[0243] If the data in the query target is only converted into one word, step 212 can be directly executed (the files recorded in the inverted list corresponding to this word are the candidate files). If the data in the query target is converted into multiple words, step 211 can be executed.

[0244] Step 211: The query device 100 determines the inverted list corresponding to each word, intersects the inverted lists corresponding to multiple words, and determines the candidate files recorded in each inverted list.

[0245] After obtaining multiple words, the query device 100 can determine the inverted list corresponding to each word from the pre-constructed inverted list. Intersecting the inverted lists corresponding to multiple words means finding the files recorded in all the inverted lists corresponding to these multiple words, and the files recorded in all the inverted lists corresponding to these multiple words are the candidate files.

[0246] Here, it is assumed that the query device 100 obtains two word segments, and the inverted lists corresponding to each word segment determined from the pre-constructed inverted list are Table 1 and Table 2 respectively. Table 1 records File 1, File 2, and File 3, and Table 2 records File 2, File 3, and File 4. Then, the candidate files obtained after intersecting Table 1 and Table 2 are File 3 and File 2.

[0247] Since the query device 100 constructs inverted lists in the form of skip lists and bitmaps, the query device 100 can perform step 211 on the skip list and the bitmap synchronously, that is, the query device 100 intersects the bitmaps corresponding to each word to determine the candidate files recorded in each inverted list. Synchronously, it also intersects the skip lists corresponding to each word to determine the candidate files recorded in each inverted list. If the candidate files can be determined at a faster speed by intersecting the bitmaps corresponding to each word, then the candidate files determined after intersecting the bitmaps corresponding to each word are used to execute the subsequent steps. If the candidate files can be determined at a faster speed by intersecting the skip lists corresponding to each word, then the candidate files determined after intersecting the skip lists corresponding to each word are used to execute the subsequent steps.

[0248] The query device 100 may also perform step 211 only on one of the skip list or the bitmap. For example, when the shortest length of the inverted list corresponding to the word segmentation is greater than the first threshold, the query device 100 queries the bitmap and performs step 211 on the bitmap. When the shortest length of the inverted list corresponding to the word segmentation is not greater than the first threshold, the query device 100 queries the skip list.

[0249] Step 212: The query device 100 determines the target file from the candidate files according to the position order of the words recorded in the inverted list in the candidate files and the positions of the words in the data of the query target.

[0250] In addition to recording the identifier of the file, any index item in the inverted list also records the position of the word corresponding to the inverted list in the file. For any candidate file, the query device 100 can determine whether the position order of the words recorded in the inverted lists corresponding to each word in the candidate file and the positions of each word in the data of the query target are consistent. If they are consistent, the candidate file is the target file; otherwise, the candidate file is not the target file.

[0251] Step 213: The query device 100 feeds back the target file to the user.

[0252] It should be noted that in the embodiments of the present application, only two query scenarios are used to illustrate the query method provided by the embodiments of the present application: querying data using a path (such as querying a file using the path of the file or querying a directory using the path of the directory), and querying a file using the data in the file. In fact, the query method provided by the embodiments of the present application is also applicable to other scenarios. For example, querying a directory using the data in a file, or querying a directory or a file using other information of a file or a directory. The specific implementation process is similar, that is, constructing a prefix compression index using a combination of some key information, and then querying the prefix compression index using the received query statement. Or constructing an inverted list using some key information, and then determining candidate targets by taking the intersection of the inverted lists of each key information included in the query statement, and determining the query target from the candidate targets using the key information in the query statement. For specific details, reference can be made to the foregoing content, and details are not described herein again.

[0253] In the embodiments of the present application, the query device 100 may be a hardware device. For example, the query device 100 may be a computing device. An acceleration device is deployed in the computing device, and the functions of the query device 100 may be jointly implemented by the processor and the acceleration device in the computing device.

[0254] The following describes the hardware components included in the computing device 10, such as Figure 4As shown, it is a schematic diagram of the structure of a computing device 10 provided in an embodiment of the present application, and the computing device 10 includes an I / O interface 130, a processor 110, a memory 120, and an acceleration device 150. The I / O interface 130, the processor 110, the memory 120, and the acceleration device 150 can be connected through a system bus, which can be a peripheral component interconnect express (PCIe) bus, or a compute express link (CXL), a universal serial bus (USB) protocol, or other protocol buses.

[0255] Figure 4 An example of one of the connection methods is shown. Figure 4 In the embodiment, the acceleration device 150 can be directly inserted into a card slot on the motherboard of the computing device 10 and exchange data with the processor 110 through the PCIe bus 140.

[0256] The I / O interface 130 is used to communicate with devices outside the computing device 10. For example, data (such as first configuration information, second configuration information, query request, user metadata or data) sent by a device outside the computing device 10 is received through the I / O interface 130, or a query target is fed back to a device outside the computing device 10 through the I / O interface 130.

[0257] The processor 110 is the computing core and control core of the computing device 10. It can be a central processing unit (CPU) or other specific integrated circuits. The processor 110 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0258] Memory 120 is generally used to store computer program instructions. Memory 120 can also be used to temporarily store metadata, data, indexes, etc. Memory 120 has the advantage of fast access speed. Memory 120 generally adopts dynamic random access memory (DRAM). In addition to DRAM, Memory 120 can also be other random access memories, such as static random access memory (SRAM), storage class memory (SCM), etc. In addition, Memory 120 can also be read only memory (ROM). For read only memory, for example, it can be programmable read only memory (PROM), erasable programmable read only memory (EPROM), etc. Memory 120 can also be a dual in-line memory module or dual inline memory module (DIMM), flash media (FLASH), hard disk drive (HDD), or solid state disk (SSD), etc.

[0259] Processor 110 is connected to Memory 120 through a double data rate (DDR) bus or other types of buses. Memory 120 is understood as the internal memory 120 of computing device 10, and internal memory 120 is also called main memory.

[0260] Processor 110 can execute all or part of the steps performed by query device 100100 in the embodiments shown by Figure 2 invoking the computer program instructions in this internal memory 120.

[0261] Although not shown, computing device 10 also includes a persistent memory, or there is a memory that computing device 10 can remotely access. Whichever persistent memory it is, this persistent memory can be used to store data, metadata, or built indexes.

[0262] Among them, the memory that the computing device 10 can remotely access can be a memory located outside the computing device 10 and connected to the computing device 10 through a network. This memory can be a volatile memory, such as RAM, DRAM, SCM, SRAM. It can also be a non-volatile memory, such as ROM, flash memory, HDD, SSD, SCM, etc.

[0263] The persistent memory included in the computing device 10 can be connected to the computing device 10 through the system bus. This memory can be a non-volatile memory such as ROM, flash memory, HDD, SSD, etc.

[0264] The acceleration device 150 is connected to the computing device 10. The acceleration device 150 can be an external device of the computing device 10; the acceleration device 150 can also be deployed inside the computing device 10, such as the acceleration device 150 is located on the motherboard or backplane of the computing device 10. Figure 4 It is a schematic diagram of the acceleration device 150 deployed inside the computing device 10.

[0265] The acceleration device 150 can be a module with data processing functions attached to the computing device 10, undertaking part of the functions of the computing device 10. That is to say, part of the functions of the computing device 10 are offloaded to the acceleration device 150, and the acceleration device 150 replaces the computing device 10 (such as the processor 110 in the computing device 10) to process data and execute part of the tasks, so as to relieve the pressure on the processor 110 in the computing device 10 and release the computing power of the processor 110.

[0266] The embodiments of the present application do not limit the specific functions undertaken by the acceleration device 150. For example, the acceleration device 150 can carry the query pre-filtering function to determine whether the query request can continue to be processed. If the query request fails the query pre-filtering, the acceleration device 150 can directly reject the query request; if the query request passes the query pre-filtering, the acceleration device 150 can forward the query request to the processor 110 of the computing device 10, and the processor 110 continues to process the query request. Another example is that the acceleration device 150 can undertake the data processing functions of the query device 100, complete the processing of metadata and the construction of indexes (after the acceleration device stores the metadata and constructs the indexes, it can inform the processor 110 in the computing device 10 of the storage locations of the metadata and indexes). The processor 110 in the computing device 10 undertakes the data query function of the query device 100 and processes the received query request.

[0267] For another example, the acceleration device 150 and the processor 110 in the computing device 10 can cooperate to implement the data query function of the query device 100. As for other functions of the query device 100, in this scenario, they can be implemented by the acceleration device 150 and / or the processor 110 in the computing device 10. The embodiments of the present application do not limit the implementation manner in which the acceleration device 150 and the processor 110 cooperate to implement the data query function. The following lists a possible implementation manner:

[0268] Implementation manner 1: The constructed prefix compression index is a prefix compression index applicable to scenarios of finding directories using directory paths or finding files using file paths, and the constructed inverted list is an inverted list for scenarios of querying files using data in files.

[0269] Then, for a query request carrying the path of the query target, the processor 110 of the computing device 10 processes the query request. The specific processing manner can refer to steps 207 to 209. For a query request carrying the data in the query target, the acceleration device 150 processes the query request. The specific processing manner can refer to steps 210 to 213.

[0270] Implementation manner 2: The constructed prefix compression index is a prefix compression index applicable to scenarios of finding directories using directory paths or finding files using file paths, and the directory file combination corresponding to the prefix compression index must include the root directory of the file system. The constructed inverted list is an inverted list applicable to scenarios of finding directories using directory paths or finding files using file paths.

[0271] Then, the computing device 10 can process a query request carrying the path of the query target. If the path of the query target carried by the query request includes the root directory, the processor 100 of the computing device 10 processes the query request. The specific processing manner can refer to steps 207 to 209. If the path of the query target carried by the query request does not include the root directory, the acceleration device 150 processes the query request. The specific processing manner is similar to the manner described in steps 210 to 213, with the difference being that for different types of query statements carried in the query request, the basic processing manner is similar: determining the inverted lists corresponding to each word segment (i.e., directory name or file name), taking the intersection of the inverted lists, determining candidate targets, and then determining the query target using the positions of each word segment in the path of the query target and the positions of the word segments recorded in the inverted lists in the candidate targets.

[0272] Implementation manner 3: The constructed prefix compression index is a prefix compression index applicable to scenarios of querying files using data in files, and the constructed inverted list is an inverted list applicable to scenarios of finding directories using directory paths or finding files using file paths.

[0273] Then, the computing device 10 can process a query request carrying data in a query target. If the data in the query target carried by the query request includes multiple word segments, the processor 110 of the computing device 10 processes the query request. The specific processing method is similar to steps 207 to 209, but the difference lies in the type of the query statement carried in the query request. The basic processing method is similar: determine the target prefix compression index corresponding to the character combination (i.e., the character combination formed by one or more words after word segmentation of the data in the query target) from the prefix compression index. Then, according to the positions of each word segment in the data in the query target and the word segment positions recorded in the target prefix compression index, determine the first target index entry, and further determine the identifier of the query target. Obtain the query target according to the identifier of the query target. If the path of the query target carried by the query request only includes one word segment, the acceleration device 150 processes the query request. The specific processing method can refer to steps 210 to 213.

[0274] Implementation method four: The constructed inverted index exists in the form of a skip list and a bitmap. The constructed inverted index can include the inverted index for the scenario of querying a file using the data in the file and / or the inverted index applicable to the scenario of finding a directory using a directory path or finding a file using a file path.

[0275] Suppose in the scenario of querying a file using the data in the file, when the shortest length of the inverted index corresponding to the words after the data in the query target in the query request is greater than the first threshold, the acceleration device 150 processes the query request. The specific processing method can refer to steps 210 to 213, where the acceleration device 150 executes step 211 on the bitmap. When the shortest length of the inverted index corresponding to the words after the data in the query target in the query request is not greater than the first threshold, the processor 100 of the computing device 10 processes the query request. The specific processing method is similar to steps 210 to 213, where the acceleration device 150 executes step 211 on the skip list.

[0276] Suppose in the scenario of finding a directory using a directory path or finding a file using a file path, when the shortest length of the inverted index corresponding to the words in the path of the query target in the query request is greater than the first threshold, the acceleration device 150 processes the query request. The specific processing method can refer to steps 210 to 213, where the acceleration device 150 executes step 211 on the bitmap. When the shortest length of the inverted index corresponding to the words after the data in the query target in the query request is not greater than the first threshold, the processor 110 of the computing device 10 processes the query request. The specific processing method is similar to steps 210 to 213, where the acceleration device 150 executes step 211 on the skip list, with the difference that the inverted index corresponds to a directory name or a file name, rather than a certain data in the file.

[0277] In addition, in the embodiments of the present application, the acceleration device 150 is allowed to process the query request in other ways. For example, the acceleration device 150 stores a finite state machine, and the finite state machine can be used to determine whether one or more words exist in a directory (or file), or determine whether one or more words exist in the path of a directory (or file). The acceleration device 150 can use the finite state machine to process the query request.

[0278] The structure of the acceleration device 150 will be described below. The acceleration device 150 includes a processor, which can be a data processing unit (DPU) 151, a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), or other processors with data processing functions. In Figure 4 the following, only the example where the processor included in the acceleration device 150 is the DPU 151 is used for illustration. Optionally, the acceleration device 150 further includes a memory 152, a power supply circuit, etc. The DPU 151 is connected to the memory 152 through a system bus, and the system bus can be a PCIe-based line, or a bus based on CXL, USB protocol, or other protocols.

[0279] The DPU 151 is the main computing unit of the acceleration device 150 and the core unit of the acceleration device 150. The DPU 151 undertakes the main functions of the acceleration device 150.

[0280] In Figure 4 the example shown, it is illustrated by taking the processor and / or the acceleration device 150 in the computing device 10 to implement the functions of the query device 100 as an example. In practical applications, a similar structure can also be applied to the storage device 200, that is, the storage device 200 can also exist in the form of the computing device 10, and the functions possessed by the storage device 200 (such as storing data or transmitting data under the instruction of the query device 100) can be implemented by the acceleration device 150.

[0281] Based on the same inventive concept as the method embodiment, the embodiments of the present application further provide a query device, which is used to execute the method executed by the query device 100 in the above method embodiment. As Figure 5 shown, the query device 500 includes a receiving module 501, an indexing module 502, and a feedback module 503. Specifically, in the query device 500, connections are established between the modules through a communication path.

[0282] A receiving module 501 is configured to: receive a query request for obtaining a query target from a file system, where the query request carries a query statement that includes data in the query target or the path of the query target, and the query statement includes wildcards.

[0283] An indexing module 502 is configured to: query a prefix-compressed index based on the query statement to determine the identifier of the query target, where the prefix-compressed index records the identifiers of files or directories, and the file or directory satisfies any of the following conditions: the path of the file includes a directory-file combination, the path of the directory includes a directory-file combination, the file includes a character combination, or the directory includes a character combination; and obtain the query target according to the identifier of the query target.

[0284] Wherein, the directory-file combination includes at least one file name and at least one directory name, or includes at least two directory names, and the character combination includes at least two characters.

[0285] A feedback module 503 is configured to: feedback the query target.

[0286] As a possible implementation manner, the apparatus further includes a data processing module 504.

[0287] The receiving module 501 obtains first configuration information that indicates the types of information included in the metadata, and the types of information included in the metadata include some or all of the following: the identifier of the file, the path of the file, the identifier of the directory, and the path of the directory.

[0288] The data processing module 504 stores the metadata of the file or directory based on the first configuration information.

[0289] As a possible implementation manner, the receiving module 501 obtains second configuration information that indicates enabling the prefix-compressed index.

[0290] The data processing module 504 constructs a prefix-compressed index according to the second configuration information and the stored metadata of the file or directory.

[0291] As a possible implementation manner, the second configuration information further indicates the maximum level of the prefix-compressed index, and the maximum level of the prefix-compressed index describes the total number of directory names and file names in the directory-file combination, or describes the total number of characters in the character combination.

[0292] As a possible implementation manner, the prefix-compressed index includes at least one index entry, the index entry includes the identifier of the file or directory, and the index entry further includes the position of the file name of the file in the path of the file or the position of the directory name of the directory in the path of the directory.

[0293] As a possible implementation, the prefix compression index includes at least one index entry, where the index entry includes the identifier of a file or the identifier of a directory, and the index entry further includes:

[0294] The positions of the characters in the character combination in the directory and the positions of the characters in the character combination in the file.

[0295] As a possible implementation, the index module 502 determines a target prefix compression index from the prefix compression index according to the query statement, where the query statement includes the target directory file combination corresponding to the target prefix compression index. Then, according to the query statement, the first target index entry is determined from the target prefix compression index.

[0296] If the first target index includes the identifier of the target file, the identifier of the query target is the identifier of the target file. For the file name or directory name included in the target directory file combination, the first target index entry and the query statement satisfy some or all of the following:

[0297] The first position information included in the first target index entry is consistent with the position of the file name in the query statement, and the first position information is the position of the file name in the path of the target file;

[0298] The second position information included in the first target index entry is consistent with the position of the directory name in the query statement, and the second position information is the position of the directory name in the path of the target file.

[0299] If the first target index includes the identifier of the target directory, the identifier of the query target is the identifier of the target directory. For the file name or directory name included in the target directory file combination, the first target index entry and the query statement satisfy some or all of the following:

[0300] The third position information included in the first target index entry is consistent with the position of the file name in the query statement, and the first position information is the position of the file name in the path of the target directory;

[0301] The fourth position information included in the first target index entry is consistent with the position of the directory name in the query statement, and the fourth position information is the position of the directory name in the path of the target directory.

[0302] As a possible implementation, the index module 502 can also perform queries using an inverted index. The index module 502 can query the inverted index based on the query statement to determine the inverted index corresponding to each word in the query statement. According to the inverted index corresponding to each word segment, the candidate targets included in all the inverted indexes corresponding to each word segment are determined; then, according to the positions of each word in the query statement and the positions of the words in the candidate targets recorded in the inverted index corresponding to each word, the query target is determined from the candidate targets.

[0303] As a possible implementation, for a word in a query statement, the query target satisfies the following:

[0304] The position of the word in the query statement is the same as the position of the word recorded in the inverted list corresponding to the word in the query target.

[0305] As a possible implementation, the indexing module 502 can also determine that the query statement satisfies some or all of the following:

[0306] The occurrence frequency of a word in the query statement is not greater than the maximum occurrence frequency of the word;

[0307] The length of the query statement is not greater than the maximum length of the data in the file or the maximum length of the path;

[0308] The position of the word in the query statement is consistent with the preset position of the word.

[0309] The division of modules in the embodiments of the present application is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present application, each functional module can be integrated in a processor, or can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0310] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a terminal device (which can be a personal computer, a mobile phone, or a network device, etc.) or a processor to execute all or part of the steps of the method in each embodiment of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0311] The present application also provides a computing device 600 as Figure 6 shown. The computing device 600 includes a bus 601, a processor 602, a communication interface 603, and a memory 604. The processor 602, the memory 604, and the communication interface 603 communicate with each other through the bus 601.

[0312] Among them, the processor 602 can be a CPU, or can also be other general-purpose processors, DSPs, ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0313] The memory 604 can adopt DRAM. In addition to DRAM, the memory 604 can also be other random access memories, such as SRAM, etc. Additionally, the memory 602 can also be a ROM. For read-only memories, for example, it can be PROM, EPROM, etc. The memory 604 can also be a flash memory medium, HDD, or SSD, etc.

[0314] The computer program instructions are stored in the memory 604, and the processor 602 executes the computer program instructions to execute the steps performed by the query device 100 in the method described above. Figure 2 The memory 604 can also include other software modules required for other running processes such as an operating system (such as multiple modules in the query device 500). The operating system can be LINUXTM, UNIXTM, WINDOWSTM, etc.

[0315] This application also provides a computing device system, and the computing device system includes at least one Figure 7 computing device 700 as shown. The computing device 700 includes a bus 701, a processor 702, a communication interface 703, and a memory 704. The processor 702, the memory 704, and the communication interface 703 communicate with each other through the bus 701. At least one of the computing devices 700 in the computing device system communicates with each other through a communication path.

[0316] Among them, for the specific types of the processor 702 and the memory 704, reference can be made to the relevant descriptions of the processor 602 and the memory 604, which will not be elaborated here. The processor 702 executes the computer program instructions stored in the memory 704 to execute some or all of the steps performed by the detection device 100 in the method described above. The memory can also include other software modules required for other running processes such as an operating system. The operating system can be LINUXTM, UNIXTM, WINDOWSTM, etc. Figure 2 At least one of the computing devices 700 in the computing device system establishes communication with each other through a communication network, and any one or any number of modules in the query device 500 run on each computing device 700.

[0317] The descriptions of the processes corresponding to the above respective drawings have their own focuses. For parts not detailed in a certain process, reference can be made to the relevant descriptions of other processes.

[0318] The descriptions of the processes corresponding to the above respective drawings have their own focuses. For parts not detailed in a certain process, reference can be made to the relevant descriptions of other processes.

[0319] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes computer program instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. Figure 2 The processes or functions described above.

[0320] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as an SSD).

[0321] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.

Claims

1. A query method, characterized in that, Including: Receiving a query request for obtaining a query target from a file system, where the query request carries a query statement including data in the query target or a path of the query target, and the query statement includes a wildcard; Querying a prefix compression index based on the query statement to determine an identifier of the query target, where the prefix compression index records identifiers of files or directories, and the file or the directory satisfies any one of the following conditions: the path of the file includes a directory file combination, the path of the directory includes the directory file combination, the file includes a character combination, or the directory includes the character combination; Wherein, the directory file combination includes at least one file name and at least one directory name, or includes at least two directory names, and the character combination includes at least two characters; Obtaining and feeding back the query target according to the identifier of the query target.

2. The method according to claim 1, characterized in that The method further includes: Obtaining first configuration information indicating information types included in metadata, where the information types included in the metadata include some or all of the following: identifier of a file, path of a file, identifier of a directory, path of a directory; Storing the metadata of the file or the directory based on the first configuration information.

3. The method according to claim 2, wherein The method further includes: Obtaining second configuration information indicating enabling the prefix compression index; Constructing the prefix compression index according to the second configuration information and the stored metadata of the file or the directory.

4. The method according to claim 3, wherein The second configuration information further indicates a maximum level of the prefix compression index, and the maximum level of the prefix compression index describes the total number of directory names and file names in the directory file combination, or describes the total number of characters in the character combination.

5. The method according to any one of claims 1 to 4, characterized in that The prefix compression index includes at least one index entry, where the index entry includes the identifier of the file or the identifier of the directory, and the index entry further includes a position of the file name of the file in the path of the file, or a position of the directory name of the directory in the path of the directory.

6. The method according to any one of claims 1 to 4, characterized in that The prefix compression index includes at least one index entry, and the index entry includes the identifier of the file or the identifier of the directory, and the index entry further includes: A position of the characters in the character combination in the directory, or a position of the characters in the character combination in the file.

7. The method according to claim 1 or 5, characterized in that, The querying the prefix compression index based on the query statement to determine the identifier of the query target includes: Determining a target prefix compression index from the prefix compression index according to the query statement, where the query statement includes a target directory file combination corresponding to the target prefix compression index; Determining a first target index entry from the target prefix compression index according to the query statement; If the first target index includes an identifier of a target file, the identifier of the query target is the identifier of the target file, and for the file name or directory name included in the target directory file combination, the first target index entry and the query statement satisfy some or all of the following: The first position information included in the first target index item is consistent with the position of the file name in the query statement, and the first position information is the position of the file name in the path of the target file; The second position information included in the first target index item is consistent with the position of the directory name in the query statement, and the second position information is the position of the directory name in the path of the target file; If the first target index includes the identifier of the target directory, the identifier of the query target is the identifier of the target directory. For the file name or directory name included in the target directory file combination, the first target index item and the query statement satisfy some or all of the following: The third position information included in the first target index item is consistent with the position of the file name in the query statement, and the first position information is the position of the file name in the path of the target directory; The fourth position information included in the first target index item is consistent with the position of the directory name in the query statement, and the fourth position information is the position of the directory name in the path of the target directory.

8. The method according to claim 1, wherein Querying the prefix compression index based on the query statement of the query target to determine the identifier of the query target includes: Querying the inverted list based on the query statement to determine the inverted lists corresponding to the words in the query statement; Determining the candidate targets included in all the inverted lists corresponding to the words according to the inverted lists corresponding to the words; Determining the query target from the candidate targets according to the positions of the words in the query statement and the positions of the words recorded in the inverted lists corresponding to the words in the candidate targets; Feeding back the query target to the user.

9. The method according to claim 8, wherein The step of according to the positions of the words in the query statement and the positions of the words recorded in the inverted lists corresponding to the words in the candidate targets includes: For a word in the query statement, the query target satisfies: The position of the word in the query statement is the same as the position of the word recorded in the inverted list corresponding to the word in the query target.

10. The method according to any one of claims 1 to 9, characterized in that, Before querying the prefix compression index based on the query statement of the query target to determine the identifier of the query target, it further includes: Determining that the query statement satisfies some or all of the following: The occurrence frequency of the word in the query statement is not greater than the maximum occurrence frequency of the word; The length of the query statement is not greater than the maximum length of the data in the file or the maximum length of the path; The position of the word in the query statement is consistent with the preset position of the word.

11. A query device, characterized in that, It includes: A receiving module, configured to: receive a query request for obtaining a query target from a file system, where the query request carries a query statement, the query statement includes data in the query target or the path of the query target, and the query statement includes wildcards; An indexing module, configured to: query a prefix-compressed index based on the query statement to determine the identifier of the query target, where the prefix-compressed index records the identifiers of files or directories, and the file or the directory satisfies any of the following conditions: the path of the file contains a directory-file combination, the path of the directory contains the directory-file combination, the file contains a character combination, or the directory contains the character combination; obtain the query target according to the identifier of the query target. Wherein, the directory-file combination includes at least one file name and at least one directory name, or includes at least two directory names, and the character combination includes at least two characters. A feedback module, configured to: feedback the query target.

12. The device according to claim 11, characterized in that, The apparatus further includes a data processing module. The receiving module is further configured to: obtain first configuration information, where the first configuration information indicates the types of information included in the metadata, and the types of information included in the metadata include some or all of the following: the identifier of the file, the path of the file, the identifier of the directory, and the path of the directory. The data processing module is configured to: store the metadata of the file or the directory based on the first configuration information.

13. The apparatus according to claim 12, wherein The receiving module is configured to: obtain second configuration information, where the second configuration information indicates to enable the prefix-compressed index. The data processing module is configured to: construct the prefix-compressed index according to the second configuration information and the stored metadata of the file or the directory.

14. The device according to claim 13, characterized in that, The second configuration information further indicates the maximum level of the prefix-compressed index, and the maximum level of the prefix-compressed index describes the total number of directory names and file names in the directory-file combination, or describes the total number of characters in the character combination.

15. The device according to any one of claims 11 to 14, characterized in that The prefix-compressed index includes at least one index entry, the index entry includes the identifier of the file or the identifier of the directory, and the index entry further includes the position of the file name of the file in the path of the file, or the position of the directory name of the directory in the path of the directory.

16. The device according to any one of claims 11 to 14, characterized in that The prefix-compressed index includes at least one index entry, and the index entry includes the identifier of the file or the identifier of the directory, and the index entry further includes: The position of the characters in the character combination in the directory, or the position of the characters in the character combination in the file.

17. The device according to claim 11 or 15, characterized in that, The indexing module is configured to: Determine a target prefix-compressed index from the prefix-compressed index according to the query statement, where the query statement includes the target directory-file combination corresponding to the target prefix-compressed index. Determine a first target index entry from the target prefix-compressed index according to the query statement. If the first target index includes the identifier of the target file, the identifier of the query target is the identifier of the target file, and for the file name or directory name included in the target directory-file combination, the first target index entry and the query statement satisfy some or all of the following: The first position information included in the first target index item is consistent with the position of the file name in the query statement, and the first position information is the position of the file name in the path of the target file; The second position information included in the first target index item is consistent with the position of the directory name in the query statement, and the second position information is the position of the directory name in the path of the target file; If the identifier of the target directory is included in the first target index, the identifier of the query target is the identifier of the target directory. For the file name or directory name included in the target directory file combination, the first target index item and the query statement satisfy some or all of the following: The third position information included in the first target index item is consistent with the position of the file name in the query statement, and the first position information is the position of the file name in the path of the target directory; The fourth position information included in the first target index item is consistent with the position of the directory name in the query statement, and the fourth position information is the position of the directory name in the path of the target directory.

18. The device according to claim 11, characterized in that, The index module is used for: Querying the inverted list based on the query statement to determine the inverted lists corresponding to the respective words in the query statement; Determining candidate targets included in all the inverted lists corresponding to the respective words according to the inverted lists corresponding to the respective words; Determining the query target from the candidate targets according to the positions of the respective words in the query statement and the positions of the words recorded in the inverted lists corresponding to the respective words in the candidate targets.

19. The device according to claim 18, wherein For a word in the query statement, the query target satisfies: The position of the word in the query statement is the same as the position of the word recorded in the inverted list corresponding to the word in the query target.

20. The device according to any one of claims 11 to 19, characterized in that The index module is further used for: Determining that the query statement satisfies some or all of the following: The occurrence frequency of the word in the query statement is not greater than the maximum occurrence frequency of the word; The length of the query statement is not greater than the maximum length of the data in the file or the maximum length of the path; The position of the word in the query statement is consistent with the preset position of the word.

21. A computing device, characterized in that, The computing device includes a processor and a memory; The memory is used for storing computer program instructions; The processor executes by calling the computer program instructions stored in the memory to execute the method according to any one of claims 1 to 10.

22. A computer-readable storage medium, characterized in that, When the computer-readable storage medium is executed by the computing device, the computing device executes the method according to any one of claims 1 to 10 above.

Citation Information

Cited By

  • Method and system for realizing compatible file hard link and S3 shallow copy based on distributed key-value pair

    CN120994615A

  • Implementation method and system of compatible file hard link and s3 shallow copy based on distributed key-value pair

    CN120994615B