File retrieval method and device, equipment and storage medium
Through the kernel-state file operation event-driven zero-copy data capture mechanism and user-state dynamic hybrid indexing architecture, combined with natural language processing technology, the real-time and efficiency problems of existing file retrieval methods are solved, real-time update of file system metadata and multi-dimensional efficient retrieval.
Patent Information
- Application Number
- CN202510660281.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing file retrieval methods have problems such as insufficient real-time, low resource efficiency and single retrieval mode, which is difficult to meet the needs of efficient file retrieval.
It adopts the kernel-state file operation event-driven zero-copy data capture mechanism and user-state dynamic hybrid indexing architecture to update file system metadata information in real time, and supports multi-dimensional retrieval conditions, and combines natural language processing technology to realize file retrieval.
Real-time update of file system metadata information and multi-dimensional efficient retrieval, improve file retrieval efficiency and flexibility, and support natural language interaction.
Smart Images

Figure CN120256542A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the technical field of data retrieval, and in particular, to a file retrieval method, apparatus, device, and storage medium. Background Art
[0002] As is well known, file retrieval technology is a core basic function in the fields of operating systems and data management. However, the current mainstream file retrieval methods usually rely on retrieval tools such as the find tool, the mlocate tool, and the locate tool. Due to problems such as relying on a specific file system, relying on periodic full disk scans or real-time traversals, and large retrieval overheads in the above-mentioned retrieval tools, there are certain limitations in the existing file retrieval based on the index library constructed by real-time traversal and periodic scanning of the file system. Summary of the Invention
[0003] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide a file retrieval method, apparatus, device, and storage medium. By means of a zero-copy data capture mechanism driven by kernel-mode file operation events and a collaborative architecture of user-mode dynamic hybrid indexing, an efficient real-time file retrieval tool is provided, realizing the real-time update of file system metadata information and the efficient retrieval purpose of multiple retrieval dimensions. Thus, not only can the efficient file retrieval requirements be met, but also the file retrieval efficiency can be improved.
[0004] In a first aspect, the present invention provides a file retrieval method, which is applied to a computer device. The computer device is installed with an operating system, and the operating system includes a user-mode space and a kernel-mode space. The method includes: The user-mode space responds to a file retrieval instruction received at the current moment, determines a target retrieval strategy corresponding to the file retrieval instruction, and updates an event filtering strategy to the kernel-mode space; The kernel-mode space obtains target file operation events from all file operation events captured at the current moment based on the updated event filtering strategy; The user-mode space determines metadata information of a target file corresponding to the target file operation event, and updates a hybrid index database based on the metadata information. The updated hybrid index database includes a mapping relationship between the target file and the metadata information, and a mapping relationship between a historical file and historical metadata information; and the hybrid index database supports multiple types of retrieval conditions; The user-mode space performs a retrieval in the updated hybrid index database based on the target retrieval strategy and the file retrieval instruction to obtain a file retrieval result.
[0005] In combination with the first aspect, in a possible implementation manner, obtaining the target file operation event from all the file operation events captured at the current moment based on the updated event filtering policy includes: Performing event filtering at different levels on all the file operation events based on the different-level event filtering policies included in the updated event filtering policy, so as to screen out the target file operation event from all the file operation events; Wherein, the different-level event filtering policies include a first-level event filtering policy based on filtering by a preset path prefix, and a second-level event filtering policy based on filtering by a preset file type and a preset file suffix.
[0006] In combination with the first aspect, in a possible implementation manner, updating the hybrid index database based on the metadata information includes: When the hybrid index database includes an inverted index library, a prefix tree index library, and a metadata index library, Updating the keyword-file ID mapping relation table existing in the inverted index library based on the file ID and the file name keyword of the target file; Updating the file path-file ID mapping relation table existing in the prefix tree index library based on the file ID of the target file and the file path in the metadata information; Updating the metadata information-file ID mapping relation table existing in the metadata index library based on the metadata information and the file ID.
[0007] In combination with the first aspect, in a possible implementation manner, performing a search in the updated hybrid index database based on the target search policy and the file search instruction to obtain a file search result includes: Determining at least two target index libraries in the updated inverted index library, the updated prefix tree index library, and the updated metadata index library based on the search conditions corresponding to the file search instruction; Performing a collaborative search in the at least two target index libraries based on the target search policy to obtain the file search result.
[0008] In combination with the first aspect, in a possible implementation manner, updating the event filtering policy to the kernel space includes: Obtaining the preset format rule configuration file at the current moment; Parsing out the event filtering rules from the preset format rule configuration file, and performing a legality check on the event filtering rules; The event filtering rule that has passed the legality check is used as the updated event filtering policy and is updated in real time to the kernel space.
[0009] In combination with the first aspect, in a possible implementation manner, the method further includes: When the file retrieval instruction is a natural language retrieval instruction, based on a pre-trained target language processing model, the natural language retrieval instruction is processed into a structured language retrieval instruction; Wherein, the target processing model is obtained by training an NLP model based on the Transformer framework with sample natural language retrieval instructions in different file retrieval scenarios and sample structured retrieval instructions identified by each of the sample natural language retrieval instructions; and each of the sample natural language retrieval instructions has undergone data augmentation.
[0010] In combination with the first aspect, in a possible implementation manner, determining the target retrieval strategy corresponding to the file retrieval instruction includes: Parsing the retrieval conditions of the file retrieval instruction, and determining a target retrieval plan based on all the parsed retrieval conditions; Based on the mapping relationship between the retrieval plan and the retrieval strategy set in advance, determining the target retrieval strategy corresponding to the target retrieval plan.
[0011] In a second aspect, the present invention further provides a file retrieval device. The device includes: A strategy determination unit configured to determine, in the user space, the target retrieval strategy corresponding to the file retrieval instruction received at the current moment, and update the event filtering strategy to the kernel space; An event acquisition unit configured to, in the kernel space, obtain a target file operation event from all the file operation events captured at the current moment based on the updated event filtering strategy; An index library update unit configured to, in the user space, determine the metadata information of the target file corresponding to the target file operation event, and update the hybrid index database based on the metadata information; the updated hybrid index database includes the mapping relationship between the target file and the metadata information, and the mapping relationship between the historical file and the historical metadata information; and the hybrid index database supports multiple types of retrieval conditions; A file retrieval unit configured to, in the user space, perform a retrieval in the updated hybrid index database based on the target retrieval strategy and the file retrieval instruction to obtain a file retrieval result.
[0012] In a third aspect, the present invention further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0013] In a fourth aspect, the present invention further provides a computer-readable storage medium. On the computer-readable storage medium, a computer program is stored, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0014] The embodiments of the present invention provide a file retrieval method, apparatus, device, and storage medium. In the file retrieval method, when the user space responds to a file retrieval instruction received at the current moment, it first determines a target retrieval strategy corresponding to the file retrieval instruction and updates an event filtering strategy to the kernel space. The kernel space obtains target file operation events from all file operation events captured at the current moment based on the updated event filtering strategy. The user space further determines metadata information of the target file corresponding to the target file operation events, updates the hybrid index database based on the metadata information, and then performs a retrieval in the updated hybrid index database based on the target retrieval strategy and the file retrieval instruction to obtain a file retrieval result. Since the updated hybrid index database includes the mapping relationship between the target file and the metadata information, as well as the mapping relationship between the historical file and the historical metadata information; and the hybrid index database supports multiple types of retrieval conditions, in this way, through the zero-copy data capture mechanism driven by kernel space file operation events and the collaborative architecture of user space dynamic hybrid indexing, an efficient real-time file retrieval tool is provided, realizing the real-time update of file system metadata information and the efficient retrieval purpose of multiple retrieval dimensions. Thus, not only can the efficient file retrieval requirements be met, but also the file retrieval efficiency can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more apparent: Figure 1 One of the flow diagrams of the file retrieval method in an embodiment; Figure 2 Another flow diagram of the file retrieval method in an embodiment; Figure 3 Still another flow diagram of the file retrieval method in an embodiment; Figure 4 Yet another flow diagram of the file retrieval method in an embodiment; Figure 5 One more flow diagram of the file retrieval method in an embodiment; Figure 6It is a structural block diagram of a file retrieval device in an embodiment; Figure 7 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0016] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that, for the sake of convenience of description, only the parts related to the invention are shown in the drawings.
[0017] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and embodiments. Additionally, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The terms "first" and "second" in the description and claims of the embodiments of the present invention are used to distinguish different objects, rather than to describe a specific order of the objects.
[0018] As is well known, file retrieval technology is a core basic function in the fields of operating systems and data management. However, the current mainstream file retrieval methods usually perform retrieval based on retrieval tools, such as find tools, mlocate tools, and other tools like mlocate tools. Due to problems such as relying on a specific file system, relying on periodic full disk scans or real-time traversals, and large retrieval overheads in the above-mentioned retrieval tools, there are certain limitations in the existing file retrieval based on the index library constructed by real-time traversal and periodic scanning of the file system.
[0019] Exemplarily, the find tool traverses the file system in real time and checks the file attributes one by one, which is suitable for one-time retrieval but has poor retrieval performance.
[0020] The mlocate tool builds an index library by periodically scanning the file system. The index library update is delayed by about several hours. Although directly retrieving the index during retrieval makes the retrieval speed faster, its limitation is that the index library update period is fixed and cannot reflect the changes in the file system in real time.
[0021] The FSearch tool is a graphical fast file search tool based on GTK3. Although it can achieve fast search by pre - establishing a file index library, it cannot capture the dynamic changes of the file system, and the index update still needs to be triggered actively. Among them, GTK is an open - source, multi - platform graphical user interface (GUI) toolkit, whose full English name is GIMP Toolkit. The full Chinese and English name of GIMP is GNU Image Manipulation Program. GUN is the abbreviation of GNU's Not Unix, indicating that GNU is a Unix - like system but not a Unix operating system. Toolkit means toolkit or toolbox.
[0022] From this, it can be understood that the above - mentioned traditional retrieval tools have at least three technical bottlenecks: The first technical bottleneck is the real - time defect of traditional retrieval tools: relying on periodic full - disk scanning or real - time traversal, resulting in delayed index updates and being difficult to adapt to the dynamic change environment of the file system. The second technical bottleneck is the resource - efficiency bottleneck: full - disk scanning occupies a large amount of input / output (I / O) and central processing unit (CPU) resources, and retrieving metadata information based on the stat system call further increases the overhead. The third technical bottleneck is the single retrieval mode: only supporting retrieval based on structured conditions such as file name and file size, lacking natural - language interaction ability, and unable to parse file retrieval instructions for semantic requests such as "find Word files modified yesterday".
[0023] In today's digital age, the amount of data generated and processed by enterprises is huge. The traditional management method can no longer meet the efficient file retrieval needs. In related technologies, there are limitations such as low retrieval efficiency, untimely index updates, and single retrieval dimensions. Thus, a file retrieval system is needed to make the search and retrieval of information faster and more convenient, thereby significantly improving work efficiency.
[0024] To solve the above - mentioned technical problems, the present invention provides a file retrieval method, device, equipment, and storage medium. By providing an efficient real - time file retrieval tool through a zero - copy data capture mechanism driven by kernel - state file operation events and a collaborative architecture of user - state dynamic hybrid indexing, it realizes the real - time update of file system metadata information and the efficient retrieval purpose of multiple retrieval dimensions. Thus, it can not only meet the efficient file retrieval needs but also improve the file retrieval efficiency.
[0025] The following combinesFigures 1 to 7 A file retrieval method, apparatus, device, and storage medium for describing the present invention are provided. The file retrieval method can be applied to a computer device installed with an operating system, which includes a user space and a kernel space. The computer device can be a personal computer, a server, an embedded system, or other devices. The present invention does not make specific limitations in this regard. Further, the file retrieval method can also be applied to a file retrieval apparatus provided in the computer device, and the file retrieval apparatus can be implemented by software, hardware, or a combination of both. Hereinafter, taking the computer device as the execution subject of the file retrieval method as an example, the file retrieval method will be described.
[0026] To facilitate understanding of the file retrieval method provided by the embodiments of the present invention, hereinafter, the file retrieval method provided by the present invention will be described in detail through the following several exemplary embodiments. It can be understood that these several exemplary embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0027] In one embodiment, a file retrieval method is provided, as Figure 1 shown, the method includes the following steps 101 to 104.
[0028] Step 101: The user space responds to the file retrieval instruction received at the current moment, determines the target retrieval strategy corresponding to the file retrieval instruction, and updates the event filtering strategy to the kernel space.
[0029] Among them, the file retrieval instruction can be an instruction automatically generated after the user inputs a file retrieval statement, and this instruction is used to instruct the computer device to perform the corresponding file retrieval operation.
[0030] The input method of the file retrieval statement can be voice input, input on the device, or other methods. The present invention does not make specific limitations in this regard.
[0031] Exemplarily, when the user space of the computer device provides a visual human-computer interaction interface, the file retrieval instruction at the current moment can be received by the user inputting a file retrieval statement in the retrieval input area of the visual human-computer interaction interface at the current moment.
[0032] It can be understood that the user space may include a retrieval engine module. When the retrieval engine module responds to a file retrieval instruction received at the current moment, it first determines the target retrieval strategy corresponding to the file retrieval instruction. The strategy determination process here may include, but is not limited to, determining the database to be retrieved, the keywords to be retrieved, clarifying the logical relationship and search steps between the retrieval terms based on the analysis of the file retrieval instruction, so as to obtain the target retrieval strategy corresponding to the file retrieval instruction. Or, a corresponding file retrieval plan may also be formulated according to different metadata fields in the file retrieval instruction, and then the target retrieval strategy for implementing each file retrieval plan is further determined. The present invention does not specifically limit the specific determination process of the target retrieval strategy.
[0033] In addition, the user space can also respond to the event filtering strategy received at the current moment, update the event filtering strategy, and update the updated event filtering strategy to the kernel space in real time. The update process here can use the Berkeley Packet Filter (BPF) hash table shared by the kernel space to dynamically update the event filtering strategy to the kernel space, so that the updated event filtering strategy takes effect immediately in the kernel space.
[0034] It should be noted that the purpose of hot updating the event filtering strategy can be achieved through the BPF hash table, so as to quickly and specifically screen out useful target file operation events subsequently.
[0035] Step 102: The kernel space obtains the target file operation event from all the file operation events captured at the current moment based on the updated event filtering strategy.
[0036] Among them, each file operation event can specifically be an event of performing a specified operation on an initial file. The kernel space can specifically be the Linux kernel space.
[0037] Specifically, for the kernel space, it can capture all initial operation events in real time and send them to the user space to update the hybrid index database, or based on the updated event filtering strategy sent by the user space at a certain moment, filter out the target file operation events from all the initial operation events captured at the corresponding moment and send them to the user space to update the hybrid index database. Therefore, when the kernel space receives the updated event filtering strategy sent by the user space, it will filter all the initial operation events captured at the current moment. Specifically, the event capture module built in the kernel space can first capture or collect file operation events. The event capture module is responsible for capturing file operation events in real time and transmitting the filtered target file operation events to the user space through the BPF circular buffer.
[0038] Furthermore, for all file operation events captured or collected, the event capture module can further filter out irrelevant events using the updated event filtering policy, so as to obtain target file operation events that meet the actual requirements, and then push all target file operation events to the user space through the BPF ring buffer.
[0039] It should be noted that extended Berkeley Packet Filter (eBPF) programs are injected into the system call layer and the Virtual File System (VFS) layer in the kernel space, so that kernel key functions configured in advance can be dynamically hooked. In this way, when responding to an event acquisition instruction, the event capture module in the kernel space can capture in real time all file operation events that match the hooked kernel key functions. All file operation events include file creation events, file deletion events, file renaming events, file writing events, file permission modification events, file permission modification events, file owner modification events, file soft link creation events, file hard link creation events, and other file operation events.
[0040] Exemplarily, the kernel key functions related to file operation events are as follows: vfs_create: Monitor file creation events; vfs_unlink: Monitor file deletion events; vfs_rename: Monitor file name modification events; vfs_write: Monitor file writing events; sys_chmod: Monitor file permission modification events; sys_chown: Monitor file owner modification events; vfs_link: Monitor file hard link creation events; vfs_symlink: Monitor file soft link creation events.
[0041] Furthermore, an example of capturing file operation events of the hooked kernel key functions is as follows: vfs_create: Monitor file creation events. After the file is successfully created, encapsulate metadata information such as the file operation type event, the absolute path of the file, the file type, the inode number, the file size, the number of disk blocks occupied by the file, the owner of the file, the group to which the file belongs, the file permissions, the last access time, the last modification time, the last status change time, the file creation time, the number of hard links, and the identifier of the device where the file is located into a unified standardized event structure, and send it to the event filtering unit in the event capture module.
[0042] vfs_unlink: Monitor file deletion events. After a file is successfully deleted, encapsulate metadata information such as the file operation type event, the absolute file path, and the file type into an event structure and send it to the event filtering unit in the event capture module.
[0043] vfs_rename: Monitor file name modification events. After a file name is successfully modified, encapsulate metadata information such as the file operation type event, the absolute file path, the file type, the original absolute file path, and the last status change time into an event structure and send it to the event filtering unit in the event capture module.
[0044] vfs_write: Monitor file write events. Intercept the system call exit. After data is successfully written to a file, encapsulate metadata information such as the file operation type event, the absolute file path, the file type, the file size, the last modification time, and the last status change time into an event structure and send it to the event filtering unit in the event capture module.
[0045] sys_chmod: Monitor file permission modification events. After a file permission is successfully modified, encapsulate metadata information such as the file operation type event, the absolute file path, the file type, the file permission, and the last status change time into an event structure and send it to the event filtering unit in the event capture module.
[0046] sys_chown: Monitor file owner modification events. After a file owner is successfully modified, encapsulate metadata information such as the file operation type event, the absolute file path, the file type, the file owner, the file group, and the last status change time into an event structure and send it to the event filtering unit in the event capture module.
[0047] vfs_link: Monitor file hard link creation events. After a file hard link is successfully created, encapsulate the metadata information of the event as the information of the newly created hard link file. The content encapsulated in the event structure is the same as that of the file creation event and send it to the event filtering unit in the event capture module.
[0048] vfs_symlink: Monitor file soft link creation events. After a file soft link is successfully created, encapsulate the metadata information of the event as the information of the newly created soft link file. The content encapsulated in the event structure is the same as that of the file creation event, and at the same time add the absolute file path information of the target file of the soft link and send it to the event filtering unit in the event capture module.
[0049] In this embodiment, examples of relevant fields of the standardized event structure are as follows: event_type: Event type; absolute_path: Absolute file path; src_absolute_path: The absolute path of the file before modification; dst_absolute_path: The absolute path of the file after modification; softlink_absolute_path: The absolute path of the target file of the soft link; type: File type; inode: inode number; size: File size; blocks: The number of disk blocks occupied by the file; uid: The owner of the file; gid: The group to which the file belongs; mode: File permissions; access: The last access time; modify: The last modification time; change: The last status change time; birth: The file creation time; links: The number of hard links; device: The identifier of the device where the file is located.
[0050] In this way, the event filtering unit in the event capture module receives the updated event filtering policy using the BPF hash table and filters out irrelevant events from all captured or collected file operation events accordingly. Then, all the target file operation events obtained after filtering are sent to the event propagation unit in the event capture module. The event propagation unit pushes all the target file operation events to the user space through the BFP circular buffer.
[0051] The event propagation unit is responsible for transmitting valid target file operation events to the user space. File operation events continuously occur in the operating system. By accumulating multiple file operation events in the kernel space and submitting batch processing events, the switching frequency between the kernel space and the user space is reduced. The correct sorting of file operation events monitored in the kernel space is crucial. The BPF circular buffer is used to ensure the order of all target file operation events obtained based on the shared buffer, ensuring that all target file operation events are processed by the user space in the order of occurrence, and at the same time achieving zero-copy data synchronization.
[0052] It should be noted that the event capture module is a core component of the operating system, responsible for real-time capturing file operation events in the kernel space and transmitting all the filtered target file operation events to the index management module in the user space through an efficient event filtering policy, supporting dynamic configuration of the event filtering policy.
[0053] Step 103: The user space determines the metadata information of the target file corresponding to the target file operation event, and updates the hybrid index database based on the metadata information; the updated hybrid index database includes the mapping relationship between the target file and the metadata information, as well as the mapping relationship between the historical file and the historical metadata information; and the hybrid index database supports multiple types of retrieval conditions.
[0054] It should be noted that the user space may also include an index management module. The index management module receives the target file operation events sent by the event capture module to the application layer, and updates the file metadata information to the index management module in real time; in addition, as the core data processing unit of the operating system, the index management module is responsible for structuring and efficiently retrieving file data; and the index management module is based on the hybrid index structure to achieve real-time and fast associated retrieval of file metadata.
[0055] For the index management module, it first receives all the target file operation events sent by the kernel space from the BPF circular buffer, and then classifies and sorts all the target file operation events through the parallel update control mechanism of the sharding queue architecture, so as to ensure that a class of target file operation events corresponding to the same file are strictly processed in the order of occurrence, and the target file operation events corresponding to different classified files are processed in parallel, realizing parallel update of file-level serialization, operation-level merging and sharding-level parallelism; thus obtaining each processed target file; the number of target files is the same as and corresponds one by one to the classification number of the target file operation events.
[0056] After that, the hybrid index database is updated based on the metadata information of each target file; so that the updated hybrid index database includes not only the mapping relationship between each target file and its metadata information, but also the mapping relationship between all historical files and their respective historical metadata information.
[0057] It should be noted that in the case where a hybrid index database that supports conditional retrieval of multiple data types is pre-constructed in the index management module, the hybrid index database can be updated based on the metadata information of each target file; on the contrary, in the case where the hybrid index database is not retrieved in the index management module, the full-scale index database of the file system can be constructed for the first time, and then the hybrid index structure is involved in the full-scale index database to construct a hybrid index database that supports conditional retrieval of multiple data types, and then the hybrid index database is updated based on the metadata information of each target file.
[0058] The hybrid index database can support conditional retrieval of multiple data types such as file name, file path, file size, modification time, and full-scale metadata information, that is, it supports multi-dimensional efficient retrieval.
[0059] Step 104: The user space retrieves in the updated hybrid index database based on the target retrieval strategy and the file retrieval instruction, and obtains the file retrieval result.
[0060] Specifically, the retrieval engine module in the user space can receive the file retrieval instruction input by the user and determine the target retrieval strategy corresponding to the file retrieval instruction, and complete the data retrieval through the updated hybrid index database in the index management module, so as to obtain the file retrieval result. That is, the file retrieval instruction is parsed into multiple retrieval tasks that support parallelization. Exemplarily, the retrieval engine module may include a unified interface access unit, a structured condition processing unit, and a parallel retrieval execution unit. The unified interface access unit provides various types of standardized interfaces. The structured condition processing unit parses the file retrieval instruction into multiple retrieval tasks that support parallelization. The parallel retrieval execution unit schedules multiple backend data sources in parallel to cooperate in retrieving file metadata and merges multi-source data. That is, through task decomposition, asynchronous pipelining, and lock-free communication, it collaborates efficiently in parallel among multiple index data sources in the updated hybrid index database in multiple retrieval plans to obtain the file retrieval result.
[0061] The unified interface access unit supports the access of multiple clients such as the terminal command mode, the desktop, and the Web, and supports interfaces such as the local Software Development Kit (SDK), Remote Procedure Call (RPC), and Representational State Transfer Application Program Interface (RESTful API), so as to provide a unified interface access layer, with good compatibility, good usability, and improved retrieval efficiency through parallel retrieval optimization methods.
[0062] It should be noted that the user space may further include a visual interaction module, which is displayed through a visual human-computer interaction interface, not only providing good human-computer interaction and intuitively displaying the file retrieval result, but also supporting multi-mode hybrid retrieval.
[0063] The visual interaction module has a visual interaction function, and the visual interaction function is completed through communication and cooperation with the retrieval engine module, providing a good human-computer interaction environment.
[0064] For the visual interaction module, it has a multi-mode retrieval function and an intuitive result display function.
[0065] The multi - mode retrieval function supports natural language retrieval, find - command - format retrieval, structured - condition - language retrieval, and hybrid - mode retrieval.
[0066] The intuitive result display function supports presenting file metadata in a structured manner and is deeply integrated with the system file browser.
[0067] Exemplarily, in the visual human - computer interaction interface, there are three areas: a retrieval input area, a function button area, and a retrieval result display area. The retrieval input area dynamically recommends high - frequency retrieval terms based on historical retrieval logs and file - system metadata, and can also switch between different retrieval modes; the function button area provides retrieval buttons and mode - switching buttons; in the retrieval result display area, for file retrieval results, it supports keyword - matching highlighting, also supports expanding file metadata information, and realizes "zero - click" path jumping through deep system integration (double - clicking on the path column automatically invokes the operating - system default file manager to locate the file entity).
[0068] In the file retrieval method provided by the embodiments of the present invention, when the user - space responds to a file retrieval instruction received at the current moment, it first determines the target retrieval strategy corresponding to the file retrieval instruction and updates the event - filtering strategy to the kernel - space. The kernel - space obtains the target file operation event from all file operation events captured at the current moment based on the updated event - filtering strategy; the user - space further determines the metadata information of the target file corresponding to the target file operation event, updates the hybrid index database based on the metadata information, and then performs a retrieval in the updated hybrid index database based on the target retrieval strategy and the file retrieval instruction to obtain the file retrieval result. Since the updated hybrid index database includes the mapping relationship between the target file and metadata information, and the mapping relationship between historical files and historical metadata information; and the hybrid index database supports multiple types of retrieval conditions, in this way, through the zero - copy data capture mechanism driven by kernel - space file operation events and the collaborative architecture of user - space dynamic hybrid indexing, an efficient real - time file retrieval tool is provided, achieving the real - time update of file - system metadata information and the efficient retrieval purpose of multiple retrieval dimensions. Thus, it can not only meet the efficient file - retrieval requirements but also improve the file - retrieval efficiency.
[0069] Based on the above Figure 1 In one exemplary embodiment of the method shown above, in step 101, updating the event - filtering strategy to the kernel - space, the specific process in this embodiment can be achieved through Figure 2 the steps 201 to 203 shown below to implement the following steps.
[0070] Step 201: Obtain the preset - format - rule configuration file at the current moment.
[0071] Step 202: Parse the event filtering rules from the preset format rule configuration file and perform a legality check on the event filtering rules.
[0072] Step 203: Use the event filtering rules that pass the legality check as the updated event filtering policy and update it in real time to the kernel space.
[0073] It should be noted that in addition to the index management module and the retrieval engine module, the user space may also include a dynamic rule management module. The dynamic rule management module supports the user space to modify the filtering configuration file, and through the BPF hash table opened in the kernel space by the event capture module for data interaction, it dynamically updates the event filtering policy to the event capture module in the kernel space in real time, so that the filtered event filtering policy takes effect immediately.
[0074] Specifically, the dynamic rule management module loads the event filtering policy into the BPF hash table when the event capture module is initialized and run, and can also dynamically update the event filtering policy during the running process of the event capture module.
[0075] In this way, it is possible to first obtain the filtering configuration file modified by the user in the user space, and then read the preset format rule configuration file at the current moment from the filtering configuration file, such as reading a rule configuration file based on the YAML format.
[0076] Exemplarily, the content example of the rule configuration file in this embodiment is as follows: rule_set: - {prefix: [" / home", " / var / log"], suffix: [".log", ".db"], type: [1,2]}.
[0077] For the preset format rule configuration file at the current moment, the event filtering rules can be first parsed from the preset format rule configuration file, and then the legality of the event filtering rules can be further checked. The event filtering rules that pass the legality check are used as the updated event filtering policy and updated in real time to the BPF hash table in the kernel space in a hot update manner.
[0078] Based on the above Figure 1 shown method, in an exemplary embodiment, in step 102, the target file operation event is obtained from all the file operation events captured at the current moment based on the updated event filtering policy. The specific determination process in this embodiment can be implemented through the following steps.
[0079] Based on different levels of event filtering policies included in the updated event filtering policy, perform different levels of event filtering on all file operation events to screen out target file operation events from all file operation events.
[0080] Among them, the different levels of event filtering policies include a first-level event filtering policy based on filtering by a preset path prefix, and a second-level event filtering policy based on filtering by a preset file type and a preset file name suffix.
[0081] Specifically, the kernel space implements a two-level event filtering mechanism in the event capture module. Irrelevant events are filtered out based on a preset path prefix, a preset file suffix, and a preset file type. Using a BPF hash table (BPF_MAP_TYPE_HASH), the dynamic rule management module dynamically updates the event filtering policy to the event capture module. Further, a BPF ring buffer (BPF_MAP_TYPE_RINGBUF) is used to achieve efficient data transfer between the event capture module and the index management module, realizing zero-copy data synchronization. At the same time, the BPF ring buffer ensures the orderliness of events by sending target file operation events to the shared buffer. Its advantages are particularly suitable for scenarios that require high-frequency event collection and low-latency processing.
[0082] Exemplarily, the preset path prefix can be, for example, / home, the preset file name suffix can be, for example,.log, and the preset file type can be, for example, regular files, directories, etc.
[0083] It should be noted that all file operation events received by the event filtering unit in the event capture module include the absolute path and file type of the file being operated on.
[0084] In this way, for all captured file operation events, first perform the first-level filtering based on the file path, that is, filtering based on the preset path prefix, and perform the filtering through string prefix matching. Then, further perform the second-level filtering on the file operation events that match successfully, that is, filtering based on the preset file name suffix and the preset file type. Both levels of filtering here are dynamic filtering processes, and irrelevant operation events are deleted through the two-level filtering mechanism.
[0085] Furthermore, file operation events that are successfully matched in both levels of filtering are directly discarded in the kernel space. File operation events that are not successfully matched in at least one level of the two-level filtering are all used as target file operation events that need to be transferred to the user space, and are transmitted to the index management module in the user space through the BPF ring buffer.
[0086] Based on the above Figure 1The method shown, in an exemplary embodiment, in step 103, the hybrid index database is updated based on the metadata information of the target file. The specific determination process in this embodiment, when the hybrid index database includes an inverted index library, a prefix tree index library, and a metadata index library, can be achieved through Figure 3 the steps 301 to 303 shown.
[0087] Step 301: Update the keyword - file ID mapping table existing in the inverted index library based on the file ID and file name keywords of the target file.
[0088] Step 302: Update the file path - file ID mapping table existing in the prefix tree index library based on the file ID of the target file and the file path in the metadata information.
[0089] Step 303: Update the metadata information - file ID mapping table existing in the metadata index library based on the metadata information and the file ID.
[0090] Specifically, the index management module in the user - space constructs a hybrid index database of an inverted index library, a prefix tree index library, and a metadata index library based on the hybrid index structure. And the index management module can specifically include an index initialization unit, an event receiving and concurrency control unit, an inverted index unit based on file names, a prefix tree index unit based on file paths, and a file metadata database storage unit.
[0091] The construction process of the hybrid index database will be exemplarily described below in combination with the above units: The index initialization unit is used to, when the operating system runs for the first time and it is detected that there is no initialized database file, perform a full - volume index construction for the existing file system of the current operating system, that is, construct the full - volume index database of the file system for the first time. The construction process can refer to the existing construction methods and will not be specifically limited and described here.
[0092] The event receiving and concurrency control unit is used to receive all target file operation events sent from the kernel - space from the BPF circular buffer and ensure that at least one target file operation event corresponding to the same initial file is strictly processed in the order of occurrence; the specific processing process can adopt a sharded queue structure to ensure that a class of target file operation events corresponding to the same initial file is strictly processed in order and different initial files are processed in parallel with the maximum parallelism.
[0093] For a sharded queue, each shard corresponds to an independent waiting queue, and file operation events with the same file path are always routed to the same shard. And for concurrent control, there is parallelism at the shard level, serialization at the file level, and merging at the operation level. Parallelism at the shard level means that different shard queues can be processed in parallel. Serialization at the file level means that file operation events for the same file are executed strictly in order. Merging at the operation level means that multiple identical operations within the window period are merged to reduce the number of index updates.
[0094] The inverted index unit based on file names is used to tokenize file names to obtain keywords, establish a hash table, map keywords to sets of file IDs, implement a two-way mapping between file name keywords and files, and obtain an inverted index database containing a keyword - file ID mapping relationship table.
[0095] The prefix tree index unit based on file paths is used to build a prefix tree index library based on file paths and file IDs, support complex path wildcard retrieval, and also support fast retrieval from file paths to file entity attributes. And in the prefix tree index library, each node is a file, and each node is mapped to the file ID in the file metadata database, thereby obtaining a prefix tree index library containing a file path - file ID mapping relationship table.
[0096] The file metadata database storage unit is used to store all the full - volume file metadata information of each file, support conditional retrieval based on the metadata information. That is, a complete information library is established based on the file metadata information of all files. Each file has a unique file ID in this information library. This information library builds an index structure based on key metadata fields and supports fast conditional filtering and retrieval of metadata information, such as file size, modification time, etc. Thus, a full - volume file metadata database is constructed, and the file metadata includes a full - volume metadata information - file ID mapping relationship table.
[0097] It should be noted that for the index management module, in addition to constructing the index layer of the hybrid index database, it can also include a cache layer to optimize the response performance in high - frequency retrieval scenarios. The cache layer is used to cache file IDs whose retrieval frequency is greater than a preset frequency threshold within a preset duration, and sort the retrieved priorities of each cached ID based on the cumulative retrieval frequency, thereby obtaining a file ID - cumulative retrieval frequency - retrieval priority mapping relationship table. And when the storage capacity threshold is not exceeded, the corresponding file data is cached. When the cache capacity exceeds the storage capacity threshold, the file ID with the lowest cumulative retrieval frequency within the preset time, or several file IDs with relatively low cumulative retrieval frequencies and their corresponding file data can be deleted. In this way, each subunit in the index management module can, through the way of hierarchical cooperation and asynchronous update strategy, significantly reduce resource consumption while ensuring real - time performance.
[0098] Further, a full-scale file metadata database is constructed based on file metadata information at the index layer, and a unique file ID is assigned to each file in this file metadata database. In this way, after obtaining a file retrieval result using other indexing methods, the corresponding file metadata information can also be found according to this file metadata database.
[0099] In this embodiment, the content example of the file metadata database can be described as Example 1 below: There is a file named ssh-keygen. After word segmentation of this file name, two keywords ssh and keygen are obtained. Then there are 2 index keys ssh and keygen in the inverted index database. At the same time, these 2 index keys may also be associated with other files. The index structure is shown as follows: There are 3 files, and the example information stored in the file metadata database is as follows in Example 2: [{"ID":1,"PATH":" / usr / bin / ssh-keygen",……},{"ID":2,"PATH":" / usr / sbin / ntp-keygen",……}, {"ID":3,"PATH":" / usr / bin / ssh-copy-id",……},……]。
[0100] For the inverted index unit based on file names, after word segmentation of each file name, multiple word segments are used to establish an inverted index with the file ID of the corresponding file to support fast searching from file names to files.
[0101] In this embodiment, the content example of the inverted index unit is described as Example 3 below: Combined with the above example, the following inverted index structure is constructed for the files / usr / bin / ssh-keygen, / usr / bin / ssh-copy-id, and / usr / sbin / ntp-keygen: "ssh": [1, 3, ……], "keygen": [1,2, ……] In addition, in the inverted index database, a mapping can also be established by using the keywords obtained from word segmentation of the file name as the key (key) and the unique key in the file metadata database as the value (value).
[0102] Furthermore, for the prefix tree index unit based on file paths, the constructed prefix tree index library contains the prefix tree structure of the complete path, and each node stores the file ID corresponding to the file of the path from the root node to this node; moreover, the complete file path is stored in the prefix tree structure, supporting efficient prefix matching retrieval (such as / var / log / *); for example, in Example 1, after retrieving the associated files according to the keyword ssh, and then according to the corresponding file path, retrieve the file metadata information in the prefix tree structure; at the same time, it also supports directly retrieving files based on the prefix tree index library.
[0103] In this embodiment, the content example of the prefix tree index library is described as Example 4 below: Combined with the above examples, the following prefix tree structure is constructed for the files / usr / bin / ssh-keygen, / usr / bin / ssh-copy-id, / usr / sbin / ntp-keygen: / (root directory)("id": xx) └── "bin" ("id": xx) └── "ssh-keygen"("id": 1) └── "ssh-copy-id" ("id": 3) └── "sbin" ("id": xx) └── "ntp-keygen" ("id": 2) Optionally, in the embodiment of the present invention, the cache layer in the index management module is the performance acceleration layer of the index management module. By caching the file index data of high-frequency access or high-frequency retrieval, the direct access frequency to the underlying storage is reduced, thereby optimizing the retrieval response speed.
[0104] At this time, according to the file ID and file name keyword of each target file, the existing keyword - file ID mapping relationship table in the inverted index library can be updated; according to the file wildcard and the file ID of each target file, and the file path in the metadata information of each target file, the existing file path - file ID mapping relationship table in the prefix tree index library can be updated; and according to the metadata information and file ID of each target file, the existing metadata information - file ID mapping relationship table in the metadata index library can be updated; thus obtaining the updated hybrid index database.
[0105] In an embodiment of the present invention, a hybrid index database is designed in the user space, integrating an inverted index, a prefix tree, and a metadata database. The hybrid index database supports keyword matching, wildcard retrieval, and multi-condition combination retrieval, and combines a caching mechanism and parallel retrieval optimization to achieve the purpose of quickly responding to complex retrieval conditions.
[0106] Based on the above Figure 1 method shown, in an exemplary embodiment, in step 104, based on the target retrieval strategy and the file retrieval instruction, a retrieval is performed in the updated hybrid index database to obtain a file retrieval result. The specific retrieval process in this embodiment can be Figure 4 implemented by the steps 401 to 402 shown.
[0107] Step 401: Based on the retrieval conditions corresponding to the file retrieval instruction, determine at least two target index databases in the updated inverted index database, the updated prefix tree index database, and the updated metadata index database.
[0108] Step 402: Based on the target retrieval strategy, perform a collaborative retrieval in at least two target index databases to obtain a file retrieval result.
[0109] Specifically, for the retrieval engine module in the user space, it first determines whether the received file retrieval instruction is a structured condition retrieval instruction. When it is determined that the file retrieval instruction is a structured condition retrieval instruction, the structured condition retrieval instruction can be further decomposed into retrieval conditions, that is, task decomposition, combined with asynchronous pipelining and lock-free communication, to achieve parallel retrieval optimization and improve retrieval efficiency.
[0110] It should be noted that the retrieval engine module provides a unified and structured retrieval method and is responsible for efficiently processing retrieval tasks; it can specifically include a unified interface access unit, a structured condition processing unit, and a parallel retrieval execution unit.
[0111] The unified interface access unit is used to receive file retrieval instructions through various types of standardized interfaces.
[0112] The structured condition processing unit is used to process the file retrieval instruction input by the user into multiple retrieval tasks that support parallel retrieval. Each retrieval task is a task of retrieving with the corresponding retrieval condition. Specifically, it can analyze and process the structured condition retrieval instruction input by the user to generate a retrieval plan, optimize the retrieval strategy, select at least two target index databases to execute the retrieval task; thereby generating a target retrieval strategy.
[0113] A parallel retrieval execution unit for parallelly scheduling multiple backend data sources to collaboratively retrieve file metadata and merge multi-source data; that is, through task decomposition, asynchronous pipelining, and lock-free communication, multiple index data sources are coordinated in multiple retrieval plans to achieve efficient parallel collaboration.
[0114] In this embodiment, all the condition examples supported by the structured condition processing unit are as follows: 1) Metadata metrics, such as file name, file name of the soft link file associated with the file, path, file type, file size, file owner, file group, file permissions, last access time, last modification time, last status change time, file creation time; 2) Comparison conditions, not equal to, greater than, less than, greater than or equal to, less than or equal to; 3) Logical conditions, satisfying multiple conditions simultaneously, satisfying any one condition, not satisfying a certain condition; 4) Range conditions, within a certain range, not within a certain range; 5) Fuzzy matching conditions, using the wildcard * to match strings; 6) Regular expression conditions, using regular expressions to process conditional retrieval of file names and file paths.
[0115] Furthermore, the parallel retrieval execution unit executes the target retrieval strategy, parallelly schedules the index data sources of at least two target index libraries to collaboratively retrieve file metadata until a file retrieval result is obtained.
[0116] Exemplarily, an example of the retrieval process of the parallel retrieval execution unit is as follows: By parsing the retrieval conditions for the file retrieval instruction "find Word files modified in the last week", it can be known that files with extensions such as doc, docx, and odt need to be retrieved, and the last modification time is within one week. According to the parsed retrieval conditions, the updated inverted index library and the updated metadata index library are used as two target index libraries for collaborative retrieval.
[0117] First, for the first retrieval plan, specifically in the target index library of the updated inverted index library, retrieve using doc, docx, and odt as keywords, and put all the retrieved file IDs into a lock-free circular queue.
[0118] Then, for the second retrieval plan, take out all the file IDs from the lock-free circular queue, retrieve the corresponding file data in the target index library of the updated metadata index library, and perform conditional filtering based on the last modification time, file type, and file name suffix.
[0119] Finally, sort the file data after condition filtering in descending order according to the last modification time, and use the file data after descending order sorting as the file retrieval result to feedback to the visual human-computer interaction interface.
[0120] Here, it should be noted that in order to improve the retrieval efficiency and reduce the retrieval resources, the file data set matching the file retrieval instruction can be retrieved first in the updated cache space; the updated cache space is obtained by updating the existing file ID - cumulative retrieval frequency - retrieval priority mapping relationship table in the cache space based on the file ID of each target file.
[0121] If there is a file data set in the updated cache space, there is no need to perform collaborative retrieval in the above two target index libraries, and this file data set is determined as the file retrieval result; on the contrary, if there is no file data set in the updated cache space, the file retrieval result is obtained by using the collaborative retrieval method of the above two target index libraries.
[0122] Based on the above Figure 1 In the method shown, in an exemplary embodiment, the user space can also process the unstructured condition retrieval instruction input by the user into a structured condition retrieval instruction, and its specific processing process in this embodiment can be implemented through the following steps.
[0123] When the file retrieval instruction is a natural language retrieval instruction, based on the pre-trained target language processing model, the natural language retrieval instruction is processed into a structured language retrieval instruction.
[0124] Among them, the target processing model is obtained by training the NLP model based on the Transformer framework with sample natural language retrieval instructions in different file retrieval scenarios and sample structured retrieval instructions marked by each sample natural language retrieval instruction; moreover, data augmentation has been performed on each sample natural language retrieval instruction.
[0125] It should be noted that the present invention realizes file retrieval based on natural language by introducing an NLP model based on the Transformer architecture, converting natural language into structured conditions, and integrating with the traditional structured condition retrieval method to support multi-mode hybrid retrieval.
[0126] Specifically, the user space can also include a natural language processing module, introduce an NLP model based on the Transformer architecture, and train this NLP model into a target processing model capable of automatically processing the input natural language retrieval instruction into a structured language retrieval instruction. In this way, through the method of automatically identifying semantic elements, the end-to-end conversion from natural language and find command syntax to structured retrieval conditions is realized.
[0127] For an NLP model based on the Transformer architecture, its training process is carried out by using a natural language corpus and a find command corpus, achieving an end-to-end conversion from natural language to structured retrieval conditions.
[0128] The natural language corpus is constructed by collecting sample natural language retrieval instructions in different file retrieval scenarios and separately annotating the corresponding sample structured retrieval instructions, covering scenarios such as operation and maintenance, development, and user document management, and expanding to include logical retrievals. The generalization ability of the model is improved through data augmentation (i.e., injecting noise). Further, for each natural language retrieval instruction, the corresponding find retrieval instruction is marked with the corresponding structured retrieval condition as a find command for training; in this way, a training sample set is obtained. Each training sample in the training sample set is a sample natural language retrieval instruction in the corresponding file retrieval scenario, and this sample natural language retrieval instruction has undergone data augmentation and is marked with a sample structured retrieval instruction.
[0129] The natural language processing module can automatically identify semantic elements and includes a preprocessing unit, an encoding unit, a decoding unit, a vocabulary mapping unit, and a proxy retrieval unit.
[0130] The preprocessing unit is used to convert the target text instruction into a sequence of tokens (token) based on a tokenizer, and then convert the token sequence into a feature vector.
[0131] The encoding unit is used to embed the feature vector output by the preprocessing unit into a high-dimensional vector space and add position encoding to retain the sequence order information, thereby obtaining a hidden vector that can abstractly represent the token sequence. This hidden vector can be regarded as a compressed, summarized, and abstract expression of the target text instruction, facilitating subsequent information processing by the model to extract more features.
[0132] Subsequently, all high-level semantic features in the hidden vector are re-extracted through the multi-head self-attention mechanism, that is, the relationships between all high-level semantic features in the hidden vector are analyzed, and a more detailed fine-grained feature representation vector is generated based on the analyzed relationships. This abstract fine-grained feature representation vector will then be used by the decoding unit to generate the target language instruction.
[0133] The decoding unit is used to convert the fine feature representation vector generated by the encoding unit into an output token sequence, mapping the features in the intermediate hidden state space to the desired target space; overall, the decoding is an autoregressive architecture that gradually constructs a structured retrieval condition through an autoregressive mechanism; the decoder predicts and outputs the result for only one token in the output token sequence each time, and after outputting the prediction result of a certain token, it uses this prediction result as the input for the next prediction and sends it together with the fine feature representation vector output by the encoder into the decoder. Since the vector output by the decoder already contains the context connection between each token and other tokens, when using the decoder for prediction, only the currently to-be-predicted token and the previous prediction results can be used as the input.
[0134] The vocabulary mapping unit converts the output token sequence generated by decoding into a structured language retrieval instruction, including two stages: field mapping and logical mapping; the field mapping translates and converts each retrieval field and its own constraint conditions, and the logical mapping is used to translate and convert the logical relationships between multiple retrieval fields. After two-stage mapping, a structured language retrieval instruction combining conditions and logic is formed, and further, the proxy retrieval unit is used for request forwarding.
[0135] Exemplarily, for the structured language retrieval instruction "find Word files modified in the last week", parsing its keywords includes the last week, modified, and Word files, and it is necessary to determine the retrieval time range, and the retrieval time parameter is the modification time of the file meta-attribute; for Word, files with extended names such as doc, docx, odt, etc. are mapped.
[0136] In this embodiment, the structured language retrieval conditions converted from the above natural language retrieval instruction examples are described as follows: ``` { "path": [" / "], "constraints": { "id": 0, "field": "mtime", "operator": "ge", "value": timestamp(s) }, { "id": 1, "field": "suffix", "operator": "in", "value": ["doc", "docx", "odt"] }, { "id": 2, "field": "type", "operator": "eq", "value": "file" } , "logic": { "type": "and", "constraints": [0, 1, 2] } } ``` Among them, path represents the directory range restricted by the search. By default, the search starts from the root directory; constraints represent the constraint conditions related to the search, supporting all conditions supported by all structured condition processing units; id represents the unique number assigned to each constraint condition; field represents the specified constraint condition type, such as mtime to constrain the file modification time; operator represents some search operation restrictions for the specified constraint condition, such as ge representing the operation behavior of greater than or equal to; value represents the target value that the specified constraint condition needs to meet, such as the search time specified by timestamp with a precision of seconds; logic represents the logical relationship between multiple constraint conditions during the search; type represents the logical type, and the value can be and or or; constraints represent the relevant constraint conditions participating in this logical operation.
[0137] The proxy search unit is used to further encapsulate the structured language search instructions output by the thesaurus mapping unit according to the interface standard of the search engine module, forward the corresponding search requests, and finally return the file search results.
[0138] It should be noted that the descriptions of the above four units, namely the preprocessing unit, the encoding unit, the decoding unit, and the vocabulary mapping unit, can be used in the model training stage and also in the subsequent actual application stage; for the model training stage, the target text instruction can specifically be the sample natural language retrieval instruction in different file retrieval scenarios, and each sample natural language retrieval instruction has undergone data augmentation and is marked with the corresponding sample structured retrieval instruction; for the subsequent actual application stage, the target text instruction can specifically be the natural language retrieval instruction input by the user at the current moment. On the trained target language processing model, after converting the natural language retrieval instruction into a structured retrieval instruction, through proxy retrieval, a retrieval request is further sent to the retrieval engine module. That is to say, the principles of the model training stage and the actual application stage are the same. Since the present invention does not improve the structure of the NLP model based on the Transformer framework itself, and the improvement lies in the collected training sample set and training it to have the function of converting natural language into structured language, therefore, only an overview of the model training process and the subsequent actual application process is given here in combination with the training sample set.
[0139] In the embodiment of the present invention, by introducing an NLP model based on the Transformer architecture, natural language is converted into structured language, file retrieval based on natural language is realized, and it is integrated with the traditional structured condition retrieval method to support multi-mode hybrid retrieval, improving the flexibility and accuracy of file retrieval.
[0140] Based on the above Figure 1 In one exemplary embodiment of the method shown, in step 101, determining the target retrieval strategy corresponding to the file retrieval instruction, the specific determination process in this embodiment can be achieved through Figure 5 the steps 501 to 502 shown.
[0141] Step 501: Parse the retrieval conditions of the file retrieval instruction, and determine the target retrieval plan based on all the parsed retrieval conditions.
[0142] Step 502: Based on the pre-set mapping relationship between the retrieval plan and the retrieval strategy, determine the target retrieval strategy corresponding to the target retrieval plan.
[0143] Specifically, extract keywords from the file retrieval instruction, and determine the retrieval conditions based on the metadata fields mapped by all the extracted keywords. That is to say, when all the extracted keywords include metadata fields, all the keywords can be used as retrieval conditions.
[0144] For example, all keywords extracted for the file retrieval instruction "Find Word files modified in the last week" include Word files, the last week, and modified. Word files are mapped to files with extension formats such as doc, docx, odt, etc. The last week and modified are mapped to the last modification time within one week. Thus, two retrieval conditions are obtained.
[0145] At this time, the target retrieval plans corresponding to all retrieval conditions can be determined. The number of target retrieval plans can be the same as and in one-to-one correspondence with the number of retrieval conditions. Each target retrieval plan includes the corresponding retrieval condition and the processing method of the retrieval result.
[0146] Furthermore, based on the pre-set mapping relationship between the retrieval plan and the retrieval strategy, the target retrieval strategy corresponding to the target retrieval plan can be determined. The target retrieval strategy is specifically the specific implementation process for realizing the corresponding target retrieval plan.
[0147] Exemplarily, when the target retrieval plan includes a first retrieval plan and a second retrieval plan, the target retrieval strategy corresponding to the first retrieval plan is to perform a retrieval based on the corresponding retrieval condition and put the retrieved candidate file IDs into a lock-free circular queue; the target retrieval strategy corresponding to the second retrieval plan is to take out all candidate file IDs from the lock-free circular queue, retrieve the file data corresponding to each candidate file ID, and further perform conditional filtering on all retrieved file data based on the corresponding retrieval condition. After that, other processing such as descending or ascending sorting can be performed on the conditional filtering result.
[0148] The file retrieval method provided by the embodiments of the present invention works through the cooperation of the zero-copy data acquisition mechanism driven by kernel-level events, the user-level dynamic hybrid index architecture, and the natural language semantic parsing technology. That is to say, through the cooperative architecture of kernel-level event-driven and user-level hybrid index, directly listen to file operation events in the kernel level and update the index database in real time; propose a dynamic sharding queue architecture in the user level to ensure the orderliness and performance of the hybrid index in concurrent event updates and support millisecond-level response; parse natural language into a structured retrieval language through the NLP engine to achieve natural language retrieval, enabling end users to complete complex file retrievals in a daily conversation manner with zero technical threshold, thereby realizing the real-time update, efficient retrieval, and natural language interaction capabilities of the file system metadata.
[0149] It should be noted that although the operations of the method of the present invention are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the order of the steps depicted in the flowchart can be changed. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0150] In one embodiment, the embodiment of the present invention further provides a file retrieval device, as Figure 6 shown. The file retrieval device includes: a policy determination unit 601, an event acquisition unit 602, an index library update unit 603, and a file retrieval unit 604.
[0151] The policy determination unit 601 is configured to determine, in the user space, a target retrieval policy corresponding to a file retrieval instruction in response to the file retrieval instruction received at the current moment, and update an event filtering policy to the kernel space.
[0152] The event acquisition unit 602 is configured to obtain, in the kernel space, a target file operation event from all file operation events captured at the current moment based on the updated event filtering policy.
[0153] The index library update unit 603 is configured to determine, in the user space, metadata information of a target file corresponding to the target file operation event, and update a hybrid index database based on the metadata information; the updated hybrid index database includes a mapping relationship between the target file and the metadata information, and a mapping relationship between a historical file and historical metadata information; and the hybrid index database supports multiple types of retrieval conditions.
[0154] The file retrieval unit 604 is configured to perform a retrieval in the updated hybrid index database based on the target retrieval policy and the file retrieval instruction to obtain a file retrieval result.
[0155] In one embodiment, the policy determination unit 601 is specifically configured to parse retrieval conditions of the file retrieval instruction, and determine a target retrieval plan based on all the parsed retrieval conditions; and determine a target retrieval policy corresponding to the target retrieval plan based on a mapping relationship between the preset retrieval plan and the retrieval policy.
[0156] In one embodiment, the policy determination unit 601 is specifically configured to obtain a preset format rule configuration file at the current moment; parse an event filtering rule from the preset format rule configuration file, and perform a legality check on the event filtering rule; and use the event filtering rule that passes the legality check as the updated event filtering policy and update it to the kernel space in real time.
[0157] In one embodiment, the event acquisition unit 602 is specifically configured to perform different-level event filtering on all file operation events based on different-level event filtering policies included in the updated event filtering policy, so as to screen out target file operation events from all file operation events; wherein, the different-level event filtering policies include a first-level event filtering policy based on a preset path prefix filtering, and a second-level event filtering policy based on a preset file type and a preset file suffix for filtering.
[0158] In one embodiment, when the hybrid index database includes an inverted index library, a prefix tree index library, and a metadata index library, the index library update unit 603 is specifically configured to update the existing keyword-file ID mapping table in the inverted index library based on the file ID and file name keywords of the target file; update the existing file path-file ID mapping table in the prefix tree index library based on the file wildcards and file IDs of the target file, and the file path in the metadata information; update the existing metadata information-file ID mapping table in the metadata index library based on the metadata information and the file ID.
[0159] In one embodiment, the file retrieval unit 604 is specifically configured to determine at least two target index libraries in the updated inverted index library, the updated prefix tree index library, and the updated metadata index library based on the retrieval conditions corresponding to the file retrieval instruction; perform collaborative retrieval in the at least two target index libraries based on the target retrieval strategy to obtain a file retrieval result.
[0160] In one embodiment, the policy determination unit 601 is further specifically configured to, when the file retrieval instruction is a natural language retrieval instruction, process the natural language retrieval instruction into a structured language retrieval instruction based on a pre-trained target language processing model.
[0161] The target processing model is obtained by training an NLP model based on the Transformer framework with sample natural language retrieval instructions in different file retrieval scenarios and sample structured retrieval instructions identified by each sample natural language retrieval instruction; and each sample natural language retrieval instruction has undergone data augmentation.
[0162] It should be understood that the units described in the file retrieval device and the reference Figure 1corresponds to each step in the described method. Thus, the operations and features described above for the method also apply to the file retrieval device and the units included therein, and will not be elaborated here. The file retrieval device can be pre-implemented in a browser or other secure applications of a computer device, or can be loaded into the browser or its secure applications of the computer device by means of downloading, etc. The corresponding units in the file retrieval device can cooperate with the units in the computer device to implement the solution of the embodiments of the present invention.
[0163] Reference is made below to Figure 7 , which shows a schematic structural diagram of a computer system 700 of a terminal device or a server suitable for implementing the embodiments of the present invention.
[0164] As Figure 7 shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0165] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as required. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as required so that a computer program read therefrom can be installed into the storage section 708 as required.
[0166] Specifically, according to an embodiment of the present disclosure, the process described above with reference to Figure 1 can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program tangibly contained on a machine-readable medium, and the computer program includes program code for performing Figure 1 the method. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711.
[0167] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0169] The units or modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described units or modules can also be provided in a processor. Among them, the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0170] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the computer device described in the above embodiments, or may exist separately without being assembled into the computer device. The above computer-readable storage medium stores one or more programs, and when the above programs are executed by one or more processors, the methods described in the present invention are performed. For example, the steps of Figure 1 the method shown can be executed.
[0171] The embodiments of the present invention provide a computer program product, which includes instructions that, when run, cause the methods described in the embodiments of the present invention to be executed. For example, the steps of Figure 1 the method shown can be executed.
[0172] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided by the present invention can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided by the present invention can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0173] The above description is only the preferred embodiments of the present invention and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present invention.
Claims
1. A file retrieval method, characterized in that, Applied to a computer device, the computer device is installed with an operating system, and the operating system includes a user space and a kernel space; the method includes: The user space determines a target retrieval strategy corresponding to the file retrieval instruction in response to the file retrieval instruction received at the current moment, and updates the event filtering strategy to the kernel space; The kernel space obtains a target file operation event from all file operation events captured at the current moment based on the updated event filtering strategy; The user space determines the metadata information of the target file corresponding to the target file operation event, and updates the hybrid index database based on the metadata information; the updated hybrid index database includes the mapping relationship between the target file and the metadata information, and the mapping relationship between the historical file and the historical metadata information; and the hybrid index database supports multiple types of retrieval conditions; The user space retrieves in the updated hybrid index database based on the target retrieval strategy and the file retrieval instruction to obtain a file retrieval result.
2. The method according to claim 1, characterized in that The obtaining of the target file operation event from all file operation events captured at the current moment based on the updated event filtering strategy includes: Performing different-level event filtering on all file operation events based on different-level event filtering strategies included in the updated event filtering strategy to screen out the target file operation event from all file operation events; Wherein, the different-level event filtering strategies include a first-level event filtering strategy based on filtering by a preset path prefix, and a second-level event filtering strategy based on filtering by a preset file type and a preset file suffix.
3. The method according to claim 1, wherein The updating of the hybrid index database based on the metadata information includes: When the hybrid index database includes an inverted index library, a prefix tree index library, and a metadata index library, Updating the keyword-file ID mapping relationship table existing in the inverted index library based on the file ID and file name keyword of the target file; Updating the file path-file ID mapping relationship table existing in the prefix tree index library based on the file ID of the target file and the file path in the metadata information; Updating the metadata information-file ID mapping relationship table existing in the metadata index library based on the metadata information and the file ID.
4. The method according to claim 3, characterized in that, The retrieving in the updated hybrid index database based on the target retrieval strategy and the file retrieval instruction to obtain a file retrieval result includes: Determining at least two target index libraries in the updated inverted index library, the updated prefix tree index library, and the updated metadata index library based on the retrieval condition corresponding to the file retrieval instruction; Performing collaborative retrieval in the at least two target index libraries based on the target retrieval strategy to obtain the file retrieval result.
5. The method according to any one of claims 1 to 4, characterized in that, The updating of the event filtering strategy to the kernel space includes: Obtaining a preset format rule configuration file at the current moment; Parse out the event filtering rules from the preset format rule configuration file and perform a legality check on the event filtering rules; Use the event filtering rules that pass the legality check as the updated event filtering policy and update it to the kernel space in real time.
6. The method according to any one of claims 1 to 4, characterized in that, The method further includes: When the file retrieval instruction is a natural language retrieval instruction, based on a pre-trained target language processing model, process the natural language retrieval instruction into a structured language retrieval instruction; Wherein, the target processing model is obtained by training an NLP model based on the Transformer framework with sample natural language retrieval instructions in different file retrieval scenarios and the sample structured retrieval instructions identified by each of the sample natural language retrieval instructions; and data augmentation has been performed on each of the sample natural language retrieval instructions.
7. The method according to any one of claims 1 to 4, characterized in that The determining the target retrieval strategy corresponding to the file retrieval instruction includes: Parse the retrieval conditions of the file retrieval instruction and determine a target retrieval plan based on all the parsed retrieval conditions; Based on the pre-set mapping relationship between the retrieval plan and the retrieval strategy, determine the target retrieval strategy corresponding to the target retrieval plan.
8. A file retrieval device, characterized in that, The device includes: A strategy determination unit configured to determine the target retrieval strategy corresponding to the file retrieval instruction in the user space in response to the file retrieval instruction received at the current moment, and update the event filtering strategy to the kernel space; An event acquisition unit configured to obtain a target file operation event from all file operation events captured at the current moment by the kernel space in response to the updated event filtering strategy; An index library update unit configured to determine the metadata information of the target file corresponding to the target file operation event in the user space and update the hybrid index database based on the metadata information; the updated hybrid index database includes the mapping relationship between the target file and the metadata information, and the mapping relationship between the historical file and the historical metadata information; and the hybrid index database supports multiple types of retrieval conditions; A file retrieval unit configured to perform a retrieval in the updated hybrid index database based on the target retrieval strategy and the file retrieval instruction to obtain a file retrieval result.
9. A computer device, comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the file retrieval method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the file retrieval method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Container mirror image data processing method and device, equipment and medium
CN114706658A
File searching method and device, storage medium and electronic equipment
CN115221119A
Retrieval method and device, medium and computing equipment
CN116226497A
File indexing method and device, electronic equipment and computer readable storage medium
CN117033307A
File resource reuse library management method and device based on database
CN119415481A
Cited By
Kernel probe generation method and device, computer equipment and storage medium
CN120950340A
File retrieval method and related product
CN121029698A
Industrial chain data retrieval treatment method based on multi-modal deep learning
CN122087099A
A Multimodal Deep Learning-Based Approach to Industry Chain Data Retrieval and Governance
CN122087099B