File retrieval method and device based on intention recognition rule and related medium

By constructing an intent recognition rule library and semantic modification dictionary, the existing file retrieval system's high computing resources consumption and inaccurate intent recognition are solved, and fast and accurate file retrieval is achieved, which is suitable for multilingual and complex query scenarios.

CN120295975APending Publication Date: 2025-07-11AFIRSTSOFT CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510417594.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing file relies on complex artificial intelligence models, resulting in high consumption of computing resources, slow response speed, and difficulty in accurately identifying users' search intentions, affecting the accuracy of search results.

Method used

Through predefined intent types and trigger word collections, an intent recognition rule library and semantic modification dictionary are constructed, semantic analysis and matching are performed, intent confidence is generated, and file retrieval is performed in combination with modification conditions.

Benefits of technology

It realizes rapid intent recognition, reduces computing resource consumption, improves response speed and accuracy of search results, supports multi-language environments and complex query expressions, and expands the coverage of file retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295975A_ABST
    Figure CN120295975A_ABST
Patent Text Reader

Abstract

The invention discloses a file retrieval method and device based on an intention recognition rule and a related medium, and the method comprises the steps: predefining a plurality of intention types, and constructing an intention recognition rule library; configuring a corresponding trigger word set according to the intention type, and establishing a semantic modification dictionary; performing semantic analysis on input query text analysis according to the semantic modification dictionary to obtain modification conditions; matching the intention type by using the trigger word set based on the intention recognition rule base to obtain a matching result, and generating an intention confidence coefficient according to the matching result; and retrieving a target file based on the modification condition and the intention confidence to obtain a retrieval result. According to the method, the intention confidence coefficient is generated by utilizing the matching result, then the target file is retrieved based on the modification condition and the intention confidence coefficient, and finally the retrieval result is obtained, so that the file retrieval response speed is high, the retrieval result is more accurate, and the intention recognition efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a file retrieval method, device and related medium based on intention recognition rules. Background Art

[0002] With the development of informatization, the storage scale of file data has been continuously expanding. At present, file retrieval systems generally adopt artificial intelligence-based recognition technologies. However, there are still several deficiencies in the existing technical solutions. First, the existing intention recognition methods highly rely on complex artificial intelligence models, which consume a large amount of computing resources during operation, resulting in slow system response speed and difficulty in meeting application scenarios with high real-time requirements. Second, when the existing file retrieval systems process complex query expressions, there are understanding deviations, making it difficult to accurately identify the actual retrieval intention of users, thus affecting the accuracy of retrieval results. In summary, the intention recognition efficiency of the existing file retrieval technology is low, and a new solution is urgently needed to improve the intention recognition efficiency. Summary of the Invention

[0003] Embodiments of the present invention provide a file retrieval method, device and related medium based on intention recognition rules, aiming to solve the problem of low intention recognition efficiency in existing file retrieval technologies.

[0004] In a first aspect, embodiments of the present invention provide a file retrieval method based on intention recognition rules, including:

[0005] Pre-define multiple intention types and construct an intention recognition rule library;

[0006] Configure a corresponding trigger word set according to the intention type, and at the same time establish a semantic modification dictionary;

[0007] Perform semantic parsing on the input query text according to the semantic modification dictionary to obtain modification conditions;

[0008] Match the intention type with the trigger word set based on the intention recognition rule library to obtain a matching result, and at the same time generate an intention confidence according to the matching result;

[0009] Retrieve the target file based on the modification conditions and intention confidence to obtain a retrieval result.

[0010] In a second aspect, embodiments of the present invention provide a file retrieval device based on intention recognition rules, including:

[0011] A rule preset unit for pre-defining multiple intention types and constructing an intention recognition rule library;

[0012] A set configuration unit, configured to configure a corresponding trigger word set according to the intent type, and simultaneously establish a semantic modification dictionary;

[0013] A text parsing unit, configured to perform semantic parsing on the input query text according to the semantic modification dictionary to obtain a modification condition;

[0014] A word matching unit, configured to match the intent type using the trigger word set based on the intent recognition rule library to obtain a matching result, and simultaneously generate an intent confidence level according to the matching result;

[0015] A target retrieval unit, configured to retrieve the target file based on the modification condition and the intent confidence level to obtain a retrieval result.

[0016] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the file retrieval method based on the intent recognition rule in the first aspect is implemented.

[0017] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the file retrieval method based on the intent recognition rule in the first aspect is implemented.

[0018] An embodiment of the present invention provides a file retrieval method based on an intent recognition rule, including predefined multiple intent types and constructing an intent recognition rule library; configuring a corresponding trigger word set according to the intent type, and simultaneously establishing a semantic modification dictionary; performing semantic parsing on the input query text according to the semantic modification dictionary to obtain a modification condition; matching the intent type using the trigger word set based on the intent recognition rule library to obtain a matching result, and simultaneously generating an intent confidence level according to the matching result; retrieving the target file based on the modification condition and the intent confidence level to obtain a retrieval result. The present invention generates an intent confidence level using the matching result, and then retrieves the target file based on the modification condition and the intent confidence level, and finally obtains the retrieval result. In this way, the file retrieval response speed is fast and the retrieval result is more accurate, improving the intent recognition efficiency.

[0019] An embodiment of the present invention further provides a file retrieval device, a computer device, and a storage medium based on an intent recognition rule, which also have the above beneficial effects. Description of the Drawings

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 It is a schematic flowchart of a file retrieval method based on intention recognition rules provided by an embodiment of the present invention;

[0022] Figure 2 It is a schematic block diagram of a file retrieval device based on intention recognition rules provided by an embodiment of the present invention. Detailed implementation manners

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0024] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "comprises" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0025] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0026] It should be further understood that the term " / and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0027] Please refer to the following Figure 1 , Figure 1 It is a schematic flowchart of a file retrieval method based on intention recognition rules provided by an embodiment of the present invention, specifically including: steps S101 to S105.

[0028] S101. Predetermine multiple intention types and construct an intention recognition rule library;

[0029] S102. Configure the corresponding trigger word set according to the intention type, and at the same time establish a semantic modification dictionary;

[0030] S103. Perform semantic parsing on the input query text according to the semantic modification dictionary to obtain modification conditions;

[0031] S104. Match the intention type with the trigger word set based on the intention recognition rule library to obtain a matching result, and at the same time generate an intention confidence level according to the matching result;

[0032] S105. Retrieve the target file based on the modification conditions and intention confidence level to obtain a retrieval result.

[0033] In step S101, by analyzing the user's query behavior and retrieval requirements, multiple intention types are defined, and a corresponding pattern matching rule is constructed for each intention. The pattern matching rules are stored in the form of rule entries to form an intention recognition rule library for fast matching and reasoning in the subsequent intention recognition process.

[0034] In one embodiment, step S101 includes:

[0035] Define the search intention type to establish a search action set for identifying the user's search request;

[0036] Define the file attribute intention type to establish a file type set for identifying the file type request;

[0037] Define the author intention type to establish an author identification set for identifying the author query request;

[0038] Define the theme intention type to establish a theme keyword set for identifying the theme query request;

[0039] Store the search intention type, file attribute intention type, author intention type, theme intention type and their corresponding sets into the intention recognition rule library.

[0040] In this embodiment, search intent types are defined to establish a set of search actions for identifying user search requests. The search intent types include common query verbs such as "find", "search", "look for", etc. Through predefined rules, these verbs and their related semantic extensions are constructed into a complete set of search actions. This set of search actions is used to identify the search actions expressed by the user in the query text and provide support for the subsequent intent recognition process. Secondly, file attribute intent types are defined to establish a set of file types for identifying file type requests. The set of file types includes information such as file types and file formats, such as text files, PDF files, Word files, Excel sheets, etc. The file attribute intent types are defined through rules, and corresponding identification words are configured for each file type to help the system accurately identify the user's intent during the query process. Then, author intent types are defined to establish a set of author identifiers for identifying author query requests. The set of author identifiers includes identifiers or related words of the file author, which can be the name or unit name of the file creator. These identifiers are compared with the query input by the user to identify the user's retrieval requirement for the file author. In addition, topic intent types can also be defined to identify a set of topic keywords for topic query requests. This set of topic keywords includes specific keywords or key phrases representing the topic or content of the file. By parsing the query text, the topic concerned by the user is identified, and relevant files are determined through keyword matching. Finally, the defined search intent types, file attribute intent types, author intent types, and topic intent types, as well as their corresponding sets, are stored in the intent recognition rule library. The intent recognition rule library, as a core component, stores different intent types and their matching rules, providing efficient query support for the subsequent intent recognition process.

[0041] In step S102, corresponding trigger word sets are configured for the intent types, and a semantic modification dictionary is established. The trigger word sets can include multi-language words, such as multiple groups of intent trigger words are configured for languages such as Chinese and English respectively. At the same time, a modification dictionary for semantic enhancement is established, including degree modifiers and negative words.

[0042] In one embodiment, step S102 includes:

[0043] Configure Chinese trigger words and English trigger words respectively according to the intent types;

[0044] Integrate the Chinese trigger words and English trigger words into the trigger word set;

[0045] Establish a degree word dictionary and a negative word dictionary respectively, and integrate the degree word dictionary and the negative word dictionary into the semantic modification dictionary; wherein, the semantic modification dictionary is used to parse the trigger word matching, modification conditions, and negative logic in the query text.

[0046] In this embodiment, Chinese trigger words and English trigger words are configured respectively according to the defined various intent types. For each intent type, such as search intent, file attribute intent, author intent and subject intent, keywords and phrases semantically related to the intent type are collected respectively, and a group of Chinese trigger words (such as "find", "search", "find", "find", etc.) are formed in the Chinese context, and corresponding English trigger words (such as "search", "find", "lookfor", etc.) are formed in the English context. Each group of trigger words is configured through rules to ensure that effective matching can be achieved in a multilingual environment. The Chinese trigger words and the English trigger words are unified and integrated into one to form a complete set of trigger words. The trigger word set is used for matching analysis of subsequent query texts to ensure that the system can recognize the intentions expressed by users across languages ​​and improve recognition coverage and accuracy. At the same time, a degree word dictionary and a negation word dictionary are established respectively to assist in semantic analysis in the process of intent recognition. The degree word dictionary includes expressions that reflect the intensity of intention, such as "very", "especially", "extremely", "very", "relatively", "a little", "slightly", "a little bit", etc., and configures corresponding semantic intensity weights for each word; the negation word dictionary includes expressions that express exclusion or negation of semantics, such as "not", "non", "have not", "don't", "exclude", "do not include", etc.

[0047] After the dictionary is built, the degree word dictionary and the negation word dictionary are integrated into a unified semantic modification dictionary. This semantic modification dictionary plays a role in the subsequent query parsing process, and is used to assist in parsing the modification conditions, contextual semantics of trigger words, and negative logical relationships in the query text, thereby achieving accurate understanding of complex query statements. Through the above operations, the system can build a trigger word set and semantic modification dictionary that adapts to multilingual scenarios, providing semantic support and rule basis for accurate and efficient intent recognition and file retrieval.

[0048] In step S103, the query text input by the user is semantically parsed based on the semantic modification dictionary to extract modification conditions. For example, the query text input may be first segmented and POS tagged to identify the modifying words therein, and their semantic role in the query is determined based on their POS and context position.

[0049] In one embodiment, the step S103 includes:

[0050] Performing word segmentation processing on the query text to generate a word segmentation result set;

[0051] Matching each word segmentation in the word segmentation result set with the semantic modification dictionary respectively, and recording the corresponding matching position;

[0052] Traverse the matching positions to extract degree words and negative words from the semantic modification dictionary;

[0053] Set corresponding intensity levels for the degree words and set negative identifiers for the negative words;

[0054] Generate the modification conditions based on the matching positions, using the intensity levels and negative identifiers.

[0055] In this embodiment, the input query text is tokenized, splitting the query text into multiple independent words or phrases to generate a tokenization result set. This tokenization process uses natural language processing techniques, which can effectively process query texts in different languages and of different types, ensuring the accuracy and integrity of the tokenization results. Each token in the generated tokenization result set is matched with a pre-constructed semantic modification dictionary. During the matching process, each token is searched one by one for its corresponding entry in the semantic modification dictionary, and the positions of each matching word are recorded. This helps to parse the context of the modifier later, so as to provide more accurate semantic parsing for the query statement.

[0056] Furthermore, traverse all the matching positions to extract degree words and negative words from the semantic modification dictionary. Degree words can represent the intensity or scope of semantics, such as "very", "slightly", "extremely", etc.; negative words are used to represent the meaning of exclusion or negation, such as "not", "non-", "exclude", etc. By traversing the matching positions, the degree words and negative words existing in the query text can be identified, preparing for subsequent semantic processing. Set corresponding intensity levels for the extracted degree words and set negative identifiers for the negative words. The intensity levels of degree words are set by rules. For example, "very" can correspond to a higher intensity level, while "slightly" corresponds to a lower intensity level. Negative words are identified according to their grammatical functions in the query text, usually with a "negative" identifier, which is used to indicate the exclusion effect of the word in semantic analysis. Finally, based on the above-mentioned matching positions, intensity levels and negative identifiers, the system generates modification conditions. The modification conditions include degree modification conditions and negative modification conditions, which are used to further precisify the semantic understanding of the user's query and ensure that the retrieval needs of the user are accurately reflected during the retrieval process.

[0057] In step S104, according to the intention recognition rule library and the configured trigger word set, perform pattern matching on the tokenization results to determine the intention type corresponding to the query text and generate a matching result. During this process, each trigger word is compared with the predefined rules, and the number of matches, positions and intensities are counted, and based on this, the confidence of each intention is calculated. The intention confidence is used to measure the credibility of the system's recognition of the intention type, which helps to determine the main intention in a multi-intention mixed query.

[0058] In one embodiment, the step S104 includes:

[0059] Perform pattern matching on the word segmentation result set and the trigger word set to obtain the trigger word matching count and the trigger word distribution positions corresponding to the intent type; wherein, the trigger word matching count and the trigger word distribution positions are the matching results.

[0060] Calculate the intent confidence based on a predefined confidence calculation rule using the trigger word matching count and the trigger word distribution positions.

[0061] In this embodiment, each word segment in the word segmentation result set is compared one by one with the predefined trigger words in the trigger word set to detect whether there are matching items. The trigger word set includes predefined keywords and phrases according to each intent type, such as search verbs, file attribute identifiers, author names, etc. Through pattern matching, the matching items corresponding to the words in the trigger word set in the query text can be identified. After the matching is completed, the matching count of the trigger words and the distribution positions of the trigger words in the query text are counted. The trigger word matching count represents the number of trigger words successfully matched in the query text, and the trigger word distribution positions indicate the specific positions of these matching words in the query text. The matching count and the distribution positions, as the matching results, provide the basic data for further calculating the intent confidence.

[0062] Furthermore, based on a predefined confidence calculation rule, use the trigger word matching count and the trigger word distribution positions to calculate the intent confidence. The confidence calculation rule needs to consider multiple factors. For example, the more the number of trigger word matches, the stronger the recognition of the user's intent; and the distribution position of the trigger words in the query also affects the credibility of the matching. For example, if the trigger words are located at the beginning or end of the query text, it indicates the dominant position of the intent in the query. Calculate the intent confidence comprehensively according to these rules to generate a value, which represents the reliability of the currently recognized intent type, that is, the intent confidence.

[0063] In step S105, construct a search query based on the identified modification conditions and the intent confidence, and perform a retrieval operation on the target file to obtain the retrieval results. During the retrieval process, the system can adopt strategies such as Boolean query, phrase query, and range query, and combine the modification conditions to limit the search scope. For example, exclude specific types of files or emphasize the importance of certain types of files, so as to improve the relevance and accuracy of the retrieval results.

[0064] In one embodiment, the step S105 includes:

[0065] Define the storage block positioning information of the disk;

[0066] Bypass the file system layer of the disk through the underlying data access interface to read the raw data stream of the disk;

[0067] Generate an aligned read instruction based on the storage block location information, and use the aligned read instruction to batch extract multiple data blocks of the disk to obtain discrete data blocks;

[0068] Perform logical recombination on the discrete data blocks to generate the target file;

[0069] Parse the content and metadata of the target file to construct a retrievable file object;

[0070] Create a file index model, dynamically associate the file object with the file index model to retrieve the residual data on the disk.

[0071] In this embodiment, the storage block location information of the disk is defined. The storage block location information is used to identify the storage location of the target file in the physical space of the disk, and specifically includes the start offset address and data block size of each file segment (i.e., file section). The storage block location information can be obtained by scanning the free area of the disk or analyzing the remaining file system metadata. Then, directly access the disk partition through the underlying data access interface, bypassing the traditional file system layer, and read the original data stream of the disk. Preferably, this operation can be implemented through the low-level volume access interface provided by the operating system (such as the volume GUID interface under the Windows platform), and the system thus obtains the underlying access permission to the disk sectors to access the data content that has been deleted but still remains on the disk. Generate an aligned read instruction based on the storage block location information, and use this aligned read instruction to batch extract multiple data blocks from the disk. The read operation adopts a memory alignment strategy, usually aligning the read range in units of the default page size of the operating system (such as 4096 bytes) to improve the disk read efficiency and system throughput. The system executes this batch read process to obtain several discrete original data blocks.

[0072] Furthermore, perform logical recombination on the discrete data blocks. According to the predefined order or paragraph identification information, splice multiple data blocks into a continuous data stream of the original file, thereby restoring the content of the target file. This recombination process is particularly suitable for recovering deleted or damaged files and can effectively reconstruct the original appearance of the file. After the recombination is completed, perform content parsing and metadata extraction on the target file. The parsing process automatically calls the corresponding file parser according to the file type, extracts structured information such as the text content, title, creation time, author, keywords, etc. of the file, and constructs a file object including the above information. Finally, create a file index model and dynamically associate the file object with this file index model. The file index model includes full-text index fields, filtering fields, storage fields, etc., and supports multi-dimensional search and retrieval optimization. After the file object is included in the index, it can participate in the search query based on intent analysis, thereby achieving precise retrieval of the disk residual data. Through the above technical means, it is possible to effectively complete the complete process from reading the underlying disk, data recombination, content parsing to index construction, and improve the recovery ability and retrieval coverage of deleted files.

[0073] In one embodiment, the bypassing the file system layer of the disk through the underlying data access interface to read the original data stream of the disk includes:

[0074] Obtain the volume identifier in the disk and call the underlying data access interface using the volume identifier;

[0075] Create a volume reader through the underlying data access interface and mount the volume reader to the physical storage layer of the disk;

[0076] Locate the original data area of the disk based on the volume reader to obtain the original data stream.

[0077] In this embodiment, the volume identifier in the disk is obtained. The volume identifier is the unique identifier of the disk partition, and it can uniquely identify a certain logical volume on the disk, thereby providing accurate positioning for subsequent data access. The obtained volume identifier is used to call the underlying data access interface. Through the underlying data access interface, the high-level abstraction of the file system can be bypassed, and direct interaction with the physical storage layer can be achieved. At this time, a volume reader is created through the underlying data access interface. The volume reader is a tool for accessing the physical data area of the disk, which can provide the ability to directly read the underlying data of the disk, avoiding the file management layer of the traditional file system, thereby improving the data access efficiency and enabling access to deleted or damaged data. Subsequently, the created volume reader is mounted to the physical storage layer of the disk. The mounting operation enables the volume reader to access the physical storage area of the disk, further ensuring efficient data reading. Through mounting, the volume reader can locate the original data area of the disk, including the free area of the disk and the data blocks of the deleted files, ensuring that the original data not managed by the file system can be read. Finally, based on the volume reader, the system locates the original data area of the disk and obtains the original data stream. The original data stream is the original data blocks stored in the disk that have not been processed by the file system. Through the original data stream, the file data not managed by the file system can be extracted for further data reorganization and recovery.

[0078] The file retrieval method based on intention recognition rules provided by the present invention has the following beneficial effects:

[0079] There is no need to rely on complex AI models. Fast intention recognition is achieved through rule matching, significantly reducing the consumption of computing resources and improving the response speed: The present invention performs intention recognition through a predefined rule matching method, avoiding the complex calculations and resource consumption of deep learning models, thereby achieving fast response. This method reduces the demand for computing resources and is suitable for application scenarios with high real-time requirements. It can directly read the disk partition data, retrieve and recover files that have been deleted but still have traces on the disk, expanding the scope of file retrieval: Different from traditional file retrieval methods, the present invention can bypass the file system and directly access the underlying data of the disk to read the file data that has been deleted but not yet overwritten on the disk. This enables the system to not only retrieve currently existing files but also recover deleted files, significantly expanding the coverage of file retrieval.

[0080] Furthermore, it also supports multi - language scenarios, accurately understanding the user's intention through predefined pattern - matching rules, and improving the retrieval accuracy: The system supports multiple language environments and can accurately understand the intention of the user's query through the trigger - word list and pattern - matching rules configured for each language. This design enhances the system's adaptability in multi - language environments, improves the accuracy and coverage of retrieval, and meets the cross - language retrieval requirements. It supports the processing of negative words and degree words, and can understand more complex query expressions, such as "find files that do not include tables", "find very important reports", etc.: The present invention supports the processing of negative words and degree words in the query text and can accurately parse complex query expressions such as "exclude certain types of files", "find highly relevant files", etc. This enhances the system's ability to understand and process complex user intentions and improves the relevance of retrieval results.

[0081] Finally, the present invention adopts a multi - process architecture, supports parallel processing, and improves the overall performance and throughput of the system. By adopting a multi - process architecture, the system can parallel - process retrieval and recovery tasks among multiple processes, thereby improving the processing capacity and throughput of the overall system. This design enables the system to efficiently process a large number of concurrent requests and is applicable to large - scale data - processing and high - load application scenarios.

[0082] Combined Figure 2 as shown Figure 2 FIG. is a schematic block diagram of a file retrieval device based on intention recognition rules provided by an embodiment of the present invention. The file retrieval device 200 based on intention recognition rules includes:

[0083] A rule preset unit 201 for predefined multiple intention types and constructing an intention recognition rule library;

[0084] A set configuration unit 202 for configuring a corresponding trigger - word set according to the intention type and establishing a semantic modification dictionary at the same time;

[0085] A text parsing unit 203 for semantic - parsing the input query text according to the semantic modification dictionary to obtain modification conditions;

[0086] A word matching unit 204 for matching the intention type using the trigger - word set based on the intention recognition rule library to obtain a matching result and generating an intention confidence level according to the matching result at the same time;

[0087] A target retrieval unit 205 for retrieving target files based on the modification conditions and intention confidence level to obtain retrieval results.

[0088] In this embodiment, the rule preset unit 201 predefined multiple intent types and constructed an intent recognition rule library; the set configuration unit 202 configured a corresponding trigger word set according to the intent type, and at the same time established a semantic modification dictionary; the text parsing unit 203 performed semantic parsing on the input query text according to the semantic modification dictionary to obtain a modification condition; the word matching unit 204 matched the intent type by using the trigger word set based on the intent recognition rule library to obtain a matching result, and at the same time generated an intent confidence level according to the matching result; the target retrieval unit 205 retrieved the target file based on the modification condition and the intent confidence level to obtain a retrieval result.

[0089] In one embodiment, the rule preset unit 201 includes:

[0090] The first definition unit is used to define the search intent type to establish a search action set for identifying the user's search request;

[0091] The second definition unit is used to define the file attribute intent type to establish a file type set for identifying the file type request;

[0092] The third definition unit is used to define the author intent type to establish an author identification set for identifying the author query request;

[0093] The fourth definition unit is used to define the theme intent type to establish a theme keyword set for identifying the theme query request;

[0094] The type integration unit is used to store the search intent type, file attribute intent type, author intent type, and theme intent type and their corresponding sets into the intent recognition rule library.

[0095] In one embodiment, the set configuration unit 202 includes:

[0096] The type configuration unit is used to configure Chinese trigger words and English trigger words according to the intent type respectively;

[0097] The type set unit is used to integrate the Chinese trigger words and English trigger words into the trigger word set;

[0098] The dictionary integration unit is used to establish an intensifier dictionary and a negation dictionary respectively, and integrate the intensifier dictionary and the negation dictionary into the semantic modification dictionary; wherein, the semantic modification dictionary is used to parse the trigger word matching, modification condition, and negation logic in the query text.

[0099] In one embodiment, the text parsing unit 203 includes:

[0100] A word segmentation processing unit for performing word segmentation processing on the query text to generate a word segmentation result set;

[0101] A word segmentation matching unit for respectively matching each word in the word segmentation result set with the semantic modification dictionary and recording the corresponding matching positions;

[0102] A position traversal unit for traversing the matching positions to extract intensifiers and negators in the semantic modification dictionary;

[0103] A word setting unit for setting corresponding intensity levels for the intensifiers and setting negation identifiers for the negators;

[0104] A condition generation unit for generating the modification conditions based on the matching positions by using the intensity levels and negation identifiers.

[0105] In one embodiment, the word matching unit 204 includes:

[0106] A set matching unit for performing pattern matching on the word segmentation result set and the trigger word set to obtain the trigger word matching quantity corresponding to the intention type and the trigger word distribution positions; wherein, the trigger word matching quantity and the trigger word distribution positions are the matching results;

[0107] A confidence calculation unit for calculating the intention confidence based on a predefined confidence calculation rule by using the trigger word matching quantity and the trigger word distribution positions.

[0108] In one embodiment, the target retrieval unit 205 includes:

[0109] A disk definition unit for defining the storage block positioning information of the disk;

[0110] An interface access unit for bypassing the file system layer of the disk through a low-level data access interface to read the original data stream of the disk;

[0111] A data extraction unit for generating an aligned read instruction based on the storage block positioning information and batch extracting multiple data blocks of the disk by using the aligned read instruction to obtain discrete data blocks;

[0112] A logical recombination unit for performing logical recombination on the discrete data blocks to generate the target file;

[0113] An object construction unit for parsing the content and metadata of the target file to construct a retrievable file object;

[0114] A model creation unit for creating a file index model, dynamically associating the file object with the file index model to retrieve residual data on the disk.

[0115] In one embodiment, the interface access unit includes:

[0116] An identification acquisition unit for acquiring a volume identifier in the disk and invoking the underlying data access interface using the volume identifier;

[0117] An interface mounting unit for creating a volume reader through the underlying data access interface and mounting the volume reader to the physical storage layer of the disk;

[0118] A data reading unit for positioning the original data area of the disk based on the volume reader to obtain the original data stream.

[0119] Since the embodiments in the device part correspond to the embodiments in the method part, for the embodiments in the device part, please refer to the description of the embodiments in the method part, which will not be elaborated here.

[0120] An embodiment of the present invention also provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0121] An embodiment of the present invention also provides a computer device, which may include a memory and a processor. When the processor invokes the computer program in the memory, the steps provided in the above embodiments can be implemented. Of course, the computer device may also include various network interfaces, power supplies, and other components.

[0122] The embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, please refer to the description in the method part. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

[0123] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

Claims

1. A file retrieval method based on intention recognition rules, characterized in that, Including: Pre - define multiple intent types and construct an intent recognition rule library; Configure a corresponding set of trigger words according to the intent types, and at the same time establish a semantic modification dictionary; Perform semantic parsing on the input query text according to the semantic modification dictionary to obtain modification conditions; Based on the intent recognition rule library, use the set of trigger words to match the intent types to obtain a matching result, and at the same time generate an intent confidence level according to the matching result; Retrieve the target file based on the modification conditions and the intent confidence level to obtain a retrieval result.

2. The file retrieval method based on intention recognition rules according to claim 1, wherein, The retrieving the target file based on the modification conditions and the intent confidence level to obtain a retrieval result includes: Define the storage block location information of the disk; Bypass the file system layer of the disk through the underlying data access interface to read the original data stream of the disk; Generate an aligned read instruction based on the storage block location information, and use the aligned read instruction to batch - extract multiple data blocks of the disk to obtain discrete data blocks; Perform logical recombination on the discrete data blocks to generate the target file; Parse the content and metadata of the target file to construct a retrievable file object; Create a file index model, dynamically associate the file object with the file index model to retrieve the residual data of the disk.

3. The file retrieval method based on an intent recognition rule according to claim 1, wherein The pre - defining multiple intent types and constructing an intent recognition rule library includes: Define a search intent type to establish a set of search actions for identifying user search requests; Define a file attribute intent type to establish a set of file types for identifying file type requests; Define an author intent type to establish a set of author identifiers for identifying author query requests; Define a theme intent type to establish a set of theme keywords for identifying theme query requests; Store the search intent type, file attribute intent type, author intent type, theme intent type and their corresponding sets into the intent recognition rule library.

4. The file retrieval method based on an intention recognition rule according to claim 1, wherein The configuring a corresponding set of trigger words according to the intent types and at the same time establishing a semantic modification dictionary includes: Configure Chinese trigger words and English trigger words respectively according to the intent types; Integrate the Chinese trigger words and English trigger words into the set of trigger words; Respectively establish a degree word dictionary and a negation word dictionary, and integrate the degree word dictionary and the negation word dictionary into the semantic modification dictionary; wherein, the semantic modification dictionary is used to parse the trigger word matching, modification conditions and negation logic in the query text.

5. The file retrieval method based on an intention recognition rule according to claim 1, wherein The performing semantic parsing on the input query text according to the semantic modification dictionary to obtain modification conditions includes: Perform word - segmentation processing on the query text to generate a set of word - segmentation results; Match each word in the set of word - segmentation results with the semantic modification dictionary respectively, and record the corresponding matching positions; Traverse the matching positions to extract degree words and negation words in the semantic modification dictionary; Set corresponding intensity levels for the degree words and set negation identifiers for the negation words; Based on the matching positions, generate the modification conditions using the intensity levels and negation identifiers.

6. The method for file retrieval based on an intent recognition rule according to claim 5, characterized in that, Using the trigger word set to match the intent type based on the intent recognition rule library to obtain a matching result, and generating an intent confidence level according to the matching result, including: Performing pattern matching between the word segmentation result set and the trigger word set to obtain the number of trigger word matches corresponding to the intent type and the distribution positions of the trigger words; wherein, the number of trigger word matches and the distribution positions of the trigger words are the matching results; Calculating the intent confidence level based on the predefined confidence level calculation rule using the number of trigger word matches and the distribution positions of the trigger words.

7. The method for retrieving a document based on an intention recognition rule according to claim 2, wherein Bypassing the file system layer of the disk through the underlying data access interface to read the original data stream of the disk, including: Obtaining the volume identifier in the disk and calling the underlying data access interface using the volume identifier; Creating a volume reader through the underlying data access interface and mounting the volume reader to the physical storage layer of the disk; Locating the original data area of the disk based on the volume reader to obtain the original data stream.

8. A file retrieval device based on intention recognition rules, characterized in that Including: A rule preset unit for predefined multiple intent types and constructing an intent recognition rule library; A set configuration unit for configuring a corresponding trigger word set according to the intent type and establishing a semantic modification dictionary at the same time; A text parsing unit for performing semantic parsing on the input query text parsing according to the semantic modification dictionary to obtain a modification condition; A word matching unit for using the trigger word set to match the intent type based on the intent recognition rule library to obtain a matching result, and generating an intent confidence level according to the matching result; A target retrieval unit for retrieving a target file based on the modification condition and the intent confidence level to obtain a retrieval result.

9. A computer device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the file retrieval method based on intent recognition rules according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the file retrieval method based on intent recognition rules according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Multi-modal data semantic retrieval method and device, equipment and storage medium

    CN120950705A

  • A multi-modal data semantic retrieval method, device, equipment and storage medium

    CN120950705B