File retrieval method and device, computer program product and electronic equipment

Through semantic vectorized encoding and vector database retrieval, the problem of inefficiency of users when finding specific documents is solved, and efficient document retrieval and positioning is achieved.

CN119938609APending Publication Date: 2025-05-06INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411974378.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, users can find specific documents through document search tools inefficiently, especially when the file name is blurred or the document content is complicated.

Method used

By receiving the object file description content uploaded by the user, semantic vectorized encoding is performed, and similar encoding is retrieved in the vector database to obtain key information of the candidate files, and finally determine its file location based on the target file selected by the user.

Benefits of technology

The technical effect of efficiently finding target documents from a large number of documents stored in the device is realized, and the efficiency of users in finding specific documents is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938609A_ABST
    Figure CN119938609A_ABST
Patent Text Reader

Abstract

The invention discloses a file retrieval method and device, a computer program product and electronic equipment. Relates to the field of artificial intelligence, and the method comprises the following steps: receiving description content of a target file uploaded by a user through a client, and carrying out semantic vectorization coding on the description content to obtain a description code; retrieving all vectorized codes which have the same part with the description codes in a vector database to obtain candidate codes; acquiring key information of a file to which the candidate code belongs, and sending the key information to the client; receiving a file selection instruction sent by a user through the client, and obtaining a target file selected by the user from the file selection instruction; and obtaining the file position of the target file from the structured database according to the file identifier of the target file, and sending the file position to the client. Through the method and the device, the problem of low efficiency of searching a specific document by a user through a document search tool in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and more specifically, to a file retrieval method, device, computer program product, and electronic device. Background Art

[0002] With the popularization of digital office, the number of documents stored in personal computers has increased dramatically, resulting in severe challenges in document management and retrieval. Specifically, most existing file search tools can only perform simple matching based on file names or keywords, which makes the search process time-consuming and laborious when the file name is vague or the document content is complex. On the one hand, the increase in local document storage makes file name-based searches inefficient. Once the file name is not remembered accurately, it is difficult to find the required document. On the other hand, even if the relevant file is found, it is often too long, which further increases the user's time and energy burden.

[0003] With regard to the problem in related art that users are inefficient in finding specific documents through document search tools, no effective solution has been proposed yet. Summary of the invention

[0004] The main purpose of the present application is to provide a file retrieval method, device, computer program product and electronic device to solve the problem of low efficiency of users finding specific documents through document search tools in the related art.

[0005] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a file retrieval method is provided. The method includes: receiving the description content of the target file uploaded by the user through the client, and performing semantic vectorization encoding on the description content to obtain the description code; retrieving all vectorized codes in the vector database that have the same part as the description code to obtain candidate codes, wherein the vector database stores the semantic vectorized codes of the files in the device drive letter; obtaining key information of the files to which the candidate codes belong, and sending the key information to the client; receiving the file selection instruction sent by the user through the client, and obtaining the target file selected by the user from the file selection instruction; obtaining the file location of the target file from the structured database according to the file identifier of the target file, and sending the file location to the client, wherein the structured database stores the structured information of the files in the device drive letter, and the structured information at least includes the file location.

[0006] Optionally, obtaining key information of the file to which the candidate encoding belongs includes: calculating the similarity score between the file to which the candidate encoding belongs and the description content, and when there are multiple candidate encodings, calculating the similarity ranking based on the similarity scores associated with the multiple candidate encoding files; calling the large language model to summarize the file to which the candidate encoding belongs to obtain a document summary of the file to which the candidate encoding belongs; and determining at least one of the similarity score, similarity ranking and document summary as key information.

[0007] Optionally, calculating the similarity score between the file to which the candidate code belongs and the description content includes: setting a quantity weight coefficient and a ranking weight coefficient, wherein the quantity weight coefficient represents the degree of influence of the number of similar file fragments between the two files on the similarity score, and the ranking weight represents the degree of influence of the similarity ranking score of a single file fragment on the similarity score; for a candidate code, determining the number of similar file fragments between the file to which the candidate code belongs and the description content, determining the similarity ranking score of each similar file fragment, and performing weighted summation of the number of similar file fragments and the similarity ranking score of each similar file fragment based on the quantity weight coefficient and the ranking weight coefficient to obtain the similarity score between the file to which the candidate code belongs and the description content.

[0008] Optionally, before receiving the description content of the target file uploaded by the user through the client, the method also includes: scanning the files in the device drive letter, and judging whether the file is an incremental file based on the numerical summary of the file; if the file is an incremental file, determining the semantic vectorization encoding of the incremental file, determining the file fragments of the incremental file, and storing the vectorization encoding results and the associated file fragments in a vector database; reading the structured information of the incremental file, and storing the structured information in a structured database; establishing a dictionary mapping between the vector database and the structured database, wherein the dictionary mapping maps the vectorization encoding and structured information of the same file through a file identifier.

[0009] Optionally, determining whether a file is an incremental file based on the numerical summary of the file includes: determining whether the numerical summary has appeared in the device drive letter; if the numerical summary has appeared in the device drive letter, determining that the file is not an incremental file; if the numerical summary has not appeared in the device drive letter, determining that the file is an incremental file.

[0010] Optionally, determining the semantic vectorization encoding of the incremental file includes: parsing the content of the incremental file and mapping the content of the incremental file to a vector space to obtain a mapping content vector; extracting deep semantic features of the mapping content vector to obtain the semantic vectorization encoding of the incremental file.

[0011] Optionally, determining the file fragments of the incremental file includes: parsing the content and structure of the incremental file, wherein the structure includes at least one of the following: a title, a number of paragraphs, and a directory; and splitting the content according to the structure to obtain the file fragments of the incremental file.

[0012] In order to achieve the above-mentioned purpose, according to another aspect of the present application, a file retrieval device is provided. The device includes: an encoding module, which is used to receive the description content of the target file uploaded by the user through the client, and semantically vectorize the description content to obtain the description code; a retrieval module, which is used to retrieve all vectorized codes in the vector database that have the same part as the description code, and obtain candidate codes, wherein the vector database stores the semantic vectorized codes of the files in the device drive letter; a sending module, which is used to obtain the key information of the file to which the candidate code belongs, and send the key information to the client; an instruction receiving module, which is used to receive the file selection instruction sent by the user through the client, and obtain the target file selected by the user from the file selection instruction; a location sending module, which is used to obtain the file location of the target file from the structured database according to the file identifier of the target file, and send the file location to the client, wherein the structured database stores the structured information of the files in the device drive letter, and the structured information at least includes the file location.

[0013] In order to achieve the above object, according to another aspect of the present application, a computer program product is provided, which includes: a non-volatile computer-readable storage medium storing a computer program, and the computer program implements the file retrieval method when executed by a processor.

[0014] In order to achieve the above object, according to another aspect of the present application, an electronic device is provided. The device includes: a computer program is stored in a memory, and a processor is configured to execute a file retrieval method through the computer program.

[0015] In an embodiment of the present application, the description content of the target file uploaded by the user through the client is received, and the description content is semantically vectorized and encoded to obtain the description code; all vectorized codes in the vector database that have the same part as the description code are retrieved to obtain candidate codes, wherein the vector database stores the semantic vectorized codes of the files in the device drive letter; key information of the files to which the candidate codes belong is obtained, and the key information is sent to the client; a file selection instruction sent by the user through the client is received, and the target file selected by the user is obtained from the file selection instruction; the file location of the target file is obtained from the structured database according to the file identifier of the target file, and the file location is sent to the client, wherein the structured database stores the structured information of the files in the device drive letter, and the structured information at least includes the file location. The semantic vectorized codes of similar files are searched from the vector database through the semantic vectorized codes of the description content, and for the target file selected by the user, the file location is determined based on the semantic vectorized codes of the target file combined with the structured information, so as to realize efficient retrieval and positioning of the target file, and achieve the technical effect of efficiently finding the target document from a large number of documents stored in the device, thereby solving the problem of low efficiency of users finding specific documents through document search tools in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0017] Figure 1 It is a hardware structure block diagram of a computer terminal for implementing a file retrieval method;

[0018] Figure 2 is a flowchart of a file retrieval method provided according to an embodiment of the present application;

[0019] Figure 3 is a schematic diagram of an optional file retrieval method provided according to an embodiment of the present application;

[0020] Figure 4 The following is a flow chart of an optional file retrieval method provided in accordance with an embodiment of the present application. Figure 1 ;

[0021] Figure 5 is a schematic diagram of a file retrieval device provided according to an embodiment of the present application;

[0022] Figure 6 It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data are in compliance with relevant laws, regulations and standards, necessary confidentiality measures are taken, and public order and good customs are not violated, and corresponding operation entrances are provided for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions to provide users with corresponding operation entrances for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, the expert decision-making process will be entered.

[0026] Example 1

[0027] According to an embodiment of the present application, a method embodiment of file retrieval is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0028] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 It is a hardware structure block diagram of a computer terminal used to implement a file retrieval method. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0029] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0030] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the file retrieval method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned file retrieval method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0031] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0032] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0033] Under the above operating environment, this application provides Figure 2 The file retrieval method shown. Figure 2 It is a flowchart of a file retrieval method provided according to an embodiment of the present application.

[0034] Step S201, receiving the description content of the target file uploaded by the user through the client, and performing semantic vectorization encoding on the description content to obtain a description code.

[0035] Specifically, the executor of this embodiment may be a retrieval system, and the description content of the target file may refer to the descriptive information related to the content of the target file input by the user through the client software (such as a special retrieval tool or application) when the user needs to find the target file, wherein the client may be at least one of the user's computer application, mobile phone application and web page interface. The description content can be presented in the form of keywords, sentences or paragraphs to express the user's retrieval needs. Semantic vectorization encoding refers to the process by which the retrieval system converts the description content input by the user into a vector representation. This process relies on natural language processing technology and can extract deep semantic features of the description content. Compared with simple character matching, semantic vectorization encoding can more comprehensively reflect the semantic relevance of the text and provide more accurate reference data for subsequent retrieval.

[0036] Step S202, retrieve all vectorized codes in the vector database that have the same parts as the description code to obtain candidate codes, wherein the vector database stores the semantic vectorized codes of the files in the device drive letter.

[0037] Specifically, the retrieval system uses the description code to perform similarity matching in the vector database, looking for vectorized codes that have overlapping or similar features with the description code, for example, calculating their distance or similarity in the vector space, thereby determining the similarity between them.

[0038] The vector database is a data structure based on high-dimensional vector storage and retrieval, which is used to save the codes generated by the files in the device drive letter after semantic vectorization processing. Semantic vectorization coding refers to converting the text content of the files in the device drive letter into a vector expression that can represent semantic features through a natural language processing model. This expression can capture the deep semantic features of the file content in the vector space.

[0039] Obtaining candidate codes means that the retrieval system calculates the similarity between the description code and all codes in the vector database, and selects a group of vectorized codes that are closest to the description code, namely, candidate codes.

[0040] Step S203, obtaining key information of the file to which the candidate code belongs, and sending the key information to the client.

[0041] Specifically, each candidate code corresponds to a file in a device drive letter, which is the file to which the candidate code belongs. The key information of the file to which the candidate code belongs may refer to at least one of the similarity score, similarity ranking, and document summary. The key information can help users quickly understand the similarity and general content of the files, thereby improving user friendliness and user selection efficiency. The retrieval system summarizes the key information and sends it to the client, so that the user can view it and proceed to the next step.

[0042] Step S204: receiving a file selection instruction sent by the user through the client, and obtaining a target file selected by the user from the file selection instruction.

[0043] Specifically, before the retrieval system receives the instruction, the user can view the key information of the candidate file through the interactive interface of the client, and then issue a file selection instruction to clearly indicate the desired target file. The file selection instruction can be a data request, including the identification information of the target file selected by the user, which can be transmitted to the retrieval system in the form of a unique file identifier (e.g., a unique number).

[0044] Step S205, obtaining the file location of the target file from the structured database according to the file identifier of the target file, and sending the file location to the client, wherein the structured database stores structured information of the file in the device drive letter, and the structured information at least includes the file location.

[0045] Specifically, after the user selects the target file, the retrieval system performs a query based on the unique identifier contained in the file selection instruction. In the structured database, the file identifier is used as a key field and mapped with the structured information of the file. The structured information may include data such as the storage location and modification time of the file, and can provide the specific storage location of the target file in the device.

[0046] The retrieval system retrieves the structured database to obtain structured information, locates the storage location of the target file selected by the user, and returns the file location information to the client as response data. After the user receives the storage location information through the client, he can directly access the target file. For example, the user can quickly open the file according to the file path and directly view or edit it.

[0047] In an embodiment of the present application, the description content of the target file uploaded by the user through the client is received, and the description content is semantically vectorized and encoded to obtain the description code; all vectorized codes in the vector database that have the same part as the description code are retrieved to obtain candidate codes, wherein the vector database stores the semantic vectorized codes of the files in the device drive letter; key information of the files to which the candidate codes belong is obtained, and the key information is sent to the client; a file selection instruction sent by the user through the client is received, and the target file selected by the user is obtained from the file selection instruction; the file location of the target file is obtained from the structured database according to the file identifier of the target file, and the file location is sent to the client, wherein the structured database stores the structured information of the files in the device drive letter, and the structured information at least includes the file location. The semantic vectorized codes of similar files are searched from the vector database through the semantic vectorized codes of the description content, and for the target file selected by the user, the file location is determined based on the semantic vectorized codes of the target file combined with the structured information, so as to realize efficient retrieval and positioning of the target file, and achieve the technical effect of efficiently finding the target document from a large number of documents stored in the device, thereby solving the problem of low efficiency of users finding specific documents through document search tools in the related technology.

[0048] In order to enable users to filter out desired target files, optionally, in the file retrieval method provided in the embodiment of the present application, obtaining key information of the file to which the candidate encoding belongs includes: calculating the similarity score between the file to which the candidate encoding belongs and the description content, and when the number of candidate encodings is multiple, calculating the similarity ranking based on the similarity scores associated with the files to which the multiple candidate encodings belong; calling the large language model to summarize the files to which the candidate encoding belongs to obtain the document summary of the file to which the candidate encoding belongs; and determining at least one of the similarity score, similarity ranking and document summary as key information.

[0049] Specifically, after the retrieval system obtains the candidate code, it evaluates the semantic similarity between the file to which the candidate code belongs and the user's description content through a vector similarity calculation algorithm (e.g., a calculation method based on semantic vectorization), thereby generating a similarity score. The similarity score reflects the degree of match between the file content and the user's needs.

[0050] If the system retrieves multiple candidate files related to the description, it will sort them based on the similarity score and calculate the similarity ranking of the files. Then, the retrieval system calls the large language model to generate document summaries for the candidate files with the highest similarity ranking. The large language model can parse the document content, extract the core information in the file, and generate a document summary, so that users can quickly understand the content of the file and determine whether it meets their needs.

[0051] The embodiment of the present application introduces the steps of calculating similarity scores, rankings and generating document summaries, which not only reduces the time for reading and screening, but also improves the interactivity and intelligence of file retrieval and shortens the time it takes for users to find target files.

[0052] In order to make the similarity score and ranking have sufficient reference value, optionally, in the file retrieval method provided in the embodiment of the present application, calculating the similarity score between the file to which the candidate code belongs and the description content includes: setting a quantity weight coefficient and a ranking weight coefficient, wherein the quantity weight coefficient represents the degree of influence of the number of similar file fragments between the two files on the similarity score, and the ranking weight represents the degree of influence of the similarity ranking score of a single file fragment on the similarity score; for a candidate code, determining the number of similar file fragments between the file to which the candidate code belongs and the description content, determining the similarity ranking score of each similar file fragment, and performing weighted summation of the number of similar file fragments and the similarity ranking score of each similar file fragment based on the quantity weight coefficient and the ranking weight coefficient to obtain the similarity score between the file to which the candidate code belongs and the description content.

[0053] Specifically, before calculating the similarity score, two important weight coefficients must be set first, namely the "quantity weight coefficient" and the "ranking weight coefficient". The quantity weight coefficient (denoted as w1) reflects the degree of influence of the number of similar segments shared between two files on the final similarity score, which means that if two documents have more similar semantic segments, the similarity between them will be given a higher evaluation. The ranking weight coefficient (denoted as w2) measures the contribution of the ranking position of a single similar segment in all retrieval results to the final score, that is, the higher the ranking of the similar segment in the retrieval results, the greater the effect on improving the overall similarity of the document.

[0054] In actual calculation, the weight coefficient of the number of fragments (w1) and the weight coefficient of the ranking of the fragments in the search results (w2) can be combined to perform weighted summation on the similarity score of each file, and finally obtain the comprehensive similarity ranking of each file. The weighted calculation formula can be:

[0055] similitude=w1×Num+w2×∑rank_score

[0056] similitude is the similarity score, w1 represents the set quantity weight coefficient, that is, the weight coefficient of the number of similar fragments in the calculation of the final similarity. The larger the setting, the greater the influence of more similar fragments on the final similarity ranking; Num represents the number of similar fragments of an article retrieved from the vector database, for example, it can be 100; w2 represents the set ranking weight coefficient, that is, the weight coefficient of the total ranking score of similar fragments in 100 fragments in the calculation of the final document similarity. The larger the setting, the higher the ranking of similar fragments in the top 100 fragments, the greater the influence on the final document similarity ranking; rank_score represents the ranking score of a single fragment. For example, from top to bottom, the score of the first-ranked fragment can be set to 100, the score of the 100th-ranked fragment can be set to 1, and so on.

[0057] The embodiment of the present application introduces the concepts of quantity weight coefficient and ranking weight coefficient, which comprehensively considers the impact of the number of similar file fragments and their similarity ranking on the total similarity score, so that the similarity score more comprehensively and accurately reflects the degree of match between the candidate file and the description content, thereby improving the reference value of the retrieval results and the accuracy of the sorting.

[0058] In order to grasp the changes of incremental files in the device drive letter and speed up the file query speed, optionally, in the file retrieval method provided in the embodiment of the present application, before receiving the description content of the target file uploaded by the user through the client, the method also includes: scanning the files in the device drive letter, and judging whether the file is an incremental file based on the numerical summary of the file; if the file is an incremental file, determining the semantic vectorization encoding of the incremental file, determining the file fragments of the incremental file, and storing the vectorization encoding results and the associated file fragments in the vector database; reading the structured information of the incremental file, and storing the structured information in the structured database; establishing a dictionary mapping between the vector database and the structured database, wherein the dictionary mapping maps the vectorization encoding and structured information of the same file through the file identifier.

[0059] Specifically, starting from the startup or regular update of the retrieval system, the retrieval system performs a full scan of all drive letters on the user's device and checks the numerical summary of each file one by one. The numerical summary is generated by calculating the hash value or checksum of the file content. By comparing the numerical summary of the current file with the existing numerical summary, it is determined whether the file is an incremental file, that is, whether it is a newly added or updated file. If it is an incremental file, it will be included in the database update process of the retrieval system.

[0060] For documents that are judged to be incremental files, they are further processed to generate their semantic vectorized codes. Semantic vectorized codes use large language models, such as natural language processing technology, to deeply analyze the file content and convert text information into semantic vector representations. This vector form not only retains the semantic features of the text, but also facilitates semantic matching in subsequent retrieval. At the same time, the retrieval system will split the document into multiple file segments and generate vectorized codes for each segment separately to improve retrieval accuracy and flexibility.

[0061] After completing the vectorization encoding, the retrieval system stores the vectorization encoding results and the corresponding file fragment information in the vector database to form a dynamically updated semantic vector set. In addition, the retrieval system will also extract the structured information of the incremental file, including file name, storage location, modification time and other structural data, and store this information in the structured database to assist in the location and display of the retrieval results.

[0062] The retrieval system establishes a dictionary mapping mechanism between the vector database and the structured database to ensure the consistency of the semantic vector and the file metadata. The relationship between the semantic vectorized code and the structured information is recorded by sharing the ID method, and the semantic vectorized code is associated with the corresponding structured information to form a one-to-one relationship. Each ID can be unique in the retrieval system and can be called a unique code. When the user searches, the system first matches based on the semantic vector, and after obtaining the user's selection result, it quickly locates the storage location of the file through dictionary mapping to achieve efficient retrieval response.

[0063] The embodiment of the present application performs file scanning and incremental updates on the device before the user initiates a search, determines the file update status through a numerical summary, and at the same time, performs semantic vectorization and structured information extraction on the incremental files. Finally, through a dictionary mapping between the vector database and the structured database, the correspondence between the semantic matching and the file storage location is achieved, making file search both fast and accurate.

[0064] In order to ensure that the file retrieval method can reflect the changes of files on the computer in real time, optionally, in the file retrieval method provided in the embodiment of the present application, judging whether a file is an incremental file based on the numerical summary of the file includes: judging whether the numerical summary has appeared in the device drive letter; if the numerical summary has appeared in the device drive letter, determining that the file is not an incremental file; if the numerical summary has not appeared in the device drive letter, determining that the file is an incremental file.

[0065] Specifically, the numerical summary is a digital fingerprint generated by calculating the hash value of the file content, which can accurately reflect the integrity of the file content. Even if the file name or storage location changes, as long as the file content does not change, its numerical summary remains unchanged.

[0066] The retrieval system maintains a record list of file numerical summaries for comparison with historical file status. If the numerical summary of the currently scanned file matches any summary in the list, it means that the file content is consistent with the existing files in the system and has not changed. The system determines that the file is not an incremental file, so there is no need to update the information in the vector database and structured database. If the numerical summary of the file does not find a match in the record list, it means that the file content is new or modified, and it is marked as an incremental file. For incremental files, the system will execute the steps of vectorization encoding and structured information extraction, perform semantic vectorization on the file content, generate a semantic vector representation of the file, and store it in the vector database; read the structured information of the file, and store it in the structured database.

[0067] The embodiment of the present application monitors and processes file changes in a personal computer in real time, quickly identifies whether a file is an incremental file, avoids repeated processing of unchanged files, saves computing resources, and improves system efficiency. For the determined incremental files, the retrieval system further performs vectorized encoding and structured information extraction to ensure that the information of the incremental files is updated to the retrieval database in a timely manner, thereby improving the accuracy of file matching and retrieval efficiency.

[0068] In order to generate a vector code that accurately matches the file content, optionally, in the file retrieval method provided in the embodiment of the present application, determining the semantic vector code of the incremental file includes: parsing the content of the incremental file, and mapping the content of the incremental file to the vector space to obtain a mapped content vector; extracting the deep semantic features of the mapped content vector to obtain the semantic vector code of the incremental file.

[0069] Specifically, when the retrieval system performs a full scan of the device, it will identify new or updated incremental files, and perform semantic parsing and encoding on these files. The natural language processing model is used to vectorize the file content. The model can capture the semantic features of the text and convert the text into a numerical vector representation to facilitate subsequent similarity calculation and analysis. On the basis of generating the mapping content vector, the mapping content vector is deeply processed to extract high-level features that reflect the semantics of the text, thereby generating more expressive semantic vectorized codes. These codes contain not only the basic meaning of the text, but also deep information such as context and semantic associations.

[0070] The embodiment of the present application ensures the dynamic update of the vector database and the retrieval accuracy of the retrieval system by performing real-time and accurate semantic analysis and encoding of incremental files. It only performs semantic analysis on new or modified files, reduces repeated processing of unchanged files, saves computing resources, and can significantly improve the accuracy of file retrieval.

[0071] In order to obtain file fragments that are easy to record and compare, optionally, in the file retrieval method provided in the embodiment of the present application, determining the file fragments of the incremental file includes: parsing the content and structure of the incremental file, wherein the structure includes at least one of the following: title, number of paragraphs and directory; splitting the content according to the structure to obtain file fragments of the incremental file.

[0072] Specifically, after the retrieval system scans the incremental files on the personal computer, it first needs to parse these files to understand their internal structure and detailed content. The file structure can include elements such as title, number of paragraphs, directory, etc. These structural information can help the system understand the organization and important parts of the document, so as to be more accurate in the subsequent file segmentation.

[0073] Split a file into multiple file segments based on the content and structure information of the file. The file segments can be divided according to structural elements such as title, directory or number of paragraphs. For example, when processing a document, the document can be split into multiple segments based on the granularity of paragraphs, number of characters or number of pages. Each segment can be expressed as a string similar to the following format: {"title":xxx,"paragraph":xxx,"index":2}, where "title" indicates the title of the segment, "paragraph" indicates the content of the segment, and "index" indicates the number of the segment.

[0074] The embodiment of the present application parses the content and structure of the incremental file, combines the title, number of paragraphs, directory and other information to perform precise splitting, and generates file fragments that are easy to record and compare, thereby optimizing file organization and retrieval efficiency and facilitating subsequent file comparison.

[0075] According to an embodiment of the present application, an optional file retrieval method is also provided. Figure 3 is a schematic diagram of an optional file retrieval method provided according to an embodiment of the present application, such as Figure 3 As shown: The following steps are included:

[0076] Incremental document loading: Receive target files uploaded by users through the client, which contain the documents that need to be retrieved.

[0077] Document splitting: Split the uploaded target file into multiple segments for subsequent semantic processing. Each segment contains a certain amount of text information, which facilitates more fine-grained analysis of the file content.

[0078] Document storage: stores the semantic vector and document location of the file: Semantic vector storage: saves the vector representation of the file fragment generated by semantic vectorization encoding for subsequent similarity calculation; Document location storage: saves the specific storage path of the target file on the device to facilitate rapid positioning by the retrieval system.

[0079] Document retrieval: Receive user query information and match the user's description with the stored semantic vectors. User descriptions can include keywords, sentences, or paragraph descriptions. The retrieval system achieves efficient retrieval through vectorized matching.

[0080] Call the large model summary comparison module: Generate a candidate file set based on the matching results, further analyze the candidate files through the large language model, calculate the similarity score between the candidate files and the user query, and generate a similarity ranking.

[0081] Result: The candidate files with the highest similarity ranking and their key information (including similarity score, ranking and document summary) are returned to the client. The user can use the key information to determine whether the file meets the requirements and obtain the file location to open the target file.

[0082] According to an embodiment of the present application, an optional file retrieval method is also provided. Figure 4 The following is a flow chart of an optional file retrieval method provided in accordance with an embodiment of the present application. Figure 1 ,like Figure 4 As shown, it shows the processing process from user input description to the system returning the file location:

[0083] Step S401: The retrieval system scans all files in the personal computer and builds a document vector database and a document information structured database. The main purpose of this step is to obtain the file content and perform preliminary processing on it by scanning the entire file to provide basic data support for subsequent retrieval.

[0084] Step S402: The user describes the content of the document to be searched. The user inputs a search request through the search system and provides descriptive text to describe the content characteristics of the target document.

[0085] Step S403: The retrieval system vectorizes the description text and performs vector matching with the content in the document vector database. By calculating similarity, documents with similar semantics are found and these documents are sorted by relevance.

[0086] Step S404: The retrieval system further processes the N similarity ranking (TOP N) documents selected in the first layer of filtering, and calls the Large Language Model (LLM) to generate a document summary for each document and display it to the user, who then determines which document best meets his or her needs. The core of this step is to use the intelligent model to assist the user in quickly narrowing down the scope of the target document.

[0087] Step S405: The retrieval system returns the address of the user's document. Finally, the system returns the storage path of the document selected by the user to the user, achieving efficient document location and access.

[0088] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0089] Example 2

[0090] The present application also provides a file retrieval device. It should be noted that the file retrieval device of the present application can be used to execute the file retrieval method provided by the present application. The file retrieval device provided by the present application is introduced below.

[0091] According to an embodiment of the present application, a device for implementing the above-mentioned file retrieval method is also provided. Figure 5 is a schematic diagram of a file retrieval device provided according to an embodiment of the present application, such as Figure 5 As shown, the device comprises:

[0092] The encoding module 501 is used to receive the description content of the target file uploaded by the user through the client, and perform semantic vectorization encoding on the description content to obtain a description code.

[0093] The retrieval module 502 is used to retrieve all vectorized codes in the vector database that have the same parts as the description code to obtain candidate codes, wherein the vector database stores the semantic vectorized codes of the files in the device drive letter.

[0094] The sending module 503 is used to obtain key information of the file to which the candidate code belongs, and send the key information to the client.

[0095] The instruction receiving module 504 is used to receive a file selection instruction sent by a user through a client, and obtain a target file selected by the user from the file selection instruction.

[0096] The location sending module 505 is used to obtain the file location of the target file from the structured database according to the file identifier of the target file, and send the file location to the client, wherein the structured database stores the structured information of the file in the device drive letter, and the structured information at least includes the file location.

[0097] In an embodiment of the present application, an encoding module 501 receives a description of a target file uploaded by a user through a client, and performs semantic vectorization encoding on the description to obtain a description code; a retrieval module 502 retrieves all vectorized codes that have the same parts as the description code in a vector database to obtain a candidate code, wherein the vector database stores the semantic vectorized codes of files in a device drive letter; a sending module 503 obtains key information of the file to which the candidate code belongs, and sends the key information to the client; an instruction receiving module 504 receives a file selection instruction sent by a user through the client, and obtains a target file selected by the user from the file selection instruction; a location sending module 505 obtains a file location of the target file from a structured database according to a file identifier of the target file, and sends the file location to the client, wherein the structured database stores structured information of the files in the device drive letter, and the structured information at least includes the file location. By using the semantic vectorized codes that describe the content, the semantic vectorized codes of similar files are searched from the vector database. For the target file selected by the user, the file location is determined based on the semantic vectorized code of the target file combined with the structured information, thereby achieving efficient retrieval and positioning of the target file, achieving the technical effect of efficiently finding the target document from a large number of documents stored in the device, and thus solving the problem of low efficiency in users finding specific documents through document search tools in related technologies.

[0098] Optionally, in the file retrieval device provided in the embodiment of the present application, the sending module 503 includes: a calculation module, used to calculate the similarity score between the file to which the candidate code belongs and the description content, and when the number of candidate codes is multiple, calculate the similarity ranking based on the similarity scores associated with the files to which the multiple candidate codes belong; a summary module, used to summarize the files to which the candidate codes belong using a large language model to obtain a document summary of the files to which the candidate codes belong; a key information determination module, used to determine at least one of the similarity score, similarity ranking and document summary as key information.

[0099] Optionally, in the file retrieval device provided in the embodiment of the present application, the calculation module includes: a weight setting module, used to set a quantity weight coefficient and a ranking weight coefficient, wherein the quantity weight coefficient represents the degree of influence of the number of similar file fragments between two files on the similarity score, and the ranking weight represents the degree of influence of the similarity ranking score of a single file fragment on the similarity score; a ranking determination module, used to determine, for a candidate code, the number of similar file fragments between the file to which the candidate code belongs and the description content, determine the similarity ranking score of each similar file fragment, and perform weighted summation of the number of similar file fragments and the similarity ranking score of each similar file fragment based on the quantity weight coefficient and the ranking weight coefficient to obtain the similarity score between the file to which the candidate code belongs and the description content.

[0100] Optionally, in the file retrieval device provided in the embodiment of the present application, the encoding module 501 includes: a scanning module, used to scan the files in the device drive letter, and determine whether the file is an incremental file based on the numerical summary of the file; a vectorized encoding module, used to determine the semantic vectorized encoding of the incremental file when the file is an incremental file, determine the file fragments of the incremental file, and store the vectorized encoding results and the associated file fragments in a vector database; a structured information storage module, used to read the structured information of the incremental file, and store the structured information in a structured database; a dictionary mapping module, used to establish a dictionary mapping between the vector database and the structured database, wherein the dictionary mapping maps the vectorized encoding and structured information of the same file through a file identifier.

[0101] Optionally, in the file retrieval device provided in the embodiment of the present application, the scanning module includes: a judgment module, used to judge whether the numerical summary has appeared in the device drive letter; an incremental file determination module one, used to determine that the file is not an incremental file when the numerical summary has appeared in the device drive letter; an incremental file determination module two, used to determine that the file is an incremental file when the numerical summary has not appeared in the device drive letter.

[0102] Optionally, in the file retrieval device provided in the embodiment of the present application, the vectorized encoding module includes: a parsing module, used to parse the content of the incremental file and map the content of the incremental file to the vector space to obtain a mapping content vector; an extraction module, used to extract deep semantic features of the mapping content vector to obtain a semantic vectorized encoding of the incremental file.

[0103] Optionally, in the file retrieval device provided in the embodiment of the present application, the vectorized encoding module also includes: an incremental file parsing module 1, used to parse the content and structure of the incremental file, wherein the structure includes at least one of the following: title, number of paragraphs and directory; an incremental file parsing module 2, used to split the content according to the structure to obtain file fragments of the incremental file.

[0104] It should be noted that the above modules correspond to steps S201 to S205 in Example 1, and the examples and application scenarios implemented by the modules and corresponding steps are the same, but are not limited to the contents disclosed in the above Example 1. It should be noted that the above modules or units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n), and the above modules may also be part of the device and may be run in the computer terminal 10 provided in Example 1.

[0105] Example 3

[0106] An embodiment of the present application may provide an electronic device, Figure 6 is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 (only one is shown) processor 602, memory 604, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0107] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0108] The processor can call the information and application stored in the memory through the transmission device to execute the following steps: receive the description content of the target file uploaded by the user through the client, and semantically vectorize the description content to obtain a description code; retrieve all vectorized codes that have the same parts as the description code in the vector database to obtain candidate codes, wherein the vector database stores the semantic vectorized codes of the files in the device drive letter; obtain key information of the files to which the candidate codes belong, and send the key information to the client; receive the file selection instruction sent by the user through the client, and obtain the target file selected by the user from the file selection instruction; obtain the file location of the target file from the structured database according to the file identifier of the target file, and send the file location to the client, wherein the structured database stores the structured information of the files in the device drive letter, and the structured information at least includes the file location.

[0109] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: obtaining key information of the file to which the candidate code belongs, including: calculating the similarity score between the file to which the candidate code belongs and the description content, and when there are multiple candidate codes, calculating the similarity ranking based on the similarity scores associated with the multiple candidate code files; calling the large language model to summarize the file to which the candidate code belongs, and obtaining the document summary of the file to which the candidate code belongs; determining at least one of the similarity score, similarity ranking and document summary as key information.

[0110] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: calculating the similarity score between the file to which the candidate code belongs and the description content includes: setting a quantity weight coefficient and a ranking weight coefficient, wherein the quantity weight coefficient represents the degree of influence of the number of similar file fragments between the two files on the similarity score, and the ranking weight represents the degree of influence of the similarity ranking score of a single file fragment on the similarity score; for a candidate code, determining the number of similar file fragments between the file to which the candidate code belongs and the description content, determining the similarity ranking score of each similar file fragment, and performing weighted summation on the number of similar file fragments and the similarity ranking score of each similar file fragment based on the quantity weight coefficient and the ranking weight coefficient to obtain the similarity score between the file to which the candidate code belongs and the description content.

[0111] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: before receiving the description content of the target file uploaded by the user through the client, the method also includes: scanning the files in the device drive letter, and judging whether the file is an incremental file based on the numerical summary of the file; if the file is an incremental file, determining the semantic vectorization encoding of the incremental file, determining the file fragments of the incremental file, and storing the vectorization encoding results and the associated file fragments in the vector database; reading the structured information of the incremental file, and storing the structured information in the structured database; establishing a dictionary mapping between the vector database and the structured database, wherein the dictionary mapping maps the vectorization encoding and structured information of the same file through the file identifier.

[0112] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: determining whether a file is an incremental file based on the numerical summary of the file, including: determining whether the numerical summary has appeared in the device drive letter; if the numerical summary has appeared in the device drive letter, determining that the file is not an incremental file; if the numerical summary has not appeared in the device drive letter, determining that the file is an incremental file.

[0113] The processor can also call the information and applications stored in the memory through the transmission device to perform the following steps: determining the semantic vectorization encoding of the incremental file includes: parsing the content of the incremental file, and mapping the content of the incremental file to the vector space to obtain a mapping content vector; extracting the deep semantic features of the mapping content vector to obtain the semantic vectorization encoding of the incremental file.

[0114] The processor can also call the information and applications stored in the memory through the transmission device to perform the following steps: determining the file fragments of the incremental file includes: parsing the content and structure of the incremental file, wherein the structure includes at least one of the following: title, number of paragraphs and directory; splitting the content according to the structure to obtain file fragments of the incremental file.

[0115] In an embodiment of the present application, the description content of the target file uploaded by the user through the client is received, and the description content is semantically vectorized and encoded to obtain the description code; all vectorized codes in the vector database that have the same part as the description code are retrieved to obtain candidate codes, wherein the vector database stores the semantic vectorized codes of the files in the device drive letter; key information of the files to which the candidate codes belong is obtained, and the key information is sent to the client; a file selection instruction sent by the user through the client is received, and the target file selected by the user is obtained from the file selection instruction; the file location of the target file is obtained from the structured database according to the file identifier of the target file, and the file location is sent to the client, wherein the structured database stores the structured information of the files in the device drive letter, and the structured information at least includes the file location. The semantic vectorized codes of similar files are searched from the vector database through the semantic vectorized codes of the description content, and for the target file selected by the user, the file location is determined based on the semantic vectorized codes of the target file combined with the structured information, so as to realize efficient retrieval and positioning of the target file, and achieve the technical effect of efficiently finding the target document from a large number of documents stored in the device, thereby solving the problem of low efficiency of users finding specific documents through document search tools in the related technology.

[0116] It can be understood by those skilled in the art that Figure 6 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, and other terminal devices. Figure 6 The structure of the electronic device is not limited. Figure 6 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Figure 6 Different configurations are shown.

[0117] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0118] Example 4

[0119] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the file retrieval method provided in the first embodiment.

[0120] Optionally, in this embodiment, the above storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0121] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing the steps of the file retrieval method.

[0122] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0123] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0124] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0125] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0126] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0127] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.

[0128] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A file retrieval method, characterized in that: These include: Receiving the description content of the target file uploaded by the user through the client, and performing semantic vectorization encoding on the description content to obtain a description code; Retrieving all vectorized codes in a vector database that have the same parts as the description code to obtain candidate codes, wherein the vector database stores semantic vectorized codes of files in a device drive letter; Obtain key information of the file to which the candidate code belongs, and send the key information to the client; Receiving a file selection instruction sent by the user through the client, and acquiring a target file selected by the user from the file selection instruction; The file location of the target file is obtained from a structured database according to the file identifier of the target file, and the file location is sent to the client, wherein the structured database stores structured information of the file in the device drive letter, and the structured information at least includes the file location.

2. The file retrieval method according to claim 1, characterized in that: The key information of the file to which the candidate encoding belongs is obtained: Calculating a similarity score between the file to which the candidate code belongs and the description content, and, if there are multiple candidate codes, calculating a similarity ranking based on the similarity scores associated with the files to which the multiple candidate codes belong; Calling a large language model to summarize the file to which the candidate encoding belongs, to obtain a document summary of the file to which the candidate encoding belongs; At least one of the similarity score, the similarity ranking, and the document summary is determined as the key information.

3. The file retrieval method according to claim 2, characterized in that: Calculating the similarity score between the file to which the candidate encoding belongs and the description content includes: Setting a quantity weight coefficient and a ranking weight coefficient, wherein the quantity weight coefficient represents the influence of the number of similar file segments between two files on the similarity score, and the ranking weight represents the influence of the similarity ranking score of a single file segment on the similarity score; For a candidate code, the number of similar file segments between the file to which the candidate code belongs and the description content is determined, the similarity ranking score of each similar file segment is determined, and the number of similar file segments and the similarity ranking score of each similar file segment are weighted and summed based on the number weight coefficient and the ranking weight coefficient to obtain the similarity score between the file to which the candidate code belongs and the description content.

4. The file retrieval method according to claim 1, characterized in that: Before receiving the description content of the target file uploaded by the user through the client, the method further includes: Scan the files in the device drive letter, and determine whether the files are incremental files according to the numerical summary of the files; In the case where the file is the incremental file, determining the semantic vectorized encoding of the incremental file, determining the file fragments of the incremental file, and storing the vectorized encoding results and the associated file fragments in the vector database; Reading the structured information of the incremental file and storing the structured information in a structured database; A dictionary mapping is established between the vector database and the structured database, wherein the dictionary mapping maps the vectorized code and the structured information of the same file through a file identifier.

5. The file retrieval method according to claim 4, characterized in that: Judging whether the file is an incremental file according to the numerical summary of the file includes: Determine whether the numerical summary has appeared in the device drive letter; In the case where the numerical summary has appeared in the device drive letter, determining that the file is not the incremental file; In the case that the numerical summary has never appeared in the device drive letter, it is determined that the file is the incremental file.

6. The file retrieval method according to claim 4, characterized in that: Determining the semantic vectorized encoding of the incremental file includes: Parsing the content of the incremental file, and mapping the content of the incremental file to a vector space to obtain a mapping content vector; The deep semantic features of the mapping content vector are extracted to obtain the semantic vectorized encoding of the incremental file.

7. The file retrieval method according to claim 4, characterized in that: Determining the file fragment of the incremental file includes: Parsing the content and structure of the incremental file, wherein the structure includes at least one of the following: a title, a number of paragraphs, and a table of contents; The content is split according to the structure to obtain file fragments of the incremental file.

8. A file retrieval device, characterized in that: Includes the following: The encoding module is used to receive the description content of the target file uploaded by the user through the client, and perform semantic vectorization encoding on the description content to obtain a description code; A retrieval module is used to retrieve all vectorized codes in a vector database that have the same parts as the description code to obtain candidate codes, wherein the vector database stores semantic vectorized codes of files in a device drive letter; A sending module, used for obtaining key information of the file to which the candidate code belongs, and sending the key information to the client; An instruction receiving module, used for receiving a file selection instruction sent by the user through the client, and obtaining a target file selected by the user from the file selection instruction; A location sending module is used to obtain the file location of the target file from a structured database according to the file identifier of the target file, and send the file location to the client, wherein the structured database stores structured information of the file in the device drive letter, and the structured information at least includes the file location.

9. A computer program product, characterized in that The invention comprises a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the file retrieval method according to any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to execute the file retrieval method according to any one of claims 1 to 7 through the computer program.