File index result determination method and device, equipment and medium

By combining text index and vector index for file indexing, the problem of low retrieval efficiency and accuracy of file index results in the prior art is solved, and more efficient and accurate positioning of file index results is achieved.

CN120086306APending Publication Date: 2025-06-03INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510166702.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the prior art, the search efficiency and accuracy of file index results are low, and it is impossible to locate the institutional key information related to index information.

Method used

By obtaining the information to be indexed and its corresponding target index text library, the key information is indexed in combination with text index and vector index. The specific steps include: determining the text recall result in the first search recall set based on the information to be indexed, and determining the vector recall result in the second search recall set, and finally combining and processing the two to generate the target index result.

Benefits of technology

The search efficiency and accuracy of file index results are improved, and content information related to the information to be indexed can be located.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086306A_ABST
    Figure CN120086306A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for determining a file index result, equipment and a medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: acquiring to-be-indexed information and a target index text library containing a first path of search recall set corresponding to a text index and a second path of search recall set corresponding to a vector index; determining a corresponding text recall result in the first path of search recall set based on the to-be-indexed information, and determining a corresponding vector recall result in the second path of search recall set based on the to-be-indexed information; and performing combined processing on the text recall result and the vector recall result to generate a target index result corresponding to the to-be-indexed information. According to the technical scheme, the content information related to the to-be-indexed information can be positioned, and the retrieval efficiency and accuracy of the file indexing result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and medium for determining a file indexing result. Background Art

[0002] With the rapid development of enterprise applications, the number of institutional documents corresponding to each business type is also increasing. Therefore, in the actual business process, it becomes increasingly difficult to perform file indexing on a large amount of institutional documents.

[0003] In the prior art, a document-level retrieval method is usually adopted for file indexing. However, if the document-level retrieval method is used, only the institutional documents corresponding to the indexing information can be determined, and the key institutional information related to the indexing information cannot be located, reducing the retrieval efficiency and accuracy of the file indexing result. Therefore, how to locate the content information related to the information to be indexed and improve the retrieval efficiency and accuracy of the file indexing result is an urgent problem to be solved at present. Summary of the Invention

[0004] The present invention provides a method, apparatus, device, and medium for determining a file indexing result, which can solve the problem of low retrieval efficiency and accuracy of the file indexing result.

[0005] According to one aspect of the present invention, there is provided a method for determining a file indexing result, including:

[0006] Obtaining the information to be indexed and the target index text library corresponding to the information to be indexed; wherein, the target index text library includes a first-path search recall set corresponding to the text index and a second-path search recall set corresponding to the vector index;

[0007] Determining a corresponding text recall result in the first-path search recall set based on the information to be indexed, and determining a corresponding vector recall result in the second-path search recall set based on the information to be indexed;

[0008] Combining and processing the text recall result and the vector recall result to generate a target index result corresponding to the information to be indexed.

[0009] According to another aspect of the present invention, there is provided a device for determining a file indexing result, including:

[0010] A data acquisition module, configured to obtain the information to be indexed and the target index text library corresponding to the information to be indexed; wherein, the target index text library includes a first-path search recall set corresponding to the text index and a second-path search recall set corresponding to the vector index;

[0011] A search and recall module, configured to determine corresponding text recall results in the first-path search and recall set based on the information to be indexed, and determine corresponding vector recall results in the second-path search and recall set based on the information to be indexed;

[0012] A result generation module, configured to process the text recall results and the vector recall results in combination to generate a target index result corresponding to the information to be indexed.

[0013] According to another aspect of the present invention, there is provided an electronic device, including:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the method for determining a file index result according to any embodiment of the present invention.

[0017] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the method for determining a file index result according to any embodiment of the present invention when executed.

[0018] According to another aspect of the present invention, there is provided a computer program product including a computer program that implements the method for determining a file index result according to any embodiment of the present invention when executed by a processor.

[0019] The technical solution of the embodiment of the present invention obtains the information to be indexed and the target index text library corresponding to the information to be indexed; wherein, the target index text library includes a first-path search and recall set corresponding to text indexing, and a second-path search and recall set corresponding to vector indexing. Furthermore, corresponding text recall results are determined in the first-path search and recall set based on the information to be indexed, and corresponding vector recall results are determined in the second-path search and recall set based on the information to be indexed. Finally, the text recall results and the vector recall results are processed in combination to generate a target index result corresponding to the information to be indexed. Since key information is indexed through two paths of text indexing and vector indexing, the problem of low retrieval efficiency and accuracy of file index results is solved, and content information related to the information to be indexed can be located, improving the retrieval efficiency and accuracy of file index results.

[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0022] Figure 1 is a flowchart of a method for determining a file index result according to Embodiment 1 of the present invention;

[0023] Figure 2 is a flowchart of a method for determining a file index result according to Embodiment 2 of the present invention;

[0024] Figure 3 is a flowchart of an alternative method for determining a file index result according to Embodiment 2 of the present invention;

[0025] Figure 4 is a flowchart of a process for establishing an index text library according to Embodiment 2 of the present invention;

[0026] Figure 5 is a schematic structural diagram of a device for determining a file index result according to Embodiment 3 of the present invention;

[0027] Figure 6 is a schematic structural diagram of an electronic device for implementing the method for determining a file index result of the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] In order to enable those skilled in the art to better understand the solution of the present invention, the following clearly and completely describes the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0029] It should be noted that in the description of the present invention, the terms "first", "second", "target", "basis", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0030] It is worth noting that in the technical solution of this application, the information collected is information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations and standards of relevant countries and regions. Necessary confidentiality measures are taken, which do not violate public order and good customs, and a corresponding operation entry is provided for the user to choose to authorize or reject; if the user chooses to reject, the expert decision-making process will be entered.

[0031] Embodiment 1

[0032] Figure 1 The figure is a flowchart of a method for determining a file index result provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of information indexing of institutional documents. This method can be executed by a device for determining a file index result. The device for determining a file index result can be implemented in the form of hardware and / or software, and the device for determining a file index result can be configured in an electronic device. As Figure 1 shown, the method includes:

[0033] S110. Obtain the information to be indexed and the target index text library corresponding to the information to be indexed; wherein, the target index text library includes a first-path search recall set corresponding to text indexing and a second-path search recall set corresponding to vector indexing.

[0034] Among them, the information to be indexed can refer to the key information that needs to be indexed in the institutional documents. Exemplarily, the information to be indexed can be a keyword or a sentence containing a keyword, etc. The indexed text library can refer to a pre-constructed database containing the indexing relationships of institutional documents. Exemplarily, the indexed text library can contain each content word in the institutional document and the indexing position of the content word in the institutional document. For example, the indexing position can be the unique encoding (Identity Document, ID) of the institutional document, the page number and the paragraph number, etc. The target indexed text library can refer to the indexed text library that matches the information to be indexed. Usually, the target indexed text libraries corresponding to different information to be indexed are different, and the corresponding target indexed text library can be determined according to the actual business type.

[0035] Among them, the text index can refer to the indexing relationship for determining the positions of the original words in the institutional document. The first-pass search recall can refer to the set composed of each text index corresponding to the same institutional document. The first-pass search recall set can refer to the set composed of each first-pass search recall corresponding to the same information to be indexed. Usually, the number of first-pass search recalls in the first-pass search recall set is at least one.

[0036] Among them, the vector index can refer to the indexing relationship for determining the positions of the relevant semantic words in the institutional document. The second-pass search recall can refer to the set composed of each vector index corresponding to the same institutional document. The second-pass search recall set can refer to the set composed of each second-pass search recall corresponding to the same information to be indexed. Usually, the number of second-pass search recalls in the second-pass search recall set is at least one.

[0037] S120. Determine the corresponding text recall result in the first-pass search recall set based on the information to be indexed, and determine the corresponding vector recall result in the second-pass search recall set based on the information to be indexed.

[0038] Among them, the text recall result can refer to the indexing result of file indexing in the first-pass search recall set according to the information to be indexed. Exemplarily, the text recall result can be the indexing position information of the information to be indexed in the institutional document.

[0039] Among them, the vector recall result can refer to the indexing result of file indexing in the second-pass search recall set according to the information to be indexed. Exemplarily, the vector recall result can be the indexing position information of the relevant semantic words of the information to be indexed in the institutional document.

[0040] S130. Combine and process the text recall result and the vector recall result to generate the target indexing result corresponding to the information to be indexed.

[0041] Among them, the target index result can refer to the content result of the institutional document that is located and displayed based on the text recall result and the vector recall result. Exemplarily, the target index result can include information such as a content summary, keywords, and institutional documents.

[0042] In the technical solution of the embodiment of the present invention, by obtaining the information to be indexed and the target index text library corresponding to the information to be indexed; wherein, the target index text library includes a first-way search recall set corresponding to the text index and a second-way search recall set corresponding to the vector index. Furthermore, based on the information to be indexed, the corresponding text recall result is determined in the first-way search recall set, and based on the information to be indexed, the corresponding vector recall result is determined in the second-way search recall set. Finally, the text recall result and the vector recall result are combined and processed to generate the target index result corresponding to the information to be indexed. Since the key information is indexed through two ways of text index and vector index, the problem of low retrieval efficiency and accuracy of the file index result is solved, and the content information related to the information to be indexed can be located, improving the retrieval efficiency and accuracy of the file index result.

[0043] Embodiment Two

[0044] Figure 2 It is a flowchart of a method for determining a file index result provided by the second embodiment of the present invention. This embodiment is an addition based on the above embodiment. Specifically, in this embodiment, the operation before obtaining the information to be indexed and the target index text library corresponding to the information to be indexed is added. Specifically, it may further include: obtaining a preset business type and the full amount of basic institutional documents corresponding to the preset business type; formatting the basic institutional documents to obtain target institutional documents; establishing a text index for the target institutional documents based on a preset full-text retrieval model to generate a first-way search recall corresponding to the target institutional documents, and establishing a vector index for the target institutional documents based on a preset text encoding model to generate a second-way search recall corresponding to the target institutional documents; combining and processing the first-way search recall and the second-way search recall corresponding to the target institutional documents to generate a basic index text corresponding to the basic institutional documents; combining and processing the basic index texts corresponding to the full amount of basic institutional documents under the same preset business type to obtain a basic index text library corresponding to the preset business type. As Figure 2 shown, the method includes:

[0045] S210. Obtain a preset business type and the full amount of basic institutional documents corresponding to the preset business type.

[0046] Among them, the preset business type can refer to the business processing type applied by an enterprise set in advance. Exemplarily, the preset business type can be a task assignment type or an item transfer type, etc. Usually, the preset business type corresponding to each enterprise application can be determined according to the actual business scenario.

[0047] Among them, the basic system file may refer to the unprocessed original system file corresponding to the preset business type. Exemplarily, the basic system file may be a normative document containing the specific implementation process for realizing the preset business type. Generally, the basic system files corresponding to different preset business types are different, and all the basic system files corresponding to the preset business type can be determined in the document library according to the preset business type.

[0048] S220. Format the basic system file to obtain a target system file.

[0049] Among them, the formatting process may refer to an operation of initializing the data of the basic system file. Exemplarily, the formatting process may include format conversion and data storage, etc. The target system file may refer to the system file obtained after formatting the basic system file.

[0050] In an optional implementation manner, the formatting the basic system file to obtain a target system file includes: converting and processing the basic system file based on a preset format type to generate a candidate system file corresponding to the basic system file; storing the candidate system file based on a preset storage format to generate a target system file corresponding to the basic system file.

[0051] Among them, the preset format type may refer to a preset display format of the system file. Exemplarily, the preset format type may be HyperText Markup Language (HTML). The candidate system file may refer to the system file obtained by converting and processing the basic system file according to the preset format type. The preset storage format may refer to a preset storage format of the system file. Exemplarily, the preset storage format may be an encoding method of converting binary data into American Standard Code for Information Interchange (ASCII) strings.

[0052] Specifically, after obtaining the basic system file, the basic system file can be converted and processed using the preset format type to generate a candidate system file of the preset format type corresponding to the basic system file. Furthermore, information such as images or formulas in the candidate system file is stored according to the preset storage format to generate a target system file corresponding to the basic system file. Thus, by performing format conversion and storage on the basic system file, it can provide an effective basis for identifying the position of the returned information in subsequent operations and also provide effective convenience for data storage.

[0053] S230. Establish a text index for the target institutional document based on a preset full-text retrieval model to generate the first-path search recall corresponding to the target institutional document, and establish a vector index for the target institutional document based on a preset text encoding model to generate the second-path search recall corresponding to the target institutional document.

[0054] Among them, the preset full-text retrieval model can refer to a model preset for indexing all fields of institutional documents. Usually, the preset full-text retrieval model can perform word segmentation on institutional documents, and then, based on each word segmentation result, perform indexing to obtain a text index. Exemplarily, the preset full-text retrieval model can be a full-text retrieval engine, etc.

[0055] Among them, the preset text encoding model can refer to a model preset for indexing the full-text semantic information of institutional documents. Usually, the preset text encoding model can convert each word in the institutional document into a one-dimensional vector, and then, fuse the one-dimensional vectors with the full-text semantic information to generate a vector index. It should be noted that the vector index generated based on the preset text encoding model can include: a text vector used to depict the global semantic information of the text and fuse with the semantic information of single words or terms, a word vector used to represent the original fields in the institutional document, and a position vector used to represent the position of the word vector in the institutional document. Exemplarily, the preset text encoding model can be a bidirectional text encoding pre-training model.

[0056] S240. Combine and process the first-path search recall and the second-path search recall corresponding to the target institutional document to generate a basic index text corresponding to the basic institutional document.

[0057] Among them, the basic index text can refer to the index text initially generated after processing the target institutional document based on the preset full-text retrieval model and the preset text encoding model. Usually, one basic index text corresponds to one target institutional document, and this basic index text includes the first-path search recall and the second-path search recall corresponding to the target institutional document. Exemplarily, the basic index text can be the inverted index relationship between paragraphs and documents.

[0058] S250. Combine and process the basic index texts corresponding to all basic institutional documents under the same preset business type to obtain a basic index text library corresponding to the preset business type.

[0059] Specifically, before indexing the system documents, the preset business types and corresponding full-volume basic system documents can be determined first, and the basic system documents can be formatted to generate target system documents. Furthermore, a text index is established for the target system documents based on the preset full-text retrieval model to generate the first path search recall corresponding to the target system documents, and a vector index is established for the target system documents based on the preset text encoding model to generate the second path search recall corresponding to the target system documents. Further, the first path search recall and the second path search recall corresponding to the target system documents are combined to generate the basic index text corresponding to the basic system documents. Finally, the basic index texts corresponding to the full-volume basic system documents under the same preset business type are combined to obtain the basic index text library corresponding to the preset business type. Thus, by processing the basic system documents through the full-text retrieval model and the text encoding model, an index text library containing text indexes and vector indexes is pre-constructed, which can provide an effective basis for the subsequent indexing of system documents and improve the indexing efficiency of system documents.

[0060] It should be noted that in the embodiments of the present invention, if the quantity of the basic system documents corresponding to the same enterprise application does not increase, after generating the basic index text library once, the subsequent file indexing process can directly utilize the pre-constructed basic index text library without constructing the basic index text library again. The embodiments of the present invention do not give additional elaboration on this.

[0061] S260. Obtain the information to be indexed, and determine the target business type corresponding to the information to be indexed based on the preset business type division standard.

[0062] Among them, the preset business type division standard can refer to the rules preset for dividing the information to be indexed into business types. Exemplarily, the preset business type division standard can include each keyword field corresponding to the same business type. The business type corresponding to the information to be indexed can be obtained by matching through the preset business type division standard. The target business type can refer to the business type corresponding to the information to be indexed.

[0063] Specifically, if the preset business type division standard is: dividing keyword fields such as taking a taxi, taxi, or hailed into the task allocation type, and the information to be indexed is taking a taxi, then the task allocation type can be used as the target business type.

[0064] S270. Perform type matching in each basic index text library based on the target business type, and use the basic index text library that matches the target business type as the target index text library.

[0065] Among them, the target index text library includes the first path search recall set corresponding to the text index and the second path search recall set corresponding to the vector index.

[0066] Specifically, when indexing institutional documents, the information to be indexed can be obtained first, and the target business type corresponding to the information to be indexed can be determined based on a preset business type classification standard. Then, the target business type is used to perform type matching in each basic index text library to obtain the target index text library that matches the target business type. Thus, the scope of institutional documents can be narrowed down by the business type of the information to be indexed, improving the indexing efficiency of institutional documents.

[0067] S280. Determine the corresponding text recall result in the first-path search recall set based on the information to be indexed, and determine the corresponding vector recall result in the second-path search recall set based on the information to be indexed.

[0068] Specifically, after obtaining the target index text library corresponding to the information to be indexed, the information to be indexed can be used to perform indexing in the first-path search recall set to determine the corresponding text recall result, and the information to be indexed can be used to perform indexing in the second-path search recall set to determine the corresponding vector recall result. Exemplarily, if the information to be indexed is "taking a taxi", "taking a taxi" can be used to perform indexing in the first-path search recall set to determine the text recall result corresponding to "taking a taxi", and "taking a taxi" can be used to perform indexing in the second-path search recall set to determine the vector recall result corresponding to "taking a taxi". Thus, when there is no text recall result corresponding to "taking a taxi" in the first-path search recall set, the vector recall result of the word vector with the same semantic information as "taking a taxi" can be determined in the second-path search recall set according to the correlation relationship between the text vector and the word vector, improving the accuracy of the file indexing result.

[0069] S290. Merge and process the text recall result and the vector recall result to generate a merged recall result.

[0070] Among them, the merged recall result can refer to the recall result obtained by merging and processing the text recall result and the vector recall result corresponding to the same information to be indexed. Generally, the merged recall result contains each text recall result and vector recall result corresponding to the same information to be indexed.

[0071] S2100. Sort the merged recall result according to a preset sorting rule to obtain a target recall result that meets the preset screening criteria.

[0072] Among them, the preset sorting rule can refer to a rule preset for sorting each recall result in the merged recall results. Exemplarily, the preset sorting rule can be to sort in descending order according to the paragraph occurrence frequency. The preset screening criterion can refer to a rule preset for screening the sorting results. Exemplarily, the preset screening criterion can be a rule for selecting the merged recall result with the maximum occurrence frequency, or a rule for successively selecting all the merged recall results according to the occurrence frequency. The target recall result can refer to the recall result selected as the final indexed content.

[0073] S2110. Perform a reverse search for the target recall result to obtain the target index result corresponding to the information to be indexed.

[0074] Among them, the reverse search can refer to an operation of performing a reverse index on the target recall result. Exemplarily, the reverse search can be an operation of locating the target recall result in the institutional document according to the position information in the target recall result. Usually, the information to be indexed is generated in a specified institutional document. Therefore, after obtaining the target recall result, the target recall result can be reversely searched in the specified institutional document.

[0075] Specifically, after determining the corresponding text recall result in the first-way search recall set based on the information to be indexed and determining the corresponding vector recall result in the second-way search recall set based on the information to be indexed, the text recall result and the vector recall result can be merged to generate a merged recall result. Then, the merged recall result is sorted according to the preset sorting rule to obtain the target recall result that meets the preset screening criterion. Finally, a reverse search is performed on the target recall result in the institutional document to obtain the target index result corresponding to the information to be indexed. Thus, by using the reverse search method to determine the file index result, the file index result in the institutional document can be accurately located, improving the accuracy of the file index result.

[0076] S2120. Perform a highlighting rendering on the target index result based on a preset rendering rule to obtain the rendering result corresponding to the target index result.

[0077] Among them, the preset rendering rule can refer to a rule preset for rendering and displaying the index result. Exemplarily, the preset rendering rule can be a rule for rendering the abstract or keywords in the target index result. The highlighting rendering can refer to an operation of highlighting the index result based on a prominent color. Exemplarily, the highlighting rendering can be to highlight the index result with yellow. The rendering result can refer to the display result generated after performing a highlighting rendering on the target index result according to the preset rendering rule.

[0078] Specifically, after generating the target index result corresponding to the information to be indexed, the target index result can be highlighted and rendered using a preset rendering rule to obtain the rendering result corresponding to the target index result. Thus, by highlighting the file index result, it is convenient for the user to quickly locate multiple relevant content segments in the institutional document, providing an effective basis for subsequent operations.

[0079] In the technical solution of the embodiment of the present invention, the full-volume basic institutional document corresponding to the preset business type is formatted to obtain the target institutional document, and a text index is established for the target institutional document based on the preset full-text retrieval model to generate the first-pass search recall corresponding to the target institutional document. A vector index is established for the target institutional document based on the preset text encoding model to generate the second-pass search recall corresponding to the target institutional document. Furthermore, the first-pass search recall and the second-pass search recall corresponding to the target institutional document are combined and processed to generate the basic index text corresponding to the basic institutional document, and the basic index texts corresponding to the full-volume basic institutional documents under the same preset business type are combined and processed to obtain the basic index text library corresponding to the preset business type. Further, the information to be indexed is obtained, and the target business type corresponding to the information to be indexed is determined based on the preset business type classification standard. Type matching is performed in each basic index text library based on the target business type, and the basic index text library that matches the target business type is used as the target index text library. Further, based on the information to be indexed, the corresponding text recall result is determined in the first-pass search recall set, and the corresponding vector recall result is determined in the second-pass search recall set. The text recall result and the vector recall result are combined and processed to generate a combined recall result. The combined recall result is sorted according to the preset sorting rule to obtain the target recall result that meets the preset screening criteria, and the target recall result is searched backward to obtain the target index result corresponding to the information to be indexed. Finally, the target index result is highlighted and rendered based on the preset rendering rule to obtain the rendering result corresponding to the target index result. Since the key information is indexed through two paths of text index and vector index, the problem of low retrieval efficiency and accuracy of the file index result is solved, and the content information related to the information to be indexed can be located, improving the retrieval efficiency and accuracy of the file index result.

[0080] Figure 3The figure shows a flowchart of an optional method for determining a file index result provided by an embodiment of the present invention. Specifically, when receiving information to be indexed, the target business type corresponding to the information to be indexed can be determined first based on a preset business type division criterion, and type matching can be performed in each basic index text library based on the target business type, and the basic index text library that matches the target business type is used as the target index text library. Then, the corresponding text recall result is determined in the first-path search recall set of the target index text library based on the information to be indexed, and the corresponding vector recall result is determined in the second-path search recall set of the target index text library based on the information to be indexed. Further, the text recall result and the vector recall result are combined and processed to generate a combined recall result, and the combined recall result is sorted according to a preset sorting rule to obtain a target recall result that meets the preset screening criterion, and the target recall result is searched backward to obtain the target index result corresponding to the information to be indexed. Finally, the target index result is highlighted and rendered based on a preset rendering rule to obtain a rendering result corresponding to the target index result. Thus, the process of determining the file index result is completed.

[0081] Figure 4 The figure shows a flowchart of a process for establishing an index text library provided by an embodiment of the present invention. Specifically, first, a preset business type is obtained, and all basic institutional documents corresponding to the preset business type are determined in a document library. Then, the basic institutional documents are formatted to obtain target institutional documents. Further, a text index is established for the target institutional documents based on a preset full-text retrieval model to generate a first-path search recall corresponding to the target institutional documents, and a vector index is established for the target institutional documents based on a preset text encoding model to generate a second-path search recall corresponding to the target institutional documents. Finally, the first-path search recall and the second-path search recall corresponding to the target institutional documents are combined and processed to generate a basic index text corresponding to the basic institutional documents, and the basic index texts corresponding to all basic institutional documents under the same preset business type are combined and processed to obtain a basic index text library corresponding to the preset business type.

[0082] Embodiment III

[0083] Figure 5 The figure shows a structural schematic diagram of a device for determining a file index result provided by Embodiment III of the present invention. As Figure 5 shown, the device includes: a data acquisition module 310, a search recall module 320, and a result generation module 330;

[0084] Among them, the data acquisition module 310 is configured to acquire information to be indexed and the target index text library corresponding to the information to be indexed; wherein, the target index text library includes a first-path search recall set corresponding to a text index and a second-path search recall set corresponding to a vector index;

[0085] A search and recall module 320, configured to determine a corresponding text recall result in the first-path search and recall set based on the information to be indexed, and determine a corresponding vector recall result in the second-path search and recall set based on the information to be indexed;

[0086] A result generation module 330, configured to combine and process the text recall result and the vector recall result to generate a target index result corresponding to the information to be indexed.

[0087] In the technical solution of the embodiment of the present invention, by obtaining the information to be indexed and the target index text library corresponding to the information to be indexed; wherein, the target index text library includes a first-path search and recall set corresponding to the text index, and a second-path search and recall set corresponding to the vector index. Furthermore, a corresponding text recall result is determined in the first-path search and recall set based on the information to be indexed, and a corresponding vector recall result is determined in the second-path search and recall set based on the information to be indexed. Finally, the text recall result and the vector recall result are combined and processed to generate a target index result corresponding to the information to be indexed. Since the key information is indexed through two paths of text index and vector index, the problem of low retrieval efficiency and accuracy of the file index result is solved, and the content information related to the information to be indexed can be located, improving the retrieval efficiency and accuracy of the file index result.

[0088] Optionally, the device for determining the file index result may further include: an index text library establishment module, configured to, before obtaining the information to be indexed and the target index text library corresponding to the information to be indexed, obtain a preset service type and the full-volume basic system files corresponding to the preset service type; format the basic system files to obtain target system files; establish a text index for the target system files based on a preset full-text retrieval model to generate a first-path search and recall corresponding to the target system files, and establish a vector index for the target system files based on a preset text encoding model to generate a second-path search and recall corresponding to the target system files; combine and process the first-path search and recall and the second-path search and recall corresponding to the target system files to generate a basic index text corresponding to the basic system files; combine and process the basic index texts corresponding to the full-volume basic system files under the same preset service type to obtain a basic index text library corresponding to the preset service type.

[0089] Optionally, the index text library establishment module may specifically be configured to:

[0090] Perform format type conversion processing on the basic system files based on a preset format type to generate candidate system files corresponding to the basic system files;

[0091] Perform file storage on the candidate system files based on a preset storage format to generate target system files corresponding to the basic system files.

[0092] Optionally, the data acquisition module 310 may specifically be configured to:

[0093] Obtain the information to be indexed, and determine the target service type corresponding to the information to be indexed based on a preset service type division criterion;

[0094] Perform type matching in each basic index text library based on the target service type, and use the basic index text library that matches the target service type as the target index text library.

[0095] Optionally, the result generation module 330 may specifically be configured to:

[0096] Merge and process the text recall result and the vector recall result to generate a merged recall result;

[0097] Sort the merged recall result according to a preset sorting rule to obtain a target recall result that meets the preset screening criterion;

[0098] Inversely search for the target recall result to obtain the target index result corresponding to the information to be indexed.

[0099] Optionally, the apparatus for determining a file index result may further include: a result rendering module, configured to, after combining and processing the text recall result and the vector recall result to generate the target index result corresponding to the information to be indexed, perform highlight rendering on the target index result based on a preset rendering rule to obtain a rendering result corresponding to the target index result.

[0100] The apparatus for determining a file index result provided by an embodiment of the present invention may execute the method for determining a file index result provided by any embodiment of the present invention, and has function modules and beneficial effects corresponding to the execution of the method.

[0101] Embodiment 4

[0102] Figure 6 FIG. shows a schematic structural diagram of an electronic device 410 that may be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0103] As Figure 6As shown, the electronic device 410 includes at least one processor 420 and a memory communicatively connected to the at least one processor 420, such as read-only memory (ROM) 430, random access memory (RAM) 440, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 420 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 430 or the computer program loaded from the storage unit 490 into the random access memory (RAM) 440. In the RAM 440, various programs and data required for the operation of the electronic device 410 can also be stored. The processor 420, ROM 430, and RAM 440 are connected to each other through a bus 450. The input / output (I / O) interface 460 is also connected to the bus 450.

[0104] Multiple components in the electronic device 410 are connected to the I / O interface 460, including: an input unit 470, such as a keyboard, a mouse, etc.; an output unit 480, such as various types of displays, speakers, etc.; a storage unit 490, such as a disk, an optical disc, etc.; and a communication unit 4100, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 4100 allows the electronic device 410 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0105] The processor 420 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 420 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 420 executes the various methods and processes described above, such as the method for determining the file index result.

[0106] The method includes:

[0107] Obtain the information to be indexed and the target index text library corresponding to the information to be indexed; wherein, the target index text library includes a first-way search recall set corresponding to the text index and a second-way search recall set corresponding to the vector index;

[0108] Based on the information to be indexed, determine the corresponding text recall result in the first-way search recall set, and based on the information to be indexed, determine the corresponding vector recall result in the second-way search recall set;

[0109] Combined process the text recall result and the vector recall result to generate the target index result corresponding to the information to be indexed.

[0110] In some embodiments, the method for determining a file indexing result may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 490. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 410 via the ROM 430 and / or the communication unit 4100. When the computer program is loaded into the RAM 440 and executed by the processor 420, one or more steps of the method for determining a file indexing result described above may be performed. Alternatively, in other embodiments, the processor 420 may be configured to execute the method for determining a file indexing result by any other suitable means (e.g., by means of firmware).

[0111] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0112] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0113] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0114] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0115] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0116] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0117] The embodiments of the present application also disclose a computer program product, which includes a computer program that, when executed by a processor, implements the method for determining a file index result provided in any embodiment of the present application. This program product and the method for determining a file index result disclosed in each embodiment of the present application belong to the same inventive concept, and thus will not be elaborated herein.

[0118] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitations are imposed herein.

[0119] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for determining a file index result, characterized in that: include: Acquire information to be indexed and a target index text library corresponding to the information to be indexed; wherein the target index text library includes a first search recall set corresponding to a text index and a second search recall set corresponding to a vector index; Determine a corresponding text recall result in the first search recall set based on the information to be indexed, and determine a corresponding vector recall result in the second search recall set based on the information to be indexed; The text recall result and the vector recall result are combined and processed to generate a target index result corresponding to the information to be indexed.

2. The method according to claim 1, characterized in that Before obtaining the information to be indexed and the target index text library corresponding to the information to be indexed, the method further includes: Obtain the preset business type and the full set of basic system documents corresponding to the preset business type; Formatting the basic system file to obtain a target system file; Establishing a text index for the target system file based on a preset full-text retrieval model to generate a first search recall corresponding to the target system file, and establishing a vector index for the target system file based on a preset text encoding model to generate a second search recall corresponding to the target system file; Combine and process the first search recall and the second search recall corresponding to the target system file to generate a basic index text corresponding to the basic system file; The basic index texts corresponding to the full amount of basic system documents under the same preset business type are combined and processed to obtain a basic index text library corresponding to the preset business type.

3. The method according to claim 2, characterized in that The formatting process of the basic system file to obtain the target system file includes: The basic system file is converted and processed based on a preset format type to generate a candidate system file corresponding to the basic system file; The candidate system file is stored based on a preset storage format to generate a target system file corresponding to the basic system file.

4. The method according to claim 2, characterized in that: The obtaining of the information to be indexed and the target index text library corresponding to the information to be indexed includes: Acquire information to be indexed, and determine a target business type corresponding to the information to be indexed based on a preset business type classification standard; Type matching is performed in each basic index text library based on the target service type, and the basic index text library matching the target service type is used as the target index text library.

5. The method according to claim 1, characterized in that The combining and processing the text recall result and the vector recall result to generate a target index result corresponding to the information to be indexed includes: Merging the text recall result and the vector recall result to generate a merged recall result; Sorting the combined recall results according to preset sorting rules to obtain target recall results that meet preset screening criteria; The target recall result is searched in reverse order to obtain the target index result corresponding to the information to be indexed.

6. The method according to claim 1, characterized in that After the combined processing of the text recall result and the vector recall result to generate a target index result corresponding to the information to be indexed, the method further includes: The target index result is highlighted and rendered based on a preset rendering rule to obtain a rendering result corresponding to the target index result.

7. A device for determining a file index result, characterized in that: include: A data acquisition module, used to acquire information to be indexed and a target index text library corresponding to the information to be indexed; wherein the target index text library includes a first search recall set corresponding to a text index and a second search recall set corresponding to a vector index; A search and recall module, configured to determine a corresponding text recall result in the first search and recall set based on the information to be indexed, and to determine a corresponding vector recall result in the second search and recall set based on the information to be indexed; The result generation module is used to combine and process the text recall result and the vector recall result to generate a target index result corresponding to the information to be indexed.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for determining the file index result according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for determining a file index result according to any one of claims 1 to 6 when executed.

10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method for determining a file index result according to any one of claims 1 to 6.